Information processing device, information processing method, and information processing program
The information processing system enhances user interactions by supplementing input with situational context to guide an LLM, ensuring relevant responses are obtained from other users, addressing the limitations of conventional dialogue agents in time-constrained situations.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PIONEER IP
- Filing Date
- 2025-08-05
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional technologies fail to allow users to ask questions to other users without much effort, particularly in situations where time is limited, such as while driving, and often result in irrelevant or general responses from dialogue agents.
An information processing system that supplements user input with additional information based on situational context, generates an instruction statement to guide a large language model (LLM) to request a response from another user, and controls the output of this request to ensure relevance and accuracy.
Enables users to obtain targeted and relevant responses from other users without significant effort, improving the quality of information received during limited-time scenarios like driving.
Smart Images

Figure JP2025027808_07052026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, and Information Processing Program
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
[0002] Conventionally, techniques for controlling communication between users have been proposed. For example, Patent Document 1 discloses a technique for more quickly obtaining information desired by a user using a machine learning model in a system where users conduct question-and-answer sessions with each other.
[0003] Japanese Unexamined Patent Application Publication No. 2016-24600
[0004] For example, there is a known dialogue agent system that conducts various daily conversations while realizing driving support according to the situation and providing content such as route guidance through speech. Therefore, a user may input a question and request an answer to the input from the dialogue agent system.
[0005] Here, when the user is driving, there may not be enough time to consider an appropriate question that is easy for the dialogue agent system to understand, and the user may desire to be able to ask a question without much effort.
[0006] In addition, the dialogue agent system may be realized, for example, as a voice assist function by a chatbot equipped with a large language model (LLM: Large Language Models). While it is good at providing general answers to questions, some users may desire answers that only an actual person, other than themselves, can give, rather than such general answers.
[0007] From the above, there is a need to be able to ask questions to other users without much effort, but in the above conventional technologies, it is not always possible to ask questions to other users without much effort.
[0008] For example, the conventional technology described above automatically generates a question tailored to the user's current situation using learning results based on questions previously sent to other users and the circumstances under which those questions were sent, and then sends that question to other users. However, in order to improve learning accuracy, it is necessary to use questions with accurate content from among the questions that the user has previously sent to other users. For this reason, the conventional technology described above still has the problem of not allowing users to ask questions without any effort.
[0009] The present invention has been made in view of the above, and provides an information processing device, an information processing method, and an information processing program that can ask questions to other users without taking any effort.
[0010] The information processing device according to claim 1 comprises: an acquisition unit that acquires input information relating to a user's question or inquiry; a setting unit that sets additional information to be added to the input information based on situation information indicating the user's movement; a generation unit that generates an instruction statement based on the input information and the additional information, which instructs the output of a sentence requesting a response from another user; and a control unit that controls the provision of the request statement output by the generation model in response to the instruction statement to the other user.
[0011] The information processing method according to claim 10 is an information processing method performed by an information processing device, comprising: an acquisition step of acquiring input information relating to a user's question or inquiry; a setting step of setting additional information to be added to the input information based on situation information indicating the user's movement; a generation step of generating an instruction statement based on the input information and the additional information, which instructs the output of a sentence requesting a response from another user; and a control step of controlling the output of the request statement generated by the generation model in response to the instruction statement so that it is provided to the other user.
[0012] The information processing program according to claim 11 is an information processing program executed by an information processing device, wherein the information processing device is made to execute: an acquisition procedure for acquiring input information relating to a user's question or inquiry; a setting procedure for setting additional information to be added to the input information based on situation information indicating the situation relating to the user's movement; a generation procedure for generating an instruction statement based on the input information and the additional information, which instructs the output of a sentence requesting a response from another user; and a control procedure for controlling the provision of the request statement output by the generation model in response to the instruction statement to the other user.
[0013] Figure 1 is a diagram showing an example configuration of an information processing system according to Embodiment 1. Figure 2 is a diagram showing an example device configuration according to Embodiment 1. Figure 3 is a sequence diagram showing the information processing procedure according to Embodiment 1. Figure 4 is a flowchart showing the detailed information processing procedure according to Embodiment 1. Figure 5 is a diagram showing an example configuration of an information processing system according to Embodiment 2. Figure 6 is a diagram showing an example device configuration according to Embodiment 2. Figure 7 is a sequence diagram (1) showing the information processing procedure according to Embodiment 2. Figure 8 is a sequence diagram (2) showing the information processing procedure according to Embodiment 2. Figure 9 is a diagram showing specific examples of instruction statements, summarization conditions, and summarization content according to Embodiment 2. Figure 10 is a flowchart showing a method for providing summary content. Figure 11 is a flowchart showing a filtering method. Figure 12 is a sequence diagram showing a search instruction procedure according to Embodiment 2. Figure 13 is a flowchart showing an alternative method. Figure 14 is a hardware configuration diagram showing an example of a computer that realizes the functions of a server device according to the embodiment.
[0014] [Embodiments] Embodiments of the proposed technology of the present invention will be described in detail below with reference to the attached drawings. In this specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant explanations will be omitted.
[0015] The one or more embodiments (including examples, modifications, and applications) described below can each be implemented independently. On the other hand, at least some of the embodiments described below may be implemented in appropriate combination with at least some of the other embodiments. These embodiments may contain novel features that differ from each other. Therefore, these embodiments may contribute to solving different objectives or problems and may produce different effects.
[0016] Furthermore, in the following embodiment, the proposed technology of the present invention will be explained using as an example a scenario in which a user ("User U") inputs a question to another user ("Other User Ux") while driving a vehicle ("Vehicle VE"). However, the scenes to which the proposed technology of the present invention can be applied are not limited to driving scenes.
[0017] Furthermore, in the following embodiments, a Large-Scale Language Model (LLM) is given as an example of a generative model, and the instructional statements used in the LLM refer to prompts. In addition, other users Ux are an unspecified number of users from the perspective of user U, and the expression "response from other users Ux" can be changed to "answer from other users Ux". Note that other users Ux may also include people within user U's circle, or groups that share common hobbies with user U.
[0018] Furthermore, the proposed technology of the present invention includes both a means to solve the problem of asking questions to other users Ux without effort, and a means to easily grasp the content of the answers obtained from other users Ux.Therefore, one embodiment in which the former solution is realized will be described as <Embodiment 1>, and one embodiment in which the latter solution is realized will be described as <Embodiment 2>.When there is no need to distinguish between Embodiment 1 and Embodiment 2, they will simply be referred to as embodiments.
[0019] <Embodiment 1> [1. Introduction] A voice assistance function (speech agent system) using a chatbot equipped with a large-scale language model (LLM) makes it possible to support drivers through dialogue with them. For example, if user U, who is a driver, has questions or concerns about various events that occur while driving, route guidance, tourist information, etc., they may input their questions or concerns to the speech agent system, for example, by voice.
[0020] However, user U, while driving, often lacks the capacity to think of precise content that the speech agent system can easily understand, and can only input simple questions or doubts. In such cases, the LLM's output (LLM's response) to such input may be irrelevant. A similar situation can occur with other users Ux.
[0021] For example, while user U interacts with a speech agent system, they may want niche information that only another user Ux, a real human being, would know. However, if the content of the question or inquiry is inappropriate, the answer from other user Ux may be irrelevant.
[0022] Therefore, the inventors of the present invention conceived the idea that if the LLM is input with conditional information that supplements the input information (input information relating to a question or inquiry) entered by user U, and an instruction sentence that instructs the LLM to output a sentence requesting a response from another user Ux based on the input information, the LLM will be able to generate an accurate request sentence that is in line with user U's intentions.
[0023] In other words, according to the proposed technology of Embodiment 1, conditional information is set that supplements the input information (input information related to a question or inquiry) entered by user U, based on the user U's situation. Furthermore, according to the proposed technology of Embodiment 1, an instruction statement is generated as an instruction statement based on the conditional information and the input information, instructing the output of a request statement requesting a response from another user Ux. This instruction statement is input into LLM (an example of a generation model).
[0024] Furthermore, according to the proposed technology of Embodiment 1, the request statement output by the LLM in response to the instruction statement is controlled to be provided to other users' Ux.
[0025] [2. System Configuration] The configuration of the system according to the embodiment will be explained using Figure 1. Figure 1 is a diagram showing an example configuration of the information processing system 1A according to Embodiment 1. As shown in Figure 1, the information processing system 1A includes a user device 10, a chatbot device 50, and a server device 100A. The user device 10, the chatbot device 50, and the server device 100A may be connected to each other via a predetermined communication network (network N) by wired or wireless means. The voice agent system may consist of the chatbot device 50 and the server device 100A.
[0026] The user device 10 may be an information processing terminal used by user U who drives the vehicle VE. For example, the user device 10 may be a smartphone, a wearable device, a tablet device, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), etc.
[0027] As another example, the user device 10 may be implemented as a navigation device, i.e., an in-vehicle device, that is built into or mounted in the vehicle's VE. The user device 10 as an in-vehicle device may have not only a navigation function but also a recording function (drive recorder function). Figure 1 shows an example in which the user device 10 is an in-vehicle device.
[0028] The chatbot device 50 enables dialogue with user U. The chatbot device 50 has the function of generating response information to input information by being equipped with an LLM. For example, in the chatbot device 50, the LLM is repeatedly trained and the LLM is updated in order to achieve high-precision dialogue with user U. In the information processing according to this embodiment, the chatbot device 50 inputs the instruction text generated by the server device 100A into the LLM, causing it to output a request text that requests a response from another user Ux.
[0029] Furthermore, the chatbot device 50 may be equipped with so-called generative AI (artificial intelligence capable of creating various content and ideas such as conversations, stories, images, videos, and music). Also, while the user device 10 is an edge computer, the chatbot device 50 can be implemented as a cloud computer.
[0030] The server device 100 according to this embodiment is an example of an information processing device. More specifically, the server device 100A is a central information processing device that realizes the information processing according to Embodiment 1.
[0031] Therefore, the server device 100A according to Embodiment 1 acquires input information IN relating to a question or inquiry from user U, and sets additional information ADD to be added to the input information IN based on the situation information indicating the situation regarding user U's movement. The server device 100A then generates an instruction statement PP based on the input information IN and the additional information ADD, which instructs the output of a text requesting a response from another user Ux. The server device 100A then controls the LLM to provide the request text TX output in response to the instruction statement PP to the other user Ux.
[0032] Here, the situational information indicating the circumstances of user U's movement may be, for example, driving information related to user U's driving. The driving information may also include navigation settings made by user U, driving conditions, driving history, and images captured from inside the vehicle VE. Driving conditions and captured images can be acquired by sensors provided in the user device 10 or the vehicle VE.
[0033] Furthermore, the additional information ADD assumes that the input information IN entered by the user U is simple in content, and includes information that supplements the content of the input information IN, serving as a condition for generating the request statement TX to the LLM. For this reason, the additional information ADD can be rephrased as conditional information that supplements the input information IN entered by the user U.
[0034] [3. Functional Configuration] Using Figure 2, we will describe the configuration examples of the user device 10, the chatbot device 50, and the server device 100A (server device 100). Figure 2 is a diagram showing an example of the device configuration according to Embodiment 1.
[0035] [User device 10] As shown in Figure 2, the user device 10 includes a communication unit 11, a storage unit 12, an input unit 13, an output unit 14, and a control unit 15.
[0036] (Communication Unit 11) The communication unit 11 is implemented by, for example, a NIC (Network Interface Card). The communication unit 11 is connected to the network N by wire or wireless connection and transmits and receives information between, for example, the chatbot device 50 and the server device 100.
[0037] (Storage Unit 12) The storage unit 12 is implemented by, for example, a semiconductor memory element such as RAM (Random Access Memory), ROM (Read Only Memory), or flash memory, or by a storage device such as a hard disk, SSD (Solid State Drive), or optical disc. The storage unit 12 may store, for example, various data and programs related to information processing according to Embodiment 1. The storage unit 12 may also store various data and programs related to information processing according to Embodiment 2.
[0038] (Input Unit 13) The input unit 13 is an input device that receives various inputs from the outside. For example, the input unit 13 is an operating device for user U to perform various operations, such as a keyboard, mouse, or operation keys. If a touch panel is used in the user device 10, the touch panel is also included in the input unit 13. In this case, user U performs various operations by touching the touch panel. The input unit 13 also includes a microphone that receives voice input via speech.
[0039] User U may input inquiry-type information IN via the input unit 13, for example, "What is causing the traffic jam?". Input information IN may be text or voice. Input information IN may be input to the chatbot device 50 via the server device 100.
[0040] (Output Unit 14) The output unit 14 is a device that outputs various things to the outside, such as sound, light, vibration, and images. The output unit 14 outputs various things to the user U in accordance with the control of the control unit 15. The output unit 14 may be a display device that displays various information. The display device may be, for example, a liquid crystal display or an organic electroluminescent display (organic electroluminescent display). The output unit 14 may also be a touch panel type display device. In this case, the input unit 13 and the output unit 14 may be considered as a single unit. The output unit 14 may also be a speaker.
[0041] (Control Unit 15) The control unit 15 is implemented by a CPU (Central Processing Unit) or MPU (Micro Processing Unit), etc., which executes various programs stored in the storage device inside the user device 10 using RAM as the working area. The control unit 15 is also implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).
[0042] As shown in Figure 2, the control unit 15 has a transmitting / receiving unit 151 and an output control unit 152, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 15 is not limited to the configuration shown in Figure 2, and other configurations are also possible as long as they perform the information processing described later. Also, the connection relationships of the various processing units in the control unit 15 are not limited to the connection relationships shown in Figure 2, and other connection relationships are also possible.
[0043] (Transmission / Reception Unit 151) The transmission / reception unit 151 transmits and receives information. For example, the transmission / reception unit 151 receives input information IN input via the input unit 13. For example, the transmission / reception unit 151 receives voice input by speech through a microphone or touch input through a touch panel. Then, the transmission / reception unit 151 transmits the received input information IN. For example, the transmission / reception unit 151 may transmit the received input information IN to the server device 100. Also, the transmission / reception unit 151 may receive response information for the input information IN.
[0044] (Output Control Unit 152) The output control unit 152 performs output control of information. For example, the output control unit 152 may control so that the information provided from the server device 100 is output by the output unit 14.
[0045] [Chatbot Device 50] As shown in FIG. 2, the chatbot device 50 includes a communication unit 51, a storage unit 52, and a control unit 53.
[0046] (Communication Unit 51) The communication unit 51 is realized by, for example, a NIC or the like. And the communication unit 51 is connected to the network N by wire or wirelessly, and performs transmission and reception of information, for example, between the user device 10 and the server device 100.
[0047] (Storage Unit 52) The storage unit 52 is realized by, for example, a semiconductor memory element such as a RAM, ROM, flash memory, or a storage device such as a hard disk, SSD, optical disk. The storage unit 52 may store various data and programs related to the information processing according to Embodiment 1. Note that the storage unit 52 may also store various data and programs related to the information processing according to Embodiment 2. Also, the storage unit 52 may store an LLM.
[0048] (Control Unit 53) The control unit 53 is realized by a CPU, MPU, etc., when various programs stored in the storage device inside the chatbot device 50 are executed using the RAM as a work area. Also, the control unit 53 is realized by an integrated circuit such as an ASIC or FPGA.
[0049] As shown in Figure 2, the control unit 53 includes a receiving unit 531, a response information generation unit 532, and a transmission unit 533, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 53 is not limited to the configuration shown in Figure 2, and other configurations are also possible as long as they perform the information processing described later. Also, the connection relationships of the various processing units in the control unit 53 are not limited to the connection relationships shown in Figure 2, and other connection relationships are also possible.
[0050] (Reception Unit 531) The reception unit 531 receives various types of information in the information processing according to Embodiment 1. The reception unit 531 may also receive various types of information in the information processing according to Embodiment 2.
[0051] (Response Information Generation Unit 532) The response information generation unit 532 generates information using LLM in the information processing according to Embodiment 1. The response information generation unit 532 may also generate information using LLM in the information processing according to Embodiment 2.
[0052] (Transmitting Unit 533) The transmitting unit 533 transmits the information generated by the response information generation unit 532. For example, the transmitting unit 533 transmits the information generated by the response information generation unit 532 to the server device 100.
[0053] As described above, the user device 10 and the chatbot device 50 may have the same functions in Embodiment 2 as well. In other words, the basic operation of the user device 10 and the chatbot device 50 may be the same or similar between Embodiment 1 and Embodiment 2.
[0054] [Server device 100A] As shown in Figure 2, the server device 100A according to Embodiment 1 includes a communication unit 110, a storage unit 120A, and a control unit 130A.
[0055] (Communication Unit 110) The communication unit 110 is implemented by, for example, a NIC. The communication unit 110 is connected to the network N by wire or wireless and transmits and receives information between, for example, the user device 10 and the chatbot device 50.
[0056] (Storage Unit 120A) The storage unit 120A is implemented by, for example, a semiconductor memory element such as RAM, ROM, or flash memory, or a storage device such as a hard disk, SSD, or optical disc. The storage unit 120A may store various data and programs related to the information processing according to Embodiment 1, for example. Also, as shown in Figure 2, the storage unit 120A may have an operation information storage unit 121, an input information storage unit 122A, an instruction statement information storage unit 123A, and a request statement storage unit 124.
[0057] The driving information storage unit 121 may store driving information related to the user U's driving. The input information storage unit 122A may store input information IN (input information IN related to questions or inquiries) entered by the user U. The instruction statement information storage unit 123A may store the instruction statement PP generated by the information processing according to Embodiment 1. The request statement storage unit 124 may store the request statement TX output by the LLM in response to the information processing according to Embodiment 1.
[0058] (Control Unit 130A) The control unit 130A is implemented by a CPU, MPU, etc., which executes various programs (for example, the information processing program according to Embodiment 1) stored in the storage device inside the server device 100A using RAM as the working area. Alternatively, the control unit 130A may be implemented by an integrated circuit such as an ASIC or FPGA.
[0059] As shown in Figure 2, the control unit 130A includes an acquisition unit 131A, a determination unit 132, a setting unit 133, a generation unit 134A, a transmission / reception unit 135A, and a provision control unit 136, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 130A is not limited to the configuration shown in Figure 2, and other configurations are also possible as long as they perform the information processing described later. Also, the connection relationships of the various processing units in the control unit 130A are not limited to the connection relationships shown in Figure 2, and other connection relationships are also possible.
[0060] (Acquisition Unit 131A) The acquisition unit 131A acquires various types of information in the information processing according to Embodiment 1. For example, the acquisition unit 131A acquires input information IN related to a question or inquiry from user U. The input information IN is stored in the input information storage unit 122A, and the acquisition unit 131A may acquire the input information IN from the input information storage unit 122A and output the acquired input information IN to an appropriate processing unit.
[0061] (Determination Unit 132) The determination unit 132 determines the additional items to be added to the input information IN. Specifically, the determination unit 132 determines, based on the words contained in the input information IN, which location type relating to the user U's movement the input information IN is related to as an additional item. For example, the determination unit 132 determines whether the location type is the current location or the destination.
[0062] As mentioned above, the additional information ADD is conditional information that supplements the input information IN entered by the user U. Additional items can be rephrased as necessary items required to supplement the input information IN, and in the process of determining whether the input information IN is a question about the current location or the destination, "current location" and "destination" are the necessary items.
[0063] (Setting Unit 133) The setting unit 133 sets additional information ADD to be added to the input information IN based on status information indicating the status of user U's movement. For example, the setting unit 133 sets additional information ADD to be added to the input information IN based on operation information related to user U's driving. The operation information is transmitted sequentially from the user device 10 and stored in the operation information storage unit 121.
[0064] Furthermore, the setting unit 133 sets additional information ADD with content corresponding to the additional items determined by the determination unit 132. Specifically, the setting unit 133 sets additional information ADD with content corresponding to the location type determined as an additional item by the determination unit 132. For example, if the determination unit 132 determines that the location type is the current location, the setting unit 133 sets additional information ADD including location information indicating the current location and the current time. If the location type is determined to be a destination, the setting unit 133 sets additional information ADD including location information indicating the destination and the estimated time of arrival at the destination.
[0065] (Generation Unit 134A) The generation unit 134A generates an instruction statement PP based on the input information IN and the additional information ADD, which instructs the output of a message requesting a response from another user Ux. Specifically, the generation unit 134A generates an instruction statement PP that includes the input information IN and the condition information ADD, which instructs the output of a request message TX requesting a response from another user Ux. When the generation unit 134A generates an instruction statement PP, it stores the generated instruction statement PP in the instruction statement information storage unit 123A.
[0066] (Transmitting / receiving unit 135A) The transmitting / receiving unit 135A transmits and receives information. For example, the transmitting / receiving unit 135A receives input information IN from the user device 10, sends instruction text PP to the chatbot device 50, receives request text TX from the chatbot device 50, and sends request text TX to the user device 10.
[0067] Furthermore, when the transmitting / receiving unit 135A receives input information IN, it stores the received input information IN in the input information storage unit 122A. Also, when the transmitting / receiving unit 135A receives a request message TX, it stores the received request message TX in the request message storage unit 124.
[0068] (Provision Control Unit 136) The provision control unit 136 controls the provision of the request message TX output by the LLM in response to the instruction message PP to other users Ux. For example, the provision control unit 136 publishes the request message TX output by the LLM to a predetermined service used by other users Ux (for example, a social networking service).
[0069] [4. Information Processing Procedure] From here, we will explain the information processing procedure performed between the user device 10, the chatbot device 50, and the server device 100A, as described in Figure 2.
[0070] Figure 3 is a sequence diagram showing the information processing procedure according to Embodiment 1. Figure 3 shows the overall picture of the information processing according to Embodiment 1, which is realized by the information processing system 1A.
[0071] The transmitting / receiving unit 135A of the server device 100A determines whether or not it has received input information IN from the user device 10 (step S301). If the transmitting / receiving unit 135A has not received input information IN (step S301; No), it waits until it receives input information IN.
[0072] Here, Figure 3 shows an example where user U inputs input information IN into user device 10, asking "What is causing the traffic jam?" or "Are there any parking spaces available?". In this case, the transmitting / receiving unit 151 of user device 10 transmits the input information IN to server device 100A (step S302). As a result, the transmitting / receiving unit 135A determines that it has received the input information IN (step S301; Yes).
[0073] Then, the determination unit 132 of the server device 100A determines the additional items to be added to the input information IN (necessary items required to assist the input information IN) (step S303). Specifically, the determination unit 132 determines, as an additional item, which type of location related to the movement of user U the input information IN is related to, based on the words contained in the input information IN.
[0074] The setting unit 133 of the server device 100A sets additional information ADD with content corresponding to the determined additional item (step S304).
[0075] The generation unit 134A of the server device 100A generates an instruction statement PP based on the input information IN and the additional information ADD, which instructs the output of a message requesting a response from another user Ux (step S305).
[0076] Here, the processing in steps S303 to S305 (processing from the determination of additional items to the generation of instruction statements) will be explained in more detail using Figure 4. Figure 4 is a flowchart showing the detailed procedure of information processing according to Embodiment 1.
[0077] When the server device 100A receives input information IN from the transmitting / receiving unit 135A, the acquisition unit 131A acquires the received input information IN (step S401).
[0078] The determination unit 132 decomposes the input information IN into word-level units (step S402). Then, based on the words obtained from the decomposition, the determination unit 132 determines whether the input information IN is a question about the current location or the destination (step S403).
[0079] For example, if the determination unit 132 obtains a word that suggests the present time (e.g., "current location" or "now"), it can determine that the input information IN is a question about the current location. On the other hand, if the determination unit 132 obtains a word that suggests the future (e.g., "destination" or "where I'm going"), it can determine that the input information IN is a question about the destination.
[0080] The determination unit 132 may also cause the LLM to determine whether the input information IN is a question about the current location or the destination. For example, the determination unit 132 (or the generation unit 134A) may generate an instruction statement instructing the LLM to determine whether the input information IN is a question about the current location or the destination, and input the generated instruction statement to the LLM. The instruction statement may include information indicating the current location, information indicating the destination, and the input information IN itself. In the example in Figure 3, the determination unit 132 can determine that if the input information IN is "What is causing the traffic jam?", this input information IN is a question about the current location. On the other hand, if the input information IN is "Is there a parking space available?", the determination unit 132 can determine that this input information IN is a question about the destination.
[0081] Returning to the explanation, if the acquisition unit 131A determines that the input information IN is a question about the current location, it acquires location information PT1 indicating the current location of user U and date and time information TM1 indicating the current date and time (step S404a).
[0082] In such cases, the setting unit 133 sets additional information ADD1 which includes location information PT1 indicating the current location and date and time information TM1 indicating the current date and time (step S405a).
[0083] The generation unit 134A generates an instruction statement PP_1 based on the input information IN and the additional information ADD1 (step S406a). Figure 4 shows an example in which the generation unit 134A generates an instruction statement PP_1 that says, "Based on this [text], please create a text that asks a question to another person, assuming the [conditions]."
[0084] Using the example in Figure 3, the input information IN, "What is causing the traffic jam?", is inserted into the [Text] field. Additionally, the additional information ADD1 is inserted into the [Conditions] field. For example, if the current location is "National Route X, Kawagoe City A-cho" and the current date and time is "July 22, 2024, 13:00", then the location information PT1 indicating "National Route X, Kawagoe City A-cho" and the date and time information TM1 indicating "July 22, 2024, 13:00" are inserted into the [Conditions] field.
[0085] On the other hand, if the acquisition unit 131A determines that the input information IN is a question about a destination, it acquires location information PT2 indicating the user U's destination and date and time information TM2 indicating the estimated date and time of arrival at the destination (step S404b).
[0086] In such cases, the setting unit 133 sets additional information ADD2 which includes location information PT2 indicating the destination and date and time information TM2 indicating the estimated date and time of arrival (step S405b).
[0087] The generation unit 134A generates an instruction statement PP_2 based on the input information IN and the additional information ADD2 (step S406b). Figure 4 shows an example in which the generation unit 134A generates an instruction statement PP_2 that says, "Based on this [text], please create a text that asks a question to another person, assuming the [conditions]."
[0088] Using the example in Figure 3, the input information IN, "Is there a parking space available?", is inserted into the [Text] field. Additionally, the additional information ADD2 is inserted into the [Conditions] field. For example, if the destination is "Roadside Station XX" and the estimated arrival date and time is "July 22, 2024, 5:00 PM", then the location information PT2 indicating "Roadside Station XX" and the date and time information TM2 indicating "July 22, 2024, 5:00 PM" are inserted into the [Conditions] field.
[0089] Figure 4 illustrates a specific example of the process in steps S303 to S305. Return to the explanation of Figure 3.
[0090] The transmitting / receiving unit 135A transmits the generated instruction text PP to the chatbot device 50 (step S306).
[0091] When the response information generation unit 532 of the chatbot device 50 receives the instruction text PP from the reception unit 531, it inputs the instruction text PP to the LLM (step S307).
[0092] Furthermore, the response information generation unit 532 acquires the request statement TX output by the LLM in accordance with the instruction statement PP (step S308). Here, when the instruction statement PP_1 shown in Figure 4 is input to the LLM, it may output a request statement TX1 as shown in Figure 3. On the other hand, when the instruction statement PP_2 shown in Figure 4 is input to the LLM, it may output a request statement TX2 as shown in Figure 3.
[0093] The transmission unit 533 of the chatbot device 50 sends the request text TX to the server device 100A (step S309).
[0094] When the server device 100A receives a request message TX from the transmitting / receiving unit 135A, the provision control unit 136 controls the provision of the request message TX to other users Ux (step S310). For example, the provision control unit 136 publishes the request message TX to a predetermined social network service to request a response from other users Ux to the input information IN, whose content has been supplemented by additional information ADD.
[0095] For example, if there is an external device 60 (Figure 5) as an information processing device for managing a predetermined social network service, the provision control unit 136 may control the publication of request text TX to the external device 60.
[0096] [5. Modifications] From here, modifications of Embodiment 1 will be described. The server device 100A according to Embodiment 1 may be realized in various forms different from Embodiment 1 described above.
[0097] (5-1. Setting additional information considering the conditions inside the vehicle) In the above embodiment 1, the setting unit 133 was shown as an example in which additional information ADD is set according to the location type (current location or destination) determined by the determination unit 132. However, the setting unit 133 may also set the additional information ADD based on the captured image of the interior of the vehicle VE.
[0098] For example, the acquisition unit 131A may acquire information about the interior of the vehicle VE in which user U is riding, specifically, captured images showing the interior of the vehicle VE, as situational information indicating the situation of user U's movement. The setting unit 133 may then set additional information ADD, which is an attribute estimated by the analysis of the captured image and corresponds to the attribute of the object included in the captured image.
[0099] The object referred to here may be a passenger in the vehicle VE (for example, a passenger other than user U), and the setting unit 133 may set additional information ADD according to the passenger's attributes.
[0100] For example, if the setting unit 133 can detect a passenger from the captured image and estimate that the passenger is a "child," it may set additional information ADD, such as "a place a child would like." According to this process, for example, if the generation unit 134A receives input information IN asking, "Is there anywhere we can stop by near our destination?", it can generate an instruction statement PP that asks other user Ux if there are any places near the destination that a child might like to stop by, thus creating the possibility that user U can obtain special places to stop by that would not be output by the speech agent system.
[0101] In the above embodiment 1, the determination unit 132 was shown to determine whether the input information IN is a question about the current location or the destination based on the words contained in the input information IN. However, it may also determine whether the setting of additional information ADD according to the passenger's attributes is appropriate based on the words contained in the input information IN. This makes it possible to suppress the generation unit 134A from generating an instruction statement PP that instructs it to output a sentence requesting another user Ux to give a response that user U does not want, for example, when input information IN asking "What is the cause of the traffic jam?" is received.
[0102] (5-2. Instruction text generation considering the user's living area) The determination unit 132 may determine, based on the words included in the input information IN, whether user U desires a first response targeting the user U's living area or a second response targeting an area outside of user U's living area, and the generation unit 134A may generate an instruction text PP with content corresponding to the determination result.
[0103] For example, if user U desires a first response, the generation unit 134A may set user Ux, who is located within the user's living area, as the other user Ux, and generate an instruction statement PP instructing the unit to output a message requesting a first response that includes insider information within the user's living area. Through this process, user U will be able to obtain information that they could not have known before, even though it is within their own living area.
[0104] Furthermore, if user U desires a second response, the generation unit 134A may generate an instruction statement PP that instructs the system to set user Ux, who is located outside of the user's area of residence, as the other user Ux, and to output a message requesting a detailed second response. Through this process, user U can obtain local support from other user Ux, for example, if they are in an unfamiliar place such as a travel destination.
[0105] <Embodiment 2> [1. Introduction] From here, Embodiment 2 will be described. In Embodiment 1, the provision control unit 136 of the server device 100A controls the provision of the request message TX output by LLM to other users Ux, and as a process for this, it was explained that it controls the publication of the request message TX to an external device 60 that manages a predetermined social network service.
[0106] In such a case, since other user Ux posts response information to the request TX to the external device 60, if the response information is provided directly to user U as in the conventional technology described above, user U may have to perform the cumbersome task of checking each piece of response information.
[0107] For example, if user U is driving, it is difficult to check each response information received from other users Ux. Therefore, it is desirable to be able to easily understand the content of the response information received from other users Ux, even while driving.
[0108] Therefore, the inventors of the present invention conceived the idea that usability would be improved if the user were provided with the summarized content output by the LLM, which is provided as input to an instruction sentence that instructs the LLM to summarize the content of other users' response information to the response request sentence output by the LLM in response to the input information.
[0109] In other words, according to the proposed technology of Embodiment 2, response information from another user Ux to a request statement output by LLM in response to input information regarding a question or inquiry from user U is acquired, and an instruction statement is generated instructing that the contents of that response information be summarized. Furthermore, according to the proposed technology of Embodiment 2, the summary content output by LLM is provided to user U, taking the response information and the instruction statement as input.
[0110] [2. System Configuration] The configuration of the system according to the embodiment will be explained using Figure 5. Figure 5 is a diagram showing an example of the configuration of the information processing system 1B according to Embodiment 2. As shown in Figure 5, the information processing system 1B includes a user device 10, a chatbot device 50, an external device 60, and a server device 100B. The user device 10, the chatbot device 50, the external device 60, and the server device 100B may be connected to each other via a predetermined communication network (network N) by wired or wireless means.
[0111] As described above, the basic operation of the user device 10 and the chatbot device 50 is the same or similar between Embodiment 1 and Embodiment 2; therefore, the description of the user device 10 and the chatbot device 50 will be omitted in Embodiment 2.
[0112] Furthermore, the server device 100B according to Embodiment 2 is a central information processing device that realizes the information processing according to Embodiment 2. Therefore, the server device 100B is said to have new functions for performing the information processing according to Embodiment 2, in addition to the storage unit and functions of the server device 100A according to Embodiment 1.
[0113] Specifically, the server device 100B acquires the response information AN from another user Ux to the request statement TX output by LLM in response to the input information IN regarding the user U's question or inquiry, and generates an instruction statement PP1 that instructs the server device 100B to summarize the contents of the response information AN. Then, the server device 100B takes the response information AN and the instruction statement PP1 as input and provides the summarized contents SM1 output by LLM to user U.
[0114] The external device 60 has the function of managing a predetermined social network service. As shown in Figure 5, other users Ux input response information AN to the external device 60, so the server device 100B has the response information AN of each other user Ux aggregated in the LLM.
[0115] [3. Functional Configuration] Using Figure 6, we will describe the configuration examples of the user device 10, the chatbot device 50, and the server device 100B (server device 100). Figure 6 is a diagram showing the device configuration example according to Embodiment 2. In the following, only the server device 100B will be described.
[0116] [Server device 100B] As shown in Figure 6, the server device 100B according to Embodiment 2 includes a communication unit 110, a storage unit 120B, and a control unit 130B.
[0117] In the example in Figure 6, the memory unit and processing unit having the symbol "B" means that, compared to the memory unit and processing unit (Figure 2) in Embodiment 1 which have the same name and the symbol "A", new functions have been added to realize the information processing according to Embodiment 2. Also, in the example in Figure 6, the memory unit and processing unit not having the symbol "B" means that new functions have been added to realize the information processing according to Embodiment 2.
[0118] (Storage Unit 120B) The storage unit 120B is implemented by, for example, a semiconductor memory element such as RAM, ROM, or flash memory, or a storage device such as a hard disk, SSD, or optical disc. The storage unit 120B may store various data and programs related to the information processing according to Embodiment 2, for example. Also, as shown in Figure 6, the storage unit 120B may have an input information storage unit 122B, an instruction statement information storage unit 123B, a response information storage unit 125, and a summary content storage unit 126.
[0119] The input information storage unit 122B may further store new input information INa that has been additionally input to the information provided in accordance with the input information IN (input information IN whose content has been supplemented by additional information ADD) according to Embodiment 1. The instruction statement information storage unit 123B may store the instruction statement PP generated in the information processing according to Embodiment 2. The response information storage unit 125 may store the response information AN of another user Ux to the request statement TX output by LLM in accordance with the input information IN. The summary content storage unit 126 may store the summary content SM1 output by LLM in accordance with the information processing according to Embodiment 2.
[0120] (Control Unit 130B) The control unit 130B is implemented by a CPU, MPU, etc., which executes various programs (for example, the information processing program according to Embodiment 2) stored in the storage device inside the server device 100B using RAM as the working area. The control unit 130B is also implemented by an integrated circuit such as an ASIC or FPGA.
[0121] As shown in Figure 6, the control unit 130B includes an acquisition unit 131B, a generation unit 134B, a transmission / reception unit 135B, an extraction unit 137, and a provision unit 138, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 130B is not limited to the configuration shown in Figure 6, and other configurations are also possible as long as they perform the information processing described later. Also, the connection relationships of the various processing units in the control unit 130B are not limited to the connection relationships shown in Figure 6, and other connection relationships are also possible.
[0122] (Acquisition Unit 131B) The acquisition unit 131B acquires various types of information in the information processing according to Embodiment 2. For example, the acquisition unit 131B acquires the response information AN of another user Ux to the request statement TX output by LLM in response to the input information IN of user U's question or inquiry. The response information AN is stored in the response information storage unit 125, and the acquisition unit 131B may acquire the response information AN from the response information storage unit 125 and output the acquired response information AN to an appropriate processing unit.
[0123] Furthermore, the acquisition unit 131B acquires newly input information INA that has been additionally input. The input information INA is stored in the input information storage unit 122B, and the acquisition unit 131B may acquire the input information INA from the input information storage unit 122B and output the acquired input information INA to an appropriate processing unit.
[0124] Furthermore, the acquisition unit 131B may also acquire the location information PTx of the other user Ux at the time the other user Ux responds.
[0125] (Generation Unit 134B) The generation unit 134B generates an instruction statement PP1 that instructs the system to summarize the contents of the response information AN. For example, the generation unit 134B generates an instruction statement PP1 that includes a condition CD for the method of summarizing the contents of the response information AN. The condition CD may be a condition that summarizes a predetermined number of the top items of the response information AN according to a predetermined criterion, and the predetermined criterion may be, for example, in order of the speed of the response or in order of the appropriateness of the content. Alternatively, the condition CD may be a condition that the number of characters in the summarized content SM1 be kept within a predetermined range. When the generation unit 134B generates the instruction statement PP1, it stores the generated instruction statement PP1 in the instruction statement information storage unit 123B.
[0126] Furthermore, if user U inputs additional input information INa (additional information) to be queried for the summary content SM1, the generation unit 134B further generates an instruction statement PP2 that instructs the system to investigate whether there is a response information ANA among the response information AN acquired so far that is suitable as a response to the input information INa. When the generation unit 134B generates the instruction statement PP2, it stores the generated instruction statement PP2 in the instruction statement information storage unit 123B.
[0127] (Transmitting / receiving unit 135B) The transmitting / receiving unit 135B transmits and receives information. For example, the transmitting / receiving unit 135B receives response information AN from the external device 60, sends instruction text PP1 and instruction text PP2 to the chatbot device 50, receives summary content SM1 from the chatbot device 50, and sends summary content SM1 to the user device 10.
[0128] Furthermore, when the transmitting / receiving unit 135B receives response information AN, it stores the received response information AN in the response information storage unit 125. Also, when the transmitting / receiving unit 135B receives summary content SM1, it stores the received summary content SM1 in the summary content storage unit 126.
[0129] (Extraction Unit 137) The extraction unit 137 has the role of a filter function that excludes response information AN that has been acquired so far and has been determined not to be appropriate as a response to the input information IN. That is, the extraction unit 137 extracts appropriate response information from the response information AN that has been determined to be suitable as a response to the input information IN. In this case, the generation unit 134B generates an instruction statement PP1 that instructs the summarization of the contents of the appropriate response information.
[0130] For example, the extraction unit 137 may extract appropriate response information from the response information AN based on a comparison between the determination result, which determines which position type the input information IN relates to regarding the movement of user U, and the position information PTx of other user Ux at the time the other user Ux responded.
[0131] (Providing unit 138) The providing unit 138 inputs the response information AN into the LLM in multiple stages depending on the passage of time or the accumulation of response information AN, and provides the summarized content SM1 to the user U according to this stage.
[0132] For example, the acquisition unit 131B may acquire response information AN1 that was responded to by other user Ux during the first period T1, after the request message TX output by LLM has been controlled to be provided to other user Ux. The provision unit 138 may then input the response information AN1 that was responded to by other user Ux during the first period T1 to LLM.
[0133] Furthermore, the acquisition unit 131B may acquire response information AN2 received by the other user Ux during the second period T2 if a second period T2 longer than the first period T1 has elapsed since the request message TX output by the LLM was controlled to be provided to the other user Ux. The provision unit 138 may then input the response information AN2 received by the other user Ux during the second period T2 to the LLM.
[0134] Here, other users Ux respond at various times. Therefore, response information AN is received sequentially by the transmitting / receiving unit 135B according to the timing of the responses. Accordingly, the acquisition unit 131B continuously acquires response information AN in response to the sequential reception of response information AN by the transmitting / receiving unit 135B. The providing unit 138 may then determine whether the summary content SM11 provided to user U in the previous stage (for example, the first period T1) is different from the summary content SM12 output by LLM in the current stage (for example, the second period T2) based on the continuously acquired response information AN. If the providing unit 138 determines that they are different, it may provide user U with the newly output summary content SM12 from LLM.
[0135] [4. Information Processing Procedure] From here, we will explain the information processing procedure performed between the user device 10, the chatbot device 50, the external device 60, and the server device 100B, as described in Figure 6. The information processing procedure will be explained using Figures 7 and 8.
[0136] Figure 7 is a sequence diagram (1) showing the information processing procedure according to Embodiment 2. Figure 7 shows the overall picture of the information processing according to Embodiment 2 realized by the information processing system 1B.
[0137] Other users Ux input response information AN to the external device 60 at various timings. Therefore, as shown in Figure 7, the response information AN is sequentially received by the transmitting / receiving unit 135B of the server device 100B according to the timing of the response. The transmitting / receiving unit 135B then receives the response information AN and determines whether the first period T1 has elapsed since the request message TX output by the LLM was controlled to be provided to other users Ux (step S701). The time when the request message TX output by the LLM was controlled to be provided to other users Ux may be the time when the processing in step S310 (Figure 3) is executed.
[0138] The transmitting / receiving unit 135B waits until the first period T1 has elapsed (step S701; No).
[0139] On the other hand, if the first period T1 has elapsed (step S701; Yes), the acquisition unit 131B of the server device 100B acquires the response information AN1 accumulated during the first period T1 (step S702).
[0140] The generation unit 134B of the server device 100B generates an instruction statement PP11 that instructs the server to summarize the contents of the response information AN1 (step S703). At this time, the generation unit 134B may include the conditions CD of the method for summarizing the contents of the response information AN1 in the instruction statement PP11.
[0141] The transmitting / receiving unit 135B of the server device 100B transmits the generated instruction text PP11 to the chatbot device 50 (step S704).
[0142] When the response information generation unit 532 of the chatbot device 50 receives the instruction statement PP11 from the reception unit 531, it inputs the instruction statement PP11 to the LLM (step S705).
[0143] Furthermore, the response information generation unit 532 acquires the summary content SM11 output by the LLM in accordance with the instruction statement PP11 (step S706).
[0144] The transmission unit 533 of the chatbot device 50 transmits the summary content SM11 to the server device 100B (step S707).
[0145] When the server device 100B receives the summary content SM11 from the transmitting / receiving unit 135B, the providing unit 138 provides the summary content SM11 to user U (step S708). Specifically, the providing unit 138 transmits the summary content SM11 to user U's user device 10, causing the information of the summary content SM11 to be output from the output unit 14.
[0146] Next, Figure 8 will be explained. Figure 8 is a sequence diagram (2) showing the information processing procedure according to Embodiment 2. Figure 8 shows the overall picture of the information processing according to Embodiment 2 realized by the information processing system 1B.
[0147] The information processing shown in Figure 8 is performed following the information processing shown in Figure 7. Specifically, when the transmission / reception unit 135B receives the summary contents SM11 from user U, it controls the request message TX output by LLM to be provided to another user Ux, and then determines whether a second period T2, which is longer than the first period T1, has elapsed (step S801).
[0148] The transmitting / receiving unit 135B waits until the second period T2 has elapsed (step S801; No).
[0149] On the other hand, if the second period T2 has elapsed (step S801; Yes), the acquisition unit 131B acquires the response information AN2 accumulated during the second period T2 (step S802).
[0150] The generation unit 134B generates an instruction statement PP12 that instructs the system to summarize the contents of the response information AN2 (step S803). At this time, the generation unit 134B may include the conditions CD for the method of summarizing the contents of the response information AN2 in the instruction statement PP12.
[0151] The transmitting / receiving unit 135B transmits the generated instruction text PP12 to the chatbot device 50 (step S804).
[0152] When the response information generation unit 532 receives the instruction statement PP12 from the reception unit 531, it inputs the instruction statement PP12 to the LLM (step S805).
[0153] Furthermore, the response information generation unit 532 acquires the summary content SM12 output by the LLM in accordance with the instruction statement PP12 (step S806).
[0154] The transmitting unit 533 transmits the summary contents SM12 to the server device 100B (step 807).
[0155] When the server device 100B receives the summary content SM12 from the transmitting / receiving unit 135B, the providing unit 138 provides the summary content SM12 to user U (step S808). Specifically, the providing unit 138 transmits the summary content SM12 to user U's user device 10, causing the information of the summary content SM12 to be output from the output unit 14.
[0156] Here, using Figure 9, we will explain specific examples of the instruction statement PP1 (instruction statement PP11 in the example of Figure 7, instruction statement PP12 in the example of Figure 8), the conditions for summarization CD, and the summary content SM1 (summary content SM11 in the example of Figure 7, and summary content SM12 in the example of Figure 8). Figure 9 is a diagram showing specific examples of the instruction statement PP1, the conditions for summarization CD, and the summary content SM1 according to Embodiment 2.
[0157] As shown in Figure 9, the generation unit 134B may generate an instruction statement PP1 that instructs the system to consolidate the response information AN accumulated during the period T after the request statement TX output by the LLM has been controlled to be provided to another user Ux. For example, as shown in Figure 9, the generation unit 134B may generate the instruction statement PP1 by associating the response information AN accumulated during the period T with a sentence that says, "Please consolidate the following sentences." Figure 9 also shows an example of a response, which is an example of response information AN.
[0158] Furthermore, the conditions CD included in the instruction statement PP1 may include "summarize by narrowing down to the top predetermined number of response information AN" or "summarize within X characters." When narrowing down to the top predetermined number of response information AN, the ranking may be done in order of the earliest response (order of submission) or in order of the appropriateness of the content of the response information AN. In determining the order of appropriateness, a method may be used in which the appropriate response information extracted by the extraction unit 137 is given a higher rank, or a method may be used in which the rank is given according to the degree of accuracy as a response to the input information IN. By including such conditions CD in the instruction statement PP1, it becomes possible to provide the user U with a summary content SM1 that is highly accurate and easy to understand.
[0159] To give an example, in step S703 of Figure 7, the generation unit 134B may generate instruction statement PP11 by associating the response information AN1 accumulated during the first period T1 with the sentence "Please summarize the following sentences." Also, in step S803 of Figure 8, the generation unit 134B may generate instruction statement PP12 by associating the response information AN2 accumulated during the second period T2 with the sentence "Please summarize the following sentences."
[0160] Figure 9 also shows specific examples of summary content SM1 output by LLM in response to instruction PP1, such as summary content SM11 output by LLM in response to instruction PP11, and summary content SM12 output by LLM in response to instruction PP12. For example, LLM may output summary content SM1 following the introductory phrase, "To summarize your answers, it is as follows." The introductory phrase may also be generated on the server device 100B side.
[0161] Furthermore, Figures 7 and 8 show an example in which the server device 100B provides the summary content SM1 in two stages: the first stage (first period T1) and the second stage (second period T2). Through this process, the request for quick responses from other users Ux can be met, and a more accurate summary content can be provided again by summarizing more response information AN.
[0162] The method of staging by the server device 100B and the number of stages are not limited to the examples in Figures 7 and 8. For example, the server device 100B may staging based on the number of stored response information AN.
[0163] [5. Method for Providing Summary Contents] As described above, the provisioning unit 138 determines whether the summary content SM11 provided to user U in the previous stage (for example, the first period T1) and the summary content SM12 output in the current stage (for example, the second period T2) are different. If it determines that they are different, it may provide user U with the summary content SM12 newly output by LLM. An example of such provisioning method will be explained using Figure 10. Figure 10 is a flowchart showing the method for providing the summary content SM1.
[0164] Note that the procedure shown in Figure 10 may be performed, for example, from step S807 onwards in Figure 8. For example, the transmitting / receiving unit 135B determines whether or not it has received the summary content SM12 (step S1001). If the transmitting / receiving unit 135B has not received the summary content SM12 (step S1001; No), it waits until it receives the summary content SM12.
[0165] On the other hand, if the providing unit 138 receives the summary content SM12 (step S1001; Yes), it calculates the similarity between the summary content SM11 and the summary content SM12 (step S1002). The providing unit 138 may calculate the similarity using any natural language processing method. The providing unit 138 then determines whether the similarity is below a certain threshold value (step S1003).
[0166] If the similarity is below a certain threshold (step S1003; Yes), the providing unit 138 determines that the summary content SM11 and the summary content SM12 are different, specifically that the summary content SM12 contains new information compared to the summary content SM11 (step S1004a). In this case, the providing unit 138 provides the summary content SM12 to user U (step S1005a).
[0167] If the similarity is higher than the standard value (step S1003; No), the providing unit 138 determines that the summary content SM11 and the summary content SM12 are similar, specifically that there is no change in the summary content SM12 compared to the summary content SM11 (step S1004b). In this case, the providing unit 138 does not provide the summary content SM12 to user U (step S1005b).
[0168] According to the provision method shown in Figure 10, when providing the summary content SM1 in multiple stages, it is possible to control the provision so that the same summary content SM1 is not provided.
[0169] Furthermore, the provisioning unit 138 does not necessarily need to use similarity to determine whether the summary content SM11 and the summary content SM12 are different. For example, the provisioning unit 138 may determine whether the summary content SM12 contains new information by comparing it with the summary content SM11 based on the word distance calculated between the summary content SM11 and the summary content SM12.
[0170] [6. Filtering Method] As described above, the extraction unit 137 has the role of a filter function that excludes response information AN that has been acquired so far and has been determined not to be an appropriate response to the input information IN. Figure 11 illustrates the filtering method for excluding response information AN that is not an appropriate response. Figure 11 is a flowchart of the filtering method. The filtering method shown in Figure 11 may be executed between steps S702 and S703 in the example of Figure 7, and between steps S802 and S803 in the example of Figure 8.
[0171] In the filtering method, first, it may be determined whether the input information IN is a question about the current location or the destination by processing steps S401 to S403 shown in Figure 4.
[0172] If the acquisition unit 131B determines that the input information IN is a question about the current location, it acquires location information PT1 indicating the current location of user U and location information PTx of other user Ux at the time that other user Ux responded (step S1104a).
[0173] The extraction unit 137 determines, based on the location information PT1 and location information PTx, whether or not there are any items in location information PTx that are included within the area corresponding to location information PT1 (step S1105a).
[0174] If the extraction unit 137 finds that there is a location information PTx that is included within the area corresponding to location information PT1 (step S1105a; Yes), it extracts only the response information AN of other user Ux that responded at the location indicated by that location information PTx that is included within the area (step S1106a). Other response information AN is not extracted and is excluded.
[0175] As a result, the generation unit 134B generates an instruction statement PP1 that instructs the system to summarize the contents of the response information AN extracted in step S1106a (step S1107a).
[0176] On the other hand, if there is no location information PTx that is included within the area corresponding to location information PT1 (step S1105a; No), the generation unit 134B may generate an instruction statement PP1 that instructs the unit to use all the response information AN accumulated up to that point to summarize the contents (step S1108a).
[0177] Furthermore, if the acquisition unit 131B determines that the input information IN is a question about a destination, it acquires location information PT2 indicating the destination of user U and location information PTx of other user Ux at the time that user Ux responded (step S1104b).
[0178] The extraction unit 137 determines, based on the location information PT2 and location information PTx, whether or not there are any items in location information PTx that are included within the area corresponding to location information PT2 (step S1105b).
[0179] If the extraction unit 137 finds that there is a location information PTx that is included within the area corresponding to the location information PT2 (step S1105b; Yes), it extracts only the response information AN of other user Ux that responded at the location indicated by that location information PTx that is included within the area (step S1106b). Other response information AN is not extracted and is excluded.
[0180] As a result, the generation unit 134B generates an instruction statement PP1 that instructs the system to summarize the contents of the response information AN extracted in step S1106b (step S1107b).
[0181] On the other hand, if there is no location information PTx that is included within the area corresponding to location information PT2 (step S1105b; No), the generation unit 134B may generate an instruction statement PP1 that instructs the unit to use all the response information AN accumulated so far to summarize the contents (step S1108b).
[0182] According to the filtering method shown in Figure 11, response information AN that is noisy and not appropriate as a response to the input information IN can be excluded, and only response information AN that is appropriate as a response to the input information IN can be summarized, thereby improving the accuracy of the summarized content SM1.
[0183] [7. Investigation Instruction Procedure] As described above, if user U inputs additional input information INa (additional information) to be inquired about regarding the summary content SM1, the generation unit 134B further generates an instruction statement PP2 that instructs the system to investigate whether there is any response information ANA among the response information AN acquired so far that is suitable as an answer to the input information INa.
[0184] Therefore, Figure 12 illustrates the investigation instruction procedure related to the generation of instruction statement PP2. Figure 12 is a sequence diagram showing the investigation instruction procedure according to Embodiment 2. Figure 12 shows an overall view of the investigation instruction procedure according to Embodiment 2 implemented by the information processing system 1B. Note that the investigation instruction procedure shown in Figure 12 may be executed from the time the summary content SM1 is provided.
[0185] For example, the transmitting / receiving unit 135B of the server device 100B determines whether or not it has received additional input information INa from the user device 10 (step S1201). If the transmitting / receiving unit 135B has not received input information INa (step S1201; No), it waits until it receives input information INa.
[0186] Here, Figure 12 shows an example in which user U inputs input information INa into user device 10, asking, "If the cause of the traffic jam is unknown, when will it be cleared?" In this case, the transmitting / receiving unit 151 of user device 10 transmits the input information INa to server device 100B (step S1202). As a result, the transmitting / receiving unit 135B determines that it has received the input information INa (step S1201; Yes).
[0187] Then, the generation unit 134B of the server device 100B generates an instruction statement PP2 that instructs the server to investigate whether there is any response information AN among the response information AN that is suitable as a response to the input information INa (step S1203).
[0188] The transmitting / receiving unit 135B transmits the generated instruction text PP2 to the chatbot device 50 (step S1204). For example, the transmitting / receiving unit 135B may transmit the instruction text PP2 together with the response information AN that has been stored up to that point.
[0189] When the response information generation unit 532 of the chatbot device 50 receives the instruction PP2 from the reception unit 531, it inputs the instruction PP2 into the LLM (step S1205). The response information generation unit 532 may also input previously stored response information AN into the LLM.
[0190] The LLM investigates whether there is a response information ANA that is suitable as a response to the input information INa, in accordance with instruction PP2, and if a response information ANA exists, it outputs that response information ANA (step S1206).
[0191] When response information ANA is output, the transmission unit 533 of the chatbot device 50 transmits the response information ANA to the server device 100B (step S1207).
[0192] When the server device 100B receives the response information ANA from the transmitting / receiving unit 135B, the providing unit 138 provides the response information ANA to user U (step S1208). Specifically, the providing unit 138 transmits the response information ANA to user U's user device 10, causing the response information ANA to be output from the output unit 14.
[0193] Here, if user U inputs additional input information INa, it is desirable to efficiently provide user U with response information ANA that is appropriate for the input information INa. Also, there is a possibility that response information ANA may have been excluded by the filtering methods described so far. According to the process shown in Figure 12, the system searches for response information ANA that was excluded and not included in the summary content SM1, and if response information ANA is found, it is provided to user U, thus enabling the efficient provision of response information ANA to user U.
[0194] Furthermore, according to step S1206, the LLM investigates whether there is a response information ANA that is suitable as a response to the input information INa, and may not be able to output the response information ANA if no such response information ANA exists. In addition, as a result, in step S1206, the server device 100B may not be able to accept the response information ANA.
[0195] Therefore, if the LLM is unable to output response information ANA, the server device 100 may perform an alternative method to not provide the response information ANA to the user U. Figure 13 is a flowchart of the alternative method. The alternative method shown in Figure 13 may be performed according to the result of step S1206 in the example of Figure 12.
[0196] First, the transmitting / receiving unit 135B determines whether or not response information ANA has been output by LLM (step S1301). For example, the transmitting / receiving unit 135B can determine whether or not response information ANA has been output by LLM based on the output result of response information ANA transmitted from the chatbot device 50.
[0197] Then, if response information ANA is output by LLM (step S1301; Yes), the providing unit 138 provides the response information ANA to user U (step S1302).
[0198] On the other hand, if the LLM does not output response information ANA (step S1301; No), the generation unit 134B generates an instruction statement PP3 that instructs the system to extract the related user Uy who responded with response information ANb related to the input information INa from among other users Ux (step S1303). Figure 13 shows an example in which the generation unit 134B generates an instruction statement PP3 that says, "Please find the person who entered a response related to this [text]." Using the example in Figure 12, the input information INa in the [text] field may be "If the cause of the traffic jam is unknown, when will it be cleared?".
[0199] Furthermore, the generation unit 134B generates an instruction statement PP4 based on the input information INa, which instructs the output of a sentence requesting a response from the related user Uy (step S1304). Specifically, the generation unit 134B generates an instruction statement PP4 which instructs the output of a sentence requesting a response from the input information INa. Figure 13 shows an example in which the generation unit 134B generates an instruction statement PP4 that says, "Based on this [sentence], please create a sentence to ask [the person you extracted] a question." In the [sentence] field, the input information INa may be inserted as, "If the cause of the traffic jam is unknown, when will it be cleared?" Also, [the person you extracted] refers to the related user Uy extracted by LLM itself in response to instruction statement PP3.
[0200] Furthermore, the transmitting / receiving unit 135B transmits the generated instruction sentences PP3 and PP4 to the chatbot device 50, and the response information generation unit 532 of the chatbot device 50 inputs instruction sentences PP3 and PP4 to the LLM. In response to the input of instruction sentences PP3 and PP4, the LLM outputs a request sentence TX4 that requests a response from the relevant user Uy.
[0201] Therefore, the provision control unit 136 (Figure 2) controls the disclosure of the request message TX4 output by the LLM so that it is made public to the related user Uy (step S1305).
[0202] The providing unit 138 provides user U with the response information ANy input by the related user Uy (step S1306). Alternatively, the providing unit 138 may provide user U with the summarized content obtained by having the LLM summarize the response information ANy input by the related user Uy. The method for obtaining such summarized content has been explained above and will be omitted here.
[0203] According to the process shown in Figure 13, even if the LLM is unable to output response information ANA, it is possible to support the LLM in outputting response information ANy, which corresponds to (and is presumed to be just as suitable as) response information ANA.
[0204] <Other Embodiments> In Embodiment 1, an example was shown in which the server device 100A and the chatbot device 50 are separate devices, but a configuration in which the server device 100A and the chatbot device 50 are integrated may also be adopted. In Embodiment 2, an example was shown in which the server device 100B and the chatbot device 50 are separate devices, but a configuration in which the server device 100B and the chatbot device 50 are integrated may also be adopted.
[0205] <Hardware Configuration> The server device 100 according to the embodiment (server device 100A according to Embodiment 1, server device 100B according to Embodiment 2) may be implemented by a computer 1000 having a configuration such as that shown in Figure 14. Figure 14 is a hardware configuration diagram showing an example of a computer that implements the functions of the server device 100 according to the embodiment. The computer 1000 has a CPU 1100, RAM 1200, ROM 1300, HDD 1400, communication interface (I / F) 1500, input / output interface (I / F) 1600, and media interface (I / F) 1700.
[0206] The CPU 1100 operates based on programs stored in the ROM 1300 or HDD 1400, and controls various parts. The ROM 1300 stores boot programs executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.
[0207] The HDD 1400 stores programs executed by the CPU 1100, and data used by such programs. The communication interface 1500 receives data from other devices via a predetermined communication network and sends it to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the predetermined communication network.
[0208] The CPU 1100 controls output devices such as displays and input devices such as keyboards via the input / output interface 1600. The CPU 1100 acquires data from input devices via the input / output interface 1600. The CPU 1100 also outputs the generated data to output devices via the input / output interface 1600.
[0209] The media interface 1700 reads a program or data stored in the recording medium 1800 and provides it to the CPU 1100 via the RAM 1200. The CPU 1100 loads the program from the recording medium 1800 onto the RAM 1200 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0210] For example, when computer 1000 functions as a server device 100A, the CPU 1100 of computer 1000 implements the functions of control unit 130A by executing programs loaded onto RAM 1200. The CPU 1100 of computer 1000 reads and executes these programs from the recording medium 1800, but as another example, these programs may be obtained from other devices via a predetermined communication network.
[0211] For example, when computer 1000 functions as server device 100B, the CPU 1100 of computer 1000 implements the functions of control unit 130B by executing programs loaded onto RAM 1200. The CPU 1100 of computer 1000 reads and executes these programs from recording medium 1800, but as another example, these programs may be obtained from other devices via a predetermined communication network.
[0212] <Other> Furthermore, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings can be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0213] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0214] Although some embodiments of the present application have been described in detail above with reference to the drawings, these are illustrative examples, and the present invention can be implemented in various other forms with modifications and improvements based on the knowledge of those skilled in the art, including the embodiments described in the section on the present invention.
[0215] 10 User device 50 Chatbot device 60 External device 1A Information processing system 100A Server device 120A Storage unit 121 Operation information storage unit 122A Input information storage unit 123A Instruction text information storage unit 124 Request text storage unit 130A Control unit 131A Acquisition unit 132 Judgment unit 133 Setting unit 134A Generation unit 135A Sending / receiving unit 136 Provisioning control unit 1B Information processing system 100B Server device 120B Storage unit 122B Input information storage unit 123B Instruction text information storage unit 125 Response information storage unit 126 Summary content storage unit 130B Control unit 131B Acquisition unit 135B Sending / receiving unit 137 Extraction unit 138 Provisioning unit
Claims
1. An information processing device comprising: an acquisition unit that acquires input information relating to a user's question or inquiry; a setting unit that sets additional information to be added to the input information based on situation information indicating the user's movement; a generation unit that generates an instruction sentence based on the input information and the additional information, which instructs the output of a sentence requesting a response from another user; and a control unit that controls the provision of the request sentence output by the generation model in response to the instruction sentence to the other user.
2. The information processing apparatus according to claim 1, further comprising a determination unit for determining additional items to be added to the input information, wherein the setting unit sets the additional information with content corresponding to the determined additional items.
3. The information processing apparatus according to claim 2, wherein the determination unit determines, based on the words contained in the input information, which type of location relating to the user's movement the input information relates to as an additional item, and the setting unit sets the additional information with content corresponding to the determined type of location.
4. The information processing apparatus according to claim 3, wherein the determination unit determines whether the location type is the current location or a destination, and the setting unit sets the additional information including location information indicating the current location and the current time when it is determined that the location type is the current location, and sets the additional information including location information indicating the destination and the estimated time of arrival at the destination when it is determined that the location type is a destination.
5. The information processing apparatus according to claim 1, wherein the user's status information includes in-vehicle information of the vehicle the user is riding in, and the setting unit sets the additional information which indicates the in-vehicle information.
6. The information processing apparatus according to claim 5, wherein the in-vehicle information is an image captured showing the interior of the vehicle, and the setting unit sets the additional information according to the attributes of an object included in the image, which is estimated by analysis of the image.
7. The information processing apparatus according to claim 1, wherein the generation unit determines whether the user desires a first response targeting the user's living area or a second response targeting an area outside the user's living area as a response to the input information, and generates an instruction statement with content corresponding to the determination result.
8. The information processing apparatus according to claim 7, wherein the generation unit generates an instruction statement instructing the other user to be a user located within the living area and to output a text requesting the first response, which includes insider information within the living area, if the user desires the first response, and generates an instruction statement instructing the other user to be a user located outside the living area and to output a text requesting a detailed second response, if the user desires the second response.
9. The information processing apparatus according to claim 1, wherein the control unit makes the request statement output by the generation model public to a predetermined service used by an unspecified number of users.
10. An information processing method performed by an information processing device, comprising: an acquisition step of acquiring input information relating to a user's question or inquiry; a setting step of setting additional information to be added to the input information based on situation information indicating the user's movement; a generation step of generating an instruction statement based on the input information and the additional information, which instructs the output of a sentence requesting a response from another user; and a control step of controlling the provision of the request statement output by the generation model in response to the instruction statement to the other user.
11. An information processing program to be executed by an information processing device, the program causing the information processing device to execute: an acquisition procedure for acquiring input information relating to a user's question or inquiry; a setting procedure for setting additional information to be added to the input information based on situation information indicating the situation relating to the user's movement; a generation procedure for generating an instruction statement based on the input information and the additional information, which instructs the device to output a sentence requesting a response from another user; and a control procedure for controlling the provision of the request statement output by the generation model in response to the instruction statement to the other user.