Information processing device, information processing method, and program
Patent Information
- Application Number
- PCT/JP2026/009315
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-12
- Filing Date
- 2026-03-11
- Publication Date
- 2026-09-17
Smart Images

Figure JP2026009315_17092026_PF_FP_ABST
Abstract
Description
Information processing apparatus, information processing method, and program
[0001] The present disclosure relates to dialogue technology.
[0002] Systems enabling voice-based interaction are provided in vehicles. In this regard, for example, Patent Document 1 discloses an in-vehicle device that allows making and receiving telephone calls through voice utterance.
[0003] Japanese Unexamined Patent Publication No. 08-307509
[0004] It is expected that natural language dialogue services will further increase in the future with the development of machine learning.
[0005] An object of the present disclosure is to provide information at appropriate timing in a voice dialogue system.
[0006] One aspect of an embodiment of the present disclosure is an information processing apparatus including a control unit that executes: acquiring first information to be provided to occupants of a vehicle; acquiring second information related to a situation of dialogue between the occupants of the vehicle that is occurring inside the vehicle; and outputting the first information at timing that does not hinder communication between the occupants of the vehicle, the timing being determined based on the second information.
[0007] One aspect of an embodiment of the present disclosure is an information processing method executed by an information processing apparatus, the method comprising: acquiring first information to be provided to occupants of a vehicle; acquiring second information related to a situation of dialogue between the occupants of the vehicle that is occurring inside the vehicle; and outputting the first information at timing that does not hinder communication between the occupants of the vehicle, the timing being determined based on the second information.
[0008] Further, as another aspect, there may be mentioned a program for causing a computer to execute the above method, or a computer-readable storage medium non-transitorily storing the program.
[0009] According to the present disclosure, information can be provided at appropriate timing in a voice dialogue system.
[0010] A schematic diagram of the dialogue system according to the first embodiment. Hardware configuration diagram of the in-vehicle device and sensor group. Software configuration diagram of the in-vehicle device. Diagram illustrating the timing of speech utterances in dialogue between occupants. Diagram showing the flow of processing performed by the control unit of the in-vehicle device. Flowchart of processing performed by the control unit of the in-vehicle device. Flowchart of processing performed by the determination unit. Flowchart of processing performed by the agent unit in the first embodiment. Diagram illustrating the timing of speech utterances in the third embodiment. Flowchart of processing performed by the agent unit in the third embodiment.
[0011] In recent years, with the advancement of machine learning, the number of products incorporating language models has increased. For example, by incorporating a Large Language Model (LLM), whose accuracy has been improved through machine learning using large datasets, into a product, it is possible to add natural language dialogue capabilities to the product.
[0012] Automobiles are a prime example of the usefulness of natural language interaction. For instance, by equipping in-vehicle systems with LLM (Language-Based Communication), information can be obtained without operating touch panels or other devices. Such functions are particularly useful when the driver's hands are occupied, such as while driving. Furthermore, the vehicle can proactively provide information to its occupants by uttering messages.
[0013] On the other hand, when speech from an in-vehicle device is in natural language, timing must be considered. For example, if multiple people are in a vehicle and are conversing with each other, speech from the in-vehicle device may disrupt communication between the occupants. The information processing device described in this disclosure solves such problems.
[0014] An information processing device according to a first aspect of the present disclosure includes a control unit that performs the following actions: acquiring first information to be provided to the occupants of a vehicle; acquiring second information relating to the status of conversations between the occupants of the vehicle taking place inside the vehicle; and outputting the first information at a timing determined based on the second information that does not interfere with communication between the occupants of the vehicle.
[0015] The information processing device can be, for example, a device mounted in a vehicle (in-vehicle device). The control unit of the information processing device acquires first information to be provided to the occupants of the vehicle. The first information may be generated in response to a request from the occupants (for example, search results for a facility) or it may be generated proactively by the in-vehicle device (for example, spontaneously suggesting stops along the route). The control unit also acquires second information regarding the status of conversations taking place inside the vehicle. The second information represents, for example, the status of conversations taking place between occupants, and typically includes audio data of utterances made by each occupant, text data indicating the content of the utterances, or status information determined based on these.
[0016] The control unit can determine, based on the second information, a timing that does not interfere with communication between the vehicle's occupants, and output the first information at that timing. The timing that does not interfere with communication between occupants may be, for example, a timing at which it can be determined that the conversation between occupants has ended or been temporarily interrupted. If the second information is text data, the timing may be determined based on the context of the conversation between occupants. If the second information is audio data, the timing may be determined based on the audio level, etc.
[0017] Furthermore, the first information may be provided to the first occupant by a virtual agent associated with the first occupant who is riding in the vehicle, and the control unit may output the first information at a time that does not interfere with communication by the first occupant. Alternatively, the first information may be provided to the associated occupant by any of the multiple virtual agents, each associated with a multiple occupant who is riding in the vehicle, and the control unit may output the first information at a time that does not interfere with communication by the occupant associated with the virtual agent that is providing the first information.
[0018] A virtual agent is a type of program that can be executed by an information processing device and is capable of natural language interaction. If multiple users may be riding in the vehicle, a virtual agent may exist for each of those users. In this case, the virtual agent may have data on the owner's personal information and preferences and may be configured to make personalized suggestions. The program and data for running the virtual agent may be stored in the information processing device, or they may be downloaded to the information processing device from a network and then executed. Furthermore, if the information processing device is an in-vehicle device, the virtual agent may be temporarily moved to the in-vehicle device via the network from a mobile device or other device carried by a user riding in the vehicle and executed there.
[0019] If the first piece of information is provided by a virtual agent associated with a particular occupant, the information may be output at a time that does not interfere with communication by that occupant. For example, if users A, B, and C are in a vehicle, and a virtual agent owned by user A is about to speak, the information may be provided even if users B and C are in conversation, as long as user A is not participating in the conversation.
[0020] Furthermore, the control unit may acquire third information relating to the vehicle's driving status and suppress the output of the first information during a period when the vehicle's driving status satisfies predetermined conditions. The first information is information directed to the vehicle's driver, and the control unit may acquire fourth information relating to the driver's workload and suppress the output of the first information during a period when the driver's workload satisfies predetermined conditions.
[0021] For example, even at times when communication between occupants is not hindered, it may be determined that information provision should be suppressed based on the vehicle's driving conditions or the driver's workload. For instance, if sensor data indicates that the driver's workload is high, it is preferable to suppress information provision.
[0022] If it would interfere with communication between occupants, the control unit may delay the output of the first information until a time comes when communication is not interfered with. With this configuration, for example, it becomes possible to wait until the conversation between occupants is interrupted before providing information. Whether or not the conversation has been interrupted for a predetermined time may be determined based on audio data acquired by the in-vehicle microphone or text data obtained by converting said audio data.
[0023] Furthermore, the first information may be information guiding the vehicle to a first point, and the control unit may perform predetermined processing if the timing does not arrive before the vehicle passes a second point located before the first point. The second point can be, for example, a point where guidance can be provided in time before reaching the first point. In other words, the control unit may perform predetermined processing if guidance becomes impossible while information provision is suppressed.
[0024] The prescribed processing may include outputting the first information without considering the second information. This form is effective when it is necessary to provide information even if it interrupts communication between crew members (for example, when the information is of high importance).
[0025] Furthermore, the prescribed processing may include ceasing the output of the first information. This form is effective when prioritizing not interrupting communication between crew members (for example, when the importance of the information is low).
[0026] Furthermore, the prescribed processing may include re-editing the content of the first information so that the information can be provided before the vehicle passes the first point. This allows the guidance to be provided in time, for example, by reducing the amount of information in the first information.
[0027] Furthermore, the predetermined process may include deciding whether to stop outputting the first information or to output the first information without considering the second information, based on the importance of the first information. In this way, the response may be switched depending on the importance of the first information.
[0028] The following describes specific embodiments of this disclosure with reference to the drawings. Unless otherwise specified, the hardware configurations, module configurations, functional configurations, etc., described in each embodiment are not intended to limit the technical scope of the disclosure to those configurations alone.
[0029] (First Embodiment) [System Overview] An overview of the dialogue system according to the first embodiment will be described. The dialogue system according to this embodiment is configured to include an in-vehicle device 10 mounted on a vehicle 1. The in-vehicle device 10 is configured to run a virtual agent that provides a dialogue service using natural language.
[0030] A virtual agent is a program that can be executed by the in-vehicle device 10 and is capable of natural language interaction. By running and keeping the virtual agent resident in the in-vehicle device 10, the occupants of the vehicle can receive information services in natural language at any time. A virtual agent may be prepared for each of the multiple users riding in vehicle 1. For example, if there are four users who ride in vehicle 1 on a daily basis, four virtual agents may be stored in the storage device of the in-vehicle device 10. Users riding in the vehicle can call up and use any virtual agent. The virtual agent may have data on the user's personal information and preferences and may be configured to make personalized suggestions.
[0031] The in-vehicle device 10 is installed in a connected vehicle that can communicate with any device via wireless communication. The in-vehicle device 10 may include a data communication module (DCM) for connecting the vehicle's components (e.g., ECU, in-vehicle terminal, etc.) to a network. The virtual agent operating in the in-vehicle device 10 can also collect information from external devices via the network.
[0032] The virtual agent can also provide information spontaneously. For example, if the virtual agent is running a tourist information service in the background, it may spontaneously make statements guiding passengers to nearby tourist attractions. In this embodiment, the virtual agent has the characteristic of using sensors installed in the vehicle 1 to determine a timing that does not interfere with communication between the occupants of the vehicle 1, and providing information when that timing arrives.
[0033] [Hardware Configuration] Next, the hardware configuration of the devices that make up the system will be described. Figure 2 is a schematic diagram showing an example of the hardware configuration of the vehicle 1 according to this embodiment.
[0034] Vehicle 1 is composed of an on-board device 10 and a sensor group 20. The on-board device 10 can be configured as a computer having a processor (CPU, GPU, etc.), main memory (RAM, ROM, etc.), and auxiliary storage (EPROM, hard disk drive, removable media, etc.). The auxiliary storage contains an operating system (OS), various programs, various tables, etc., and by executing the programs stored therein, various functions (software modules) that match a predetermined purpose, as described later, can be realized. However, some or all of the functions may be realized as hardware modules by hardware circuits such as ASICs and FPGAs.
[0035] The in-vehicle device 10 is configured as hardware and includes a control unit 11, a storage unit 12, a communication unit 13, an input / output unit 14, a wireless communication unit 15, and a location information acquisition unit 16.
[0036] The control unit 11 is a computing unit that realizes various functions of the in-vehicle device 10 by executing a predetermined program. The control unit 11 can be implemented by a hardware processor such as a CPU. The control unit 11 may also be configured to include RAM, ROM (Read Only Memory), cache memory, etc.
[0037] The storage unit 12 is a means for storing information and is composed of storage media such as RAM, magnetic disks, and flash memory. The storage unit 12 stores programs executed by the control unit 11, data used by those programs, and so on.
[0038] The communication unit 13 is a communication interface for connecting the in-vehicle device 10 to the in-vehicle network.
[0039] The input / output unit 14 is a unit that receives input from the user of the device and presents information to the user. Typically, the input / output unit 14 includes devices for inputting and outputting sound, such as a microphone or speaker. The input / output unit 14 may also include devices that provide visual information (such as a display).
[0040] The wireless communication unit 15 includes a communication module that performs wireless communication with a predetermined network. In this embodiment, the wireless communication unit 15 is configured to communicate with a predetermined mobile communication network. The wireless communication unit 15 may be configured to have an eUICC (e.g., a SIM card).
[0041] The location information acquisition unit 16 acquires location information of the vehicle 1. The location information acquisition unit 16 includes, for example, a GPS antenna and a positioning module for determining location information. The GPS antenna is an antenna that receives positioning signals transmitted from positioning satellites (also called GNSS satellites). The positioning module is a module that calculates location information based on the signals received by the GPS antenna.
[0042] The sensor group 20 is a collection of multiple sensors for acquiring sensor data used by the in-vehicle device 10. Examples of sensors include a vehicle speed sensor, a steering sensor, and a throttle sensor. The sensor group 20 may also include a microphone for acquiring audio data and a camera for acquiring image data.
[0043] The in-vehicle device 10 and the sensor group 20 are interconnected via a network bus such as a CAN (Controller Area Network).
[0044] [Software Configuration] Next, the software configuration of each device constituting the system will be described. FIG. 3 is a diagram schematically showing the software configuration of the in-vehicle device 10 according to the present embodiment.
[0045] In the present embodiment, the control unit 11 of the in-vehicle device 10 is configured to include three software modules: a dialogue input / output unit 111, an agent unit 112, and a determination unit 113. Each software module may be realized by executing a program stored in the storage unit 12 by the control unit 11 (such as a CPU). Note that the information processing executed by the software module is synonymous with the information processing executed by the control unit 11 (such as a CPU).
[0046] The dialogue input / output unit 111 acquires, via the input / output unit 14, utterances made by a user riding in a vehicle (hereinafter also referred to as an occupant). The dialogue input / output unit 111 performs predetermined processing on the acquired audio data and executes speech recognition. Thereby, the content of the utterance is converted into text. Further, the dialogue input / output unit 111 transmits the text obtained as a result of speech recognition to the agent unit 112.
[0047] Further, the dialogue input / output unit 111 outputs a response from the language model (hereinafter referred to as an answer sentence) transmitted from the agent unit 112. The dialogue input / output unit 111 converts the answer sentence into speech and outputs it via the input / output unit 14.
[0048] The agent unit 112 provides a dialogue service using a language model stored in the own device (also referred to as a local language model). By providing a dialogue service using the local language model, the agent unit 112 can behave as a virtual agent. The virtual agent may be realized, for example, by a combination of a program and a local language model. Further, different local language models may be used for each occupant. For example, by using a local language model corresponding to user A, the agent unit 112 can behave as a virtual agent for user A, and by using a local language model corresponding to user B, the agent unit 112 can behave as a virtual agent for user B.
[0049] These local language models may be stored in the in-vehicle device 10 (storage unit 12), or may be copied from a mobile terminal owned by the occupant via wireless communication or the like every time the virtual agent is executed. Further, the local language model may be downloaded from a network (such as a cloud server) every time the virtual agent is executed.
[0050] Note that the agent unit 112 may collect information from one or more external devices available via a network when specialized knowledge is required for dialogue. For example, when it is determined that an occupant who has made an utterance is seeking information on a facility or a store, the agent unit 112 may access a language model (also referred to as a remote language model) accessible via the network and capable of providing regional information to acquire information.
[0051] The agent unit 112 can provide dialogue in response to the occupant's speech and can also perform background services. Background services are services that reside in memory and run in the background, such as route guidance services and tourist information services. In some cases, the in-vehicle device 10 may spontaneously provide information if certain conditions are met while a background service is running (for example, when approaching an intersection where a right or left turn is required). In such cases, the agent unit 112 determines the appropriate timing for speech based on information obtained from the determination unit 113, which will be described later.
[0052] The determination unit 113 provides information to the agent unit 112 for determining the timing of speech. Figure 4 is a diagram illustrating the timing of speech. Here, it is assumed that occupant A and occupant B are having a conversation inside vehicle 1. The determination unit 113 can determine which occupant is speaking by performing speech recognition using a sensor included in the sensor group 20 (for example, an in-vehicle microphone). If a predetermined period of time has passed during which neither occupant has spoken, the determination unit 113 can determine that the conversation between the occupants has ended (i.e., the virtual agent is able to speak). The determination unit 113 can provide such a determination result to the agent unit 112. The determination result may be something like "occupants are having a conversation" or "the conversation between occupants has ended," or it may further include identifiers of the occupants who are having the conversation.
[0053] In this example, we have given an example of determining whether or not a vehicle occupant is speaking based on the results of speech recognition, but the determination unit 113 may make a similar determination based on the speech level. In addition, the determination unit 113 may determine the context of the dialogue based on the results of speech recognition and determine from the determined context whether a dialogue between occupants has started or ended.
[0054] The storage unit 12 of the in-vehicle device 10 stores multiple local language models corresponding to multiple occupants. A local language model is a language model that has been trained to enable natural language dialogue tasks. The multiple local language models may be trained for each user, or they may contain the personal information of the target user. For example, the local language model corresponding to user A may know user A's preferences and personal information, and the local language model corresponding to user B may know user B's preferences and personal information.
[0055] [Processing Flow] Next, an overview of the processing performed by the control unit 11 will be explained. Figure 5 is a diagram showing the processing flow performed by the control unit 11 of the in-vehicle device 10.
[0056] The dialogue input / output unit 111 acquires speech from the vehicle occupants via the input / output unit 14. For example, the input / output unit 14 converts speech acquired via a microphone or the like into voice data, which is then acquired by the dialogue input / output unit 111. The dialogue input / output unit 111 performs a predetermined speech recognition process on the acquired voice data and converts the voice data into text. While holding off on responding to the speech, the dialogue input / output unit 111 transmits the text obtained as a result of the speech recognition to the agent unit 112. This text will hereafter be referred to as the "speech text".
[0057] Upon receiving a spoken sentence, the agent unit 112 performs two types of processing depending on the type of utterance. If the utterance made by the crew member requests a dialogue, the agent unit 112 transfers the acquired utterance to the local language model corresponding to the currently running virtual agent and obtains a response. For example, if the currently running virtual agent corresponds to user A, the agent unit 112 transfers the utterance to the local language model corresponding to user A. If the agent unit 112 determines that the local language model cannot provide sufficient information, it may transfer the utterance to an external language model (remote language model) via the network. In this case, the agent unit 112 obtains a response from the local or remote language model and outputs it as audio data.
[0058] On the other hand, utterances made by the crew may also be requests for the execution of background services, such as a request to start route guidance. In this case, the agent unit 112 executes the service in the background in response to the request. In the following explanation, services executed in the background, such as route guidance, will be referred to as background services. Background services may be executed, for example, while receiving information from the location information acquisition unit 16.
[0059] On the other hand, background services may provide information to the occupants. For example, when route guidance is being performed, the background service may issue notifications such as instructions to turn right or left. When the agent unit 112 receives information from the background service for the occupants, it outputs it by voice. By performing this action, the agent unit 112 can give the occupants of the vehicle the impression that they are receiving route guidance from a virtual agent.
[0060] On the other hand, as mentioned above, if passengers are conversing with each other inside the vehicle, the virtual agent may interrupt and disrupt the communication. Therefore, the agent unit 112 obtains information about the communication between passengers from the determination unit 113 and adjusts the timing of its speech based on this information. As a result, as shown in Figure 4, it becomes possible to provide information at a time that does not disrupt the communication between passengers.
[0061] [Flowchart] Next, the details of the processes performed by the in-vehicle device 10 will be described. Figure 6 is a flowchart of the processes performed by the in-vehicle device 10. The processes shown in Figure 6 are started when a vehicle occupant makes a speech. The start of a speech may be detected, for example, by a predetermined keyword.
[0062] First, in step S11, the dialogue input / output unit 111 recognizes the content of the utterance. The dialogue input / output unit 111 acquires the audio data output from the input / output unit 14 and performs speech recognition processing to convert the utterance into text. The converted text (utterance) is sent to the agent unit 112.
[0063] In step S12, the agent unit 112 determines whether the content of the utterance requests a dialogue or requests the execution of a background service. If the content of the utterance requests a dialogue, the process proceeds to step S13. If the content of the utterance requests the execution of a background service, the process proceeds to step S15.
[0064] In step S13, the agent unit 112 uses the local language model to obtain a response to the utterance. If a virtual agent is provided for each of multiple users, the agent unit 112 may use the local language model corresponding to the currently running virtual agent. If the local language model cannot provide sufficient information, the agent unit 112 may use the remote language model to obtain a response to the utterance. The obtained response is output as audio data via the input / output unit 14 in step S14.
[0065] If the process proceeds to step S15, the agent unit 112 starts executing the specified background service. The execution of the background service continues until it is instructed to terminate or until predetermined conditions are met (for example, in the case of a route guidance service, until the vehicle arrives at its destination).
[0066] Next, we will explain the process of providing information generated from background services to the vehicle's occupants.
[0067] As mentioned above, the determination unit 113 provides the agent unit 112 with information regarding communication between occupants in real time. Figure 7 is a flowchart of the processes executed by the determination unit 113. The illustrated processes are repeatedly executed while the vehicle is in motion.
[0068] First, in step S21, the determination unit 113 acquires sound from inside the vehicle via a sensor included in the sensor group (for example, an in-vehicle microphone), and recognizes the content of speech made by the occupants based on the acquired sound. If the determination unit 113 can identify the user who made the speech, the content of the speech may be associated with the identifier of the user who made the speech.
[0069] Next, in step S22, the determination unit 113 determines the status of a conversation between users, who are occupants of the vehicle, based on the content of the utterances, and generates a determination result (hereinafter referred to as conversation status data). The conversation status data may, for example, indicate whether or not a conversation is currently taking place between occupants. For example, if a period of time in which no occupants have spoken continues for a predetermined amount of time or longer, the determination unit 113 may determine that a conversation is not currently taking place between occupants. Alternatively, if the content of the utterances is analyzed and it is determined that the conversation has reached a certain conclusion, the determination unit 113 may determine that a conversation is not currently taking place between occupants. The conversation status data is transmitted to the agent unit 112 in step S23.
[0070] The dialogue status included in the dialogue status data may be represented by a binary value, such as "dialogue is currently taking place" or "dialogue is not currently taking place," or it may be the result obtained by analyzing the speech status for each of the multiple passengers. For example, if three users, A, B, and C, are in the vehicle, the determination unit 113 may include the elapsed time since each user last spoke in the dialogue status data and send it to the agent unit 112. The determination unit 113 may also include a flag indicating whether or not each user is in a dialogue in the dialogue status data and send it to the agent unit 112. This allows the agent unit 112 to make a determination such as, "Users B and C are in a dialogue, but User A is not participating in the dialogue."
[0071] Next, we will explain the process by which the agent unit 112 provides information based on the dialogue status data received from the determination unit 113. Figure 8 is a flowchart of the process executed by the agent unit 112 when information for the crew is generated from a background service that is currently running.
[0072] First, in step S31, the agent unit 112 receives dialogue status data from the determination unit 113. Next, in step S32, the agent unit 112 determines whether or not information can be provided based on the dialogue status data. The agent unit 112 determines that information can be provided if, for example, any of the following conditions are met.
[0073] (1) When it is determined that all crew members are not currently communicating, this is because if no crew members are communicating, it will not hinder communication between crew members. (2) When a conversation is taking place, but the agent unit 112 determines that the crew member corresponding to the running virtual agent is not participating in the conversation, even if there are crew members communicating with each other, if the owner of the running virtual agent is not participating in the conversation, it is possible to provide information to that owner.
[0074] If it is determined in step S32 that information can be provided, the process proceeds to step S33. If it is determined in step S32 that information cannot be provided, the process returns to step S31. In other words, the process described above is repeated while information provision remains suppressed.
[0075] In step S33, the agent unit 112 outputs information generated from the background service via the dialogue input / output unit 111. This provides information to the vehicle occupants.
[0076] As described above, the in-vehicle device 10 according to this embodiment determines the status of conversations between occupants inside the vehicle and provides information when no conversation is taking place. This makes it possible to provide information without interfering with communication between occupants.
[0077] (Second Embodiment) In the first embodiment, the timing of information provision was determined based on the conversation status of the occupants. On the other hand, there are times when information provision should be refrained from even if the occupants are not conversing with each other, such as when the driver's workload is high. The second embodiment is an embodiment in which information provision is suppressed when the vehicle or driver's workload conditions meet predetermined conditions.
[0078] In the second embodiment, the agent unit 112 acquires data related to the vehicle's driving status and / or the driver's load status from any sensor included in the sensor group. Examples of such data include vehicle speed, turn signal operation status, and steering angle. Furthermore, if sensors that sense the driver's gaze or biometric information are available, the driver's load status can also be estimated from the data output by these sensors.
[0079] In the second embodiment, the agent unit 112 suppresses the provision of information during periods when the vehicle's driving conditions meet predetermined conditions, or during periods when the driver's workload meets predetermined conditions. This process can be additionally performed in step S32. With this configuration, for example, information provision can be suppressed during periods when the driver's workload is estimated to be high, such as when the vehicle is turning right or left, changing lanes, or when the driver is performing safety checks, thereby improving safety.
[0080] (Third Embodiment) In the first and second embodiments, the in-vehicle device 10 delayed providing information until the timing for providing the information arrived. However, there are cases where delaying information provision is not appropriate, such as when providing guidance at turning points or branching points. In such cases, delaying information provision could cause the vehicle to pass the point related to the guidance. To address this, in the third embodiment, a deadline for providing information is set, and information is provided in a manner that takes this deadline into consideration.
[0081] Figure 9 illustrates the timeline when providing guidance to points along a route (for example, turning points or branching points). In the figure, t0 is the time when information provision should ideally begin, and t2 is the time when the vehicle passes the point related to guidance (the first point). Furthermore, t1 is the time obtained by subtracting the time required for information provision (such as voice output) from t2 (referred to as the deadline time). If time t1 is the time when vehicle 1 passes the second point, then in order to provide proper route guidance, information provision should begin before vehicle 1 passes the second point. However, there are cases where the occupants of the vehicle are conversing with each other at the time when information provision should ideally begin, and the conversation does not end by the deadline time. In this case, the agent unit 112 can take one of the following three actions.
[0082] (1) If the deadline for providing information unconditionally has passed, this method will start providing information regardless of the status of the crew's conversation. This method is effective when the importance of the information to be provided is high. (2) If the deadline for discontinuing information provision has passed, this method will discontinue providing information altogether. This method is effective when the importance of the information to be provided is low. (3) If the deadline for providing information after reducing the amount of information provided has passed, this method will provide information after reducing the amount of information provided. For example, the agent unit 112 shortens the length of its utterance so that information provision can be completed before the vehicle reaches the location related to the guidance. Alternatively, the agent unit 112 provides only the minimum amount of information so as not to interfere with communication between the crew members.
[0083] Figure 10 is a flowchart of the process performed by the agent unit 112 in the third embodiment. Steps that are the same as those shown in the flowchart in Figure 8 are shown with dashed lines and their explanations are omitted.
[0084] In this embodiment, if it is determined in step S32 that information cannot be provided, the process proceeds to step S34. In step S34, it is determined whether the information to be provided guides to a point along the route. If the information to be provided does not guide to a point along the route, no special processing is required, and the process returns to step S31. If the information to be provided guides to a point along the route, the process proceeds to step S35, where it is determined whether the deadline time has passed. If it is determined that the deadline time has passed, the process returns to step S31.
[0085] If it is determined that the deadline has passed, the process proceeds to step S36, where adjustment processing is performed. The adjustment processing can be any of the above-mentioned (1) to (3). The agent unit 112 may decide which of (1) to (3) to select based on the importance of the information to be provided. For example, it may select (1) if the information to be provided is of high importance, (3) if the importance is moderate, and (2) if the importance is low.
[0086] (Modifications) The embodiments described above are merely examples, and this disclosure may be modified as appropriate without departing from its essence. For example, the processes and means described in this disclosure can be freely combined and implemented as long as no technical inconsistencies arise.
[0087] Furthermore, a process described as being performed by a single device may be divided and executed by multiple devices. Conversely, a process described as being performed by different devices may be executed by a single device. In a computer system, the hardware configuration (server configuration) by which each function is implemented can be flexibly changed.
[0088] The present disclosure can also be realized by supplying a computer program implementing the functions described in the embodiments above to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer by a non-temporary computer-readable storage medium that can be connected to the computer's system bus, or it may be provided to the computer via a network. Non-temporary computer-readable storage mediums include, for example, any type of disk such as magnetic disks (floppy disks, hard disk drives (HDDs), etc.), optical disks (CD-ROMs, DVDs, Blu-ray discs, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, optical cards, and any type of medium suitable for storing electronic instructions.
[0089] 10...In-vehicle device 11...Control unit 12...Storage unit 13...Communication unit 14...Input / output unit 15...Wireless communication unit 16...Location information acquisition unit
Claims
1. An information processing device having a control unit that performs the following: acquiring first information to be provided to the occupants of a vehicle; acquiring second information regarding the status of conversations between the occupants of the vehicle taking place inside the vehicle; and outputting the first information at a timing determined based on the second information that does not interfere with communication between the occupants of the vehicle.
2. The information processing apparatus according to claim 1, wherein the first information is provided to the first occupant by a virtual agent associated with the first occupant who is riding in the vehicle, and the control unit outputs the first information at a timing that does not interfere with communication by the first occupant.
3. The information processing apparatus according to claim 1, wherein the first information is provided to a passenger by one of a plurality of virtual agents each associated with a plurality of passengers currently riding in the vehicle, and the control unit outputs the first information at a timing that does not interfere with communication by the passenger associated with the virtual agent that intends to provide the first information.
4. The information processing apparatus according to claim 1, wherein the control unit acquires third information relating to the driving status of the vehicle, and suppresses the output of the first information during a period in which the driving status of the vehicle satisfies predetermined conditions.
5. The information processing apparatus according to claim 1, wherein the first information is information directed to the driver of the vehicle, the control unit acquires fourth information relating to the driver's load status, and suppresses the output of the first information during a period when the driver's load status satisfies predetermined conditions.
6. The information processing apparatus according to claim 1, wherein the control unit delays the output of the first information until the timing arrives.
7. The information processing apparatus according to claim 1, wherein the first information is information that guides the vehicle to a first point through which it passes, and the control unit outputs the first information without considering the second information if the timing does not occur before the vehicle passes a second point located before the first point.
8. The information processing apparatus according to claim 1, wherein the first information is information that guides the vehicle to a first point through which it passes, and the control unit stops outputting the first information if the timing does not arrive before the vehicle passes a second point located before the first point.
9. The information processing apparatus according to claim 1, wherein the first information is information that guides the vehicle to a first point through which it passes, and the control unit re-edits the content of the first information so that it can provide information through the time the vehicle passes the first point if the timing does not arrive before the vehicle passes a second point located before the first point.
10. The information processing apparatus according to claim 1, wherein the first information is information that guides the vehicle to a first point through which it passes, and the control unit determines, based on the importance of the first information, whether to stop outputting the first information or to output the first information without considering the second information, if the timing does not arrive before the vehicle passes a second point located before the first point.
11. The information processing device according to claim 1, wherein the second information is voice data acquired by an in-vehicle microphone in the vehicle, and the control unit determines that the timing has arrived when the conversation between the occupants of the vehicle is interrupted for a predetermined period of time or longer.
12. The information processing apparatus according to claim 1, wherein the second information is text data obtained by converting audio data acquired by an in-vehicle microphone in the vehicle, and the control unit determines the timing of a break in conversation between the occupants of the vehicle based on the results of analyzing the text data.
13. An information processing method comprising: an information processing device acquiring first information to be provided to the occupants of a vehicle; acquiring second information relating to the status of conversations between the occupants of the vehicle taking place inside the vehicle; and outputting the first information at a timing determined based on the second information that does not interfere with communication between the occupants of the vehicle.
14. The information processing method according to claim 13, wherein the first information is provided to the first occupant by a virtual agent associated with the first occupant who is riding in the vehicle, and the first information is output at a time that does not interfere with communication by the first occupant.
15. The information processing method according to claim 13, wherein the first information is provided to a passenger by one of a plurality of virtual agents each associated with a plurality of passengers currently riding in the vehicle, and the first information is output at a timing that does not interfere with communication by the passenger associated with the virtual agent that intends to provide the first information.
16. The information processing method according to claim 13, which includes acquiring a third piece of information relating to the driving status of the vehicle, and suppressing the output of the first piece of information during a period in which the driving status of the vehicle satisfies predetermined conditions.
17. The information processing method according to claim 13, wherein the first information is information directed to the driver of the vehicle, a fourth piece of information relating to the driver's load status is acquired, and the output of the first information is suppressed during a period in which the driver's load status satisfies predetermined conditions.
18. The information processing method according to claim 13, wherein the output of the first information is delayed until the aforementioned timing arrives.
19. A program for causing a computer to execute the information processing method described in any one of claims 13 to 18.