Information processing method, computer, program, communication terminal, and interactive agent system
The information processing method ensures seamless interaction with a personalized AI agent in a vehicle by acquiring authentication and access information to set the AI agent's character on a vehicle-mounted computer, addressing the challenge of maintaining user preferences during the transition from a communication terminal.
Patent Information
- Application Number
- JP2024101963
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2026-01-14
AI Technical Summary
Conventional dialogue agent systems face challenges in seamlessly continuing interactions with a personalized AI agent attuned to an individual user's preferences and characteristics when transitioning from a communication terminal to a vehicle environment.
An information processing method that acquires authentication and access information from a user's communication terminal, retrieves agent attribute information, and sets a GUI or VUI on a vehicle-mounted computer to represent the AI agent's character, enabling continuous interaction using the AI agent's attributes and multimodal data processing.
Enables seamless continuation of user interactions with a personalized AI agent within a vehicle environment, allowing the user to access their regular AI agent with minimal disruption and maintaining individualized preferences and characteristics.
Smart Images

Figure 2026003870000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing method, a computer, a program, a communication terminal, and a dialogue agent system. [Background technology]
[0002] Patent Document 1 discloses an in-vehicle system in which the presence of an agent is planned. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-117302 Summary of the Invention [Problem to be solved by the invention]
[0004] Further improvements are needed in conventional dialogue agent systems. [Means for solving the problem]
[0005] An information processing method according to one embodiment of the present disclosure is an information processing method executed on a first computer installed in a vehicle in an interactive agent system capable of interacting with a user, in which the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, then acquires access information for accessing the AI agent used by the user from the communication terminal, acquires agent attribute information for representing the AI agent's character from an agent database capable of communicating with the first computer based on the access information, sets at least one of a GUI and a VUI representing the AI agent's character in the first computer based on the agent attribute information, and, while connecting to a second computer on which the AI agent is implemented based on the access information, performs interactive processing with the user using at least one of the GUI and VUI and generation processing by the AI agent. [Effects of the Invention]
[0006] According to an information processing method according to one aspect of the present disclosure, further improvements are realized in a dialogue-capable agent linked to a user. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a diagram showing an example of the overall configuration of a dialogue agent system according to this embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of the interactive agent system according to the present embodiment. [Figure 3] FIG. 3 is a sequence diagram showing an example of API exchange between an information terminal and a cloud via a wide area communication network in the dialogue agent system according to this embodiment. [Figure 4] FIG. 4 is a sequence diagram showing an example of the flow of processing in which a client downloads and uses agent attribute information in the interactive agent system according to this embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of a functional configuration of the dialogue agent system according to the present embodiment. [Figure 6] FIG. 6 is a sequence diagram showing an example of a processing flow for continuing interaction with an agent even when the client is switched in the interactive agent system according to this embodiment. [Figure 7-1] FIG. 7-1 is a sequence diagram showing an example of a process flow for continuing interaction with an agent even when the client is switched in the interactive agent system according to this embodiment. [Figure 7-2] FIG. 7-2 is a sequence diagram showing an example of a process flow for continuing interaction with an agent even when the client is switched in the interactive agent system according to this embodiment. [Figure 8] FIG. 8 is a sequence diagram showing an example of a processing flow for enhancing the interaction between the client and the agent in the interactive agent system according to this embodiment. [Figure 9] FIG. 9 is a flowchart showing an example of the flow of a process for querying and transmitting multimodal data types input to the interactive agent system according to this embodiment. [Figure 10] FIG. 10 is a diagram for explaining an example of the structure of additional data that has been converted from multimodal raw data into text data in the interactive agent system according to this embodiment. [Figure 11] FIG. 11 is a diagram for explaining an example of the structure of data converted from multimodal raw data into text data in the interactive agent system according to this embodiment. [Figure 12] FIG. 12 is a diagram for explaining an example of a dialogue agent system according to this embodiment. [Figure 13] FIG. 13 is a diagram for explaining an example of a dialogue agent system according to this embodiment. [Figure 14]FIG. 14 is a diagram for explaining an example of a dialogue agent system according to this embodiment. [Figure 15] FIG. 15 is a diagram for explaining an example of multimodal data in the dialogue agent system according to this embodiment. [Figure 16] FIG. 16 is a diagram illustrating an example of a functional configuration of the dialogue agent system according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] [Findings that formed the basis of this disclosure] In recent years, technological development of AI agents using large-scale language models (LLMs) has progressed. AI agents possess short-term and long-term memory (user usage logs or portions of their contents) and can autonomously communicate with external applications and web services over a network, launching and operating other applications and web services. This allows AI agents to set or update goals through text or voice communication with users (instructions to the AI are also called prompts), autonomously generate tasks necessary to achieve those goals, and execute the information processing of the generated tasks sequentially, either autonomously or through communication with the user, to achieve their final goal. AI is a computer system, or software program, that can process information input or output not only in a single modal (data format) such as text, but also in a combination of multiple modalities such as voice and images. This is called multimodal AI.
[0009] AI agents can be given specializations and characteristics depending on the databases they refer to when processing information and the algorithms they use to generate tasks. This allows them to be implemented as highly specialized agents specializing in specific functions, for example. On the other hand, AI agents can also be implemented as personalized AI agents that get close to individual users by learning the preferences, biometric information, and past behavioral history of the users they communicate with, and by accessing databases that store such personal user data (hereinafter referred to as user attribute information). The former type of AI agent is sometimes called a specialized agent, and the latter type of AI agent is sometimes called a partner-type agent.
[0010] Partner-type agents are thought to be particularly effective partners in mobility spaces. This is because when a user travels by vehicle to a place outside their usual range of activity and has a new (or unusual) experience, a partner-type agent can act as an appropriate navigator that is attuned to the user's individuality. For example, when navigating a vehicle's route, a partner-type agent can make selections or suggestions that reflect the user's preferences. Examples of user preferences include whether the user prefers the shortest route, roads that are easy to drive on with separated sidewalks, or whether they like to stop by tourist spots.
[0011] The inventors have studied a series of user experiences related to vehicle use and the use of an AI agent. The initial scenario envisioned is a scenario in which a user uses an AI agent using a smartphone or other device before getting into the vehicle, and then continues to use the AI agent after getting into the vehicle using an information device (e.g., in-vehicle infotainment; IVI). In such a scenario, there is a need for seamless continuation of interactions with a partner-type agent that is attuned to the individual user, even after getting into the vehicle. However, there are challenges in terms of customer value and implementation methods regarding how to appropriately configure a personal partner-type agent for a vehicle used by a user for the first time or a vehicle used by multiple users.
[0012] The following aspects of the present disclosure are based on the above findings, but the inventions described in the claims are not limited to the above findings.
[0013] [Outline of the embodiment] An information processing method according to one aspect of the present disclosure is an information processing method executed by a first computer mounted on a vehicle in an interactive agent system capable of interacting with a user. The information processing method includes the following steps: the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, and then acquires access information for accessing an AI agent used by the user from the communication terminal; acquires agent attribute information for representing a character of the AI agent from an agent database capable of communicating with the first computer based on the access information; sets at least one of a GUI and a VUI representing the character of the AI agent in the first computer based on the agent attribute information; and, while connecting to a second computer on which the AI agent is implemented, executes an interactive process with the user using at least one of the GUI and the VUI and a process generated by the AI agent based on the access information.
[0014] According to this information processing method, a first computer installed in a vehicle can obtain agent attribute information related to the character of an AI agent used by a user via the user's communication terminal and set the information in the first computer's GUI (Graphical User Interface) and / or VUI (Voice User Interface). For example, if a second computer on which an AI agent is implemented is a communication terminal, the communication terminal can function as the AI agent's brain, and the first computer can function as the AI agent's body (face and / or voice). For example, if the second computer on which an AI agent is implemented is a server, the server can function as the AI agent's brain, and the first computer can function as the AI agent's body (face and / or voice). This allows the AI agent that a user regularly uses via a communication terminal to be represented on the first computer installed in the vehicle and used within the vehicle.
[0015] The process from obtaining access information to setting the AI agent's character may be triggered by determining that the authentication information is valid. This allows the process from authentication to setting the AI agent's character to be executed smoothly. For example, even when the user is riding in a vehicle for the first time or in a vehicle that requires the use of another AI agent's character, the user can call and use their partner AI agent in the vehicle through a simple procedure, automatically, or in a short time. If the vehicle unlocking process is triggered by authentication, the AI agent's setting can be completed before or immediately after the user sits in the seat.
[0016] An information processing method according to one aspect of the present disclosure is an information processing method executed by a communications terminal capable of communicating with a first computer mounted on a vehicle and implementing an AI agent used by the user in a dialogue agent system capable of dialogue with a user. The information processing method includes: transmitting authentication information for the user to use the vehicle to the first computer; after determining that the authentication information is valid, transmitting access information for accessing the AI agent to the first computer; transmitting agent attribute information for representing a character of the AI agent to the first computer to cause the first computer to set at least one of a GUI and a VUI based on the agent attribute information; and, while connected to the first computer, executing dialogue processing with the user using at least one of the GUI and the VUI and generation processing by the AI agent.
[0017] According to this information processing method, the AI agent in the communication terminal that the user uses on a daily basis can be used while being represented on the first computer installed in the vehicle.
[0018] An information processing method according to one aspect of the present disclosure is an information processing method executed by a first computer installed in a vehicle in a dialogue agent system capable of dialogue with a user. The information processing method includes: establishing a connection to a communication terminal implementing a multimodal AI agent after the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid; converting sensing data acquired via one or more sensors installed in the vehicle while the user is using the vehicle into output data in a format and type compatible with the AI agent for each unit of data acquired over a predetermined period; periodically or intermittently transmitting the output data to the AI agent in the communication terminal and causing the AI agent to execute dialogue processing reflecting the latest output data.
[0019] According to this information processing method, an AI agent installed in a communication terminal can receive information sensed by a vehicle and carry out dialogue with the user. In this case, output data downsized to a format and type that the AI agent can handle is transmitted to the communication terminal, rather than the actual sensing data (e.g., video and audio) detected by the vehicle. This allows the AI agent in the communication terminal to carry out dialogue while appropriately incorporating information from the vehicle.
[0020] [Embodiment Mode] Hereinafter, exemplary embodiments of an information processing method, a computer, a program, a communication terminal, and a dialogue agent system according to the present disclosure will be described with reference to the drawings.
[0021] First, we will explain the definition of an AI agent (Artificial Intelligence Agent). An AI agent is software or a mechanism for achieving a predefined goal. An AI agent is designed to autonomously generate, select, and execute actions to achieve the goal based on communication with the user, the situation in the user's surrounding space, and external information acquired via a network regarding the interaction with the user. Communication with the user includes all means for conveying the user's emotions, will, and thoughts. For example, this includes one or more of the following: conversation, timing before speaking, tone of voice, visual information (GUI) such as text and symbols, physical operation such as buttons and switches, and bodily actions such as facial expressions, gaze, posture, or gestures. In this embodiment, an AI agent is also simply referred to as an agent. To facilitate communication with the user, an agent may be represented as a unique character, in which case it has the character's attribute information. The agent's attribute information includes, for example, information about one or more of the character's appearance, body, clothing, accessories, gestures, facial expressions, voice, personality, preferences, habits, knowledge, or experience (records of past interactions with users).
[0022] FIG. 1 is a diagram showing an example of the overall configuration of a dialogue agent system according to this embodiment (an agent system that communicates with a user as described above is referred to as a "dialogue agent," "dialogue agent system," or simply an "agent" in this disclosure). The dialogue agent system according to this embodiment is an example of a dialogue agent system that can dialogue with a user (the communication between a user and an agent system is not limited to dialogue, and the above-mentioned communication may also take place; however, for ease of reading, the phrase "capable of dialogue" will be used in this disclosure rather than "capable of communication"). As shown in FIG. 1, the dialogue agent system according to this embodiment includes an information terminal 1, a vehicle 2, a cloud 3, etc. The information terminal 1, the vehicle 2, and the cloud 3 are connected to each other so as to be able to communicate with each other via a network (for example, a wide area communication network such as the Internet).
[0023] The user's information terminal 1 stores information about an electronic key for unlocking and starting a vehicle 2. The agent operates through AI inference processing performed by devices such as the information terminal 1, the vehicle 2, and the cloud 3.
[0024] The cloud agent is an agent that processes requests on the cloud 3 and generates responses. The corporate agent and private agent are partner-type agents that run AI inference processing on an information terminal 1 such as a smartphone. Specifically, the corporate agent is a work partner of the user, and is an agent that is managed in association with a corporate electronic key for using the vehicle 2 for work purposes. The private agent is a private partner, and is an agent that is managed in association with a personal electronic key for using the vehicle 2 for private purposes.
[0025] A vehicle agent is an agent running on an in-vehicle system (a computer system installed in vehicle 2). A cloud agent is an agent that operates in a server-client model via the Internet. Corporate agents, private agents, and vehicle agents are agents that can operate using only the computer resources and data of the device, even when there is no Internet connection. When any agent has an Internet connection, it may perform data searches on the Internet to obtain information useful for generating a response and generate the response.
[0026] Cloud agents, corporate agents, and private agents are agents that can be used by users from any device (client) and can interact with the user after gaining a deep understanding of the user's behavior, thoughts, and experiences by working together with the user while completing daily tasks. Vehicle agents are agents that can be used in in-vehicle systems and are specialized for vehicle-related tasks with unique functions such as supporting safe driving, setting and updating routes according to traffic conditions, and answering questions and setting instructions about the vehicle.
[0027] 2 is a diagram showing an example of the hardware configuration of the interactive agent system according to this embodiment. Information terminal 1 includes a sensor unit 101 for acquiring video information, audio information, and / or physical quantities of the surrounding environment, a UI unit 103 for providing the user with video and audio information and accepting button presses, touch operations, etc., a calculation unit 104 for performing various calculations including learning and inference processing of AI models performed within information terminal 1 and information processing such as information drawing, a memory 105 for storing data and files used by calculation unit 104, and a communication unit 106 for communicating with other computers on a communication network.
[0028] The UI unit 103 has a display that displays a GUI (Graphical User Interface), and a speaker and microphone that input and output a VUI (Voice User Interface). The calculation unit 104 is an example of an AI processor that causes an agent (e.g., a partner-type agent) to execute a generation process. The memory 105 is an example of a memory that stores access information including the address of the agent and agent attribute information related to the agent. The memory 105 is also an example of a memory that stores a usage log of the agent by the user. The memory 105 is also an example of a memory that stores a usage log related to one or more agents available to the user. The memory 105 also manages (stores) information related to multiple electronic keys available to the user. Here, the multiple electronic keys include a company electronic key associated with the company to which the user belongs and the user's personal electronic key.
[0029] When an application for managing keys is installed in the information terminal 1, the program and necessary data are recorded in the memory 105 of the information terminal 1, and the program is executed by the calculation unit 104.
[0030] In this embodiment, the information terminal 1 is described as a smartphone, but is not limited to this. It may also be in the form of a wristwatch-type smart watch, smart glasses-type eyeglasses, smart earphones worn on the ears, a smart ring-type finger ring, a smart speaker operated by voice, or a robot with moving parts.
[0031] The cloud 3 includes a communication unit 301 for communicating with other computers on a communication network (wide area communication network), a memory 302 that stores information about the vehicle 2 and the user and its management program, and a calculation unit 303 that performs various data processing.
[0032] Vehicle 2 has a movable unit 201 for moving vehicle 2 and operating equipment (seats, etc.) within the vehicle cabin, a lighting unit 202 for illuminating the area around vehicle 2, a sensor unit 203 for detecting the position and status of people and cars around vehicle 2, as well as people and objects within the vehicle cabin, a UI unit 204 for providing passengers with various video and audio information and accepting input from passengers such as touch operations and voice operations, a key control unit 205 for authenticating the key to be unlocked and controlling the locking / unlocking of vehicle 2's doors, a calculation unit 206 for executing various processes related to the vehicle's core system and vehicle functions, a memory 207 for recording various data including the vehicle's core system's program and key management database, and a communication unit 208 for wireless communication with external devices.
[0033] In this embodiment, the UI unit 204, the key control unit 205, the calculation unit 206, the memory 207, and the communication unit 208 are realized by an in-vehicle system mounted on the vehicle 2. Here, the calculation unit 206 is an example of a processor, and may be a processor that executes a vehicle agent different from the agent. Also, here, the memory 207 is an example of a memory that stores a program for causing the calculation unit 206 to execute predetermined information processing. Here, the predetermined information processing includes processing executed by the acquisition unit 206a, the setting unit 206b, and the execution unit 206c (see FIG. 5), which will be described later.
[0034] The vehicle 2, the information terminal 1, and the cloud 3 may communicate with each other via a communication means other than the wide area communication network Internet. For example, the unlocking authentication process performed between the vehicle 2 and the information terminal 1 may use short-range wireless communication.
[0035] FIG. 3 is a sequence diagram showing an example of API exchanges between an information terminal and a cloud via a wide area communication network in a dialogue agent system according to this embodiment. First, the premise of the dialogue agent system according to this embodiment will be described. Due to the limited computational resources required to run large-scale language models in current AI, in many cases, an information terminal 1 (client) such as a smartphone is used as a UI terminal. A user inputs questions or requests in text or voice into an app or web browser on the information terminal 1. The app or web browser on the information terminal 1 sends the input data to the cloud 3 (such as an address indicated by a URL (API endpoint) specified to use the AI model) via an API or HTTP / HTTPS protocol.
[0036] The API (Application Programming Interface) and HTTP / HTTPS protocol are rules for communication between the app or web browser on information terminal 1 and the web app on cloud 3, and are responsible for exchanging data sent and received between them in a specified format (for example, HTTPS request / response). After receiving a request, cloud 3 returns the response to information terminal 1 in a specified format (for example, HTTPS response). The API and HTTP / HTTPS protocol are responsible for authentication to ensure security, and sending and receiving requests and responses. Cloud 3 (server) processes the data and generates a response. Cloud 3 utilizes its vast computing resources and large amounts of data to process the request. Inference processing using an AI model is performed on cloud 3, and image processing, audio processing, data processing, natural language processing, etc. are performed as appropriate depending on the request.
[0037] For example, an information terminal 1 such as a smartphone requests user authentication from the cloud 3 (step S301). If the cloud 3 succeeds in authenticating the user (step S302), it notifies the information terminal 1 of the success of user authentication (step S303). If the authentication fails, the process ends here. Upon being notified of the success of user authentication, the information terminal 1 acquires a user request (step S304) and transmits the acquired request to the cloud 3 (step S305). The cloud 3 executes data processing in response to the request received from the information terminal 1 (step S306) and notifies the information terminal 1 of the result as a response (step S307). The information terminal 1 outputs the notified response (step S308). Thereafter, the information terminal 1 and the cloud 3 repeatedly exchange requests and responses (steps S309 to S313).
[0038] Looking at the overall system configuration of the interactive agent system according to this embodiment, it can also be seen as a single device with the information terminal 1 as the UI unit, the cloud 3 as the calculation unit that generates responses, and a data communication path including the Internet as the data bus. The data exchanged here can use the same API / HTTPS protocol whether the UI unit and the calculation unit are contained in a single device or whether they are composed of two physically different devices, and does not depend on the implementation form.
[0039] In other words, if the UI unit and the processing unit are on different terminals, they communicate with each other via a network using a predetermined API / protocol to carry out processing. Even if the UI unit and the processing unit are on a single terminal, they can communicate with each other via a bus within the terminal using the same predetermined API / protocol to carry out processing, allowing for a high degree of freedom in the combination of embodiments of the interactive agent system.
[0040] For example, the processing described in FIG. 3 as being executed by the information terminal 1 may be executed by an application acting as a client running on the in-vehicle system of the vehicle 2. Alternatively, the processing described as being executed by a client and a server may be executed by a UI unit and a computing unit, respectively, within a single information terminal 1. As long as the same predetermined API and protocol are used, the function of the UI unit (or client) in the interactive agent system may be implemented or executed in the information terminal 1 or the vehicle 2, and the function of the computing unit (or server) that performs inference processing of the AI model may be implemented or executed in the cloud 3, the information terminal 1, or the vehicle 2. Of course, these may be realized by software running on the information terminal 1, the vehicle 2, or the cloud 3. In other words, all of the disclosures in the present embodiment can be realized with any of these configurations, and may be realized with any system configuration.
[0041] 4 is a sequence diagram showing an example of the flow of processing in which a client downloads and uses agent attribute information in the interactive agent system according to this embodiment. First, information terminal 1 requests user authentication from cloud 3 (step S401). If user authentication is successful (step S402), cloud 3 acquires agent attribute information of the agent designated (or default) by the user from memory 302 (step S403), and transmits the acquired agent attribute information to information terminal 1 (step S404).
[0042] Upon receiving the agent attribute information, the information terminal 1 sets the agent attribute information (step S405) and activates the agent based on the set agent attribute information (step S406). This allows the user to recognize the specified agent and understand that communication, such as talking to the user, can begin based on the agent's status displayed on the UI unit 103. Next, the activated agent acquires a request based on the user's speech (step S407) and transmits the acquired request to the cloud 3 based on a predetermined protocol / API (step S408). The cloud 3 executes data processing in response to the request received from the information terminal 1 (step S409) and notifies the information terminal 1 of the result as a response based on the predetermined protocol / API (step S410). The cloud 3 updates the user attribute information based on the interaction with the user (step S411) and updates the agent usage log (a database recording the date and time of interactions between the user and the agent, the content of the conversation, etc.) (step S412). Furthermore, the information terminal 1 responds to the user via the agent displayed on the UI unit 13 in accordance with the received response (step S413). Thereafter, the information terminal 1 and the cloud 3 repeatedly exchange requests and responses.
[0043] User attribute information is a database that includes one or more of the following: the user's name (nickname), age, gender, career history, interests, preferences, past conversation history, thoughts, experiences, schedule information, frequently used or subscribed external information / services, incomplete or unresolved tasks, unique information about the device used by the user (identification information, access information), biometric information, medical history, or behavioral history (movement history).
[0044] By storing agent attribute information on the cloud 3 (server side), the agent attribute information can be displayed on any information terminal 1 (client) when accessed from that information terminal 1. The agent attribute information is a data set used to represent an agent, including, for example, one or more of the following: 3D model data (data defining the agent's appearance and 3D physique, including texture, clothing, etc.), animation data (data on facial expressions, mouth, and gestures for reproducing natural movements), voice data (data for reproducing the characteristics of an agent's vocalizations), emotion data (data for reproducing specific behavioral patterns based on emotions), or control script (control code for ensuring consistency of the agent's actions and overall behavior when responding).
[0045] In cloud 3, a usage log (user attribute information) recording interactions between the agent and the user is stored and managed while being updated continuously in memory 302. This makes it possible to search for past events with the user and provide answers or suggestions. For example, if a user asks the agent about the final status of a specific matter, the agent can refer to the conversation history in the user attribute information to find the final status of that matter and generate an answer. This type of technology is known as RAG (Retrieval-Augmented Generation), a natural language processing technology that combines information retrieval and generative modeling.
[0046] The cloud 3 may update user attribute information (interests, preferences, thoughts, etc.) obtained through interactions between the agent and the user in memory 302. In addition, by recording, managing, and updating information about the user's knowledge system and experience as user attribute information, it becomes possible to provide replies and suggestions based on the user's knowledge system (which allows interactions based on the user's knowledge in areas in which the user is knowledgeable) and experience (recorded data such as what the user has experienced in the past and events that have occurred).
[0047] The contents of the request and response may be text data such as text chat, or may be data such as images, videos, and audio. If the information terminal 1 (client) on which the agent that interacts with the user operates is equipped with various sensors, modal data other than the above-mentioned text, images, and audio may be included.
[0048] 5 is a diagram showing an example of the functional configuration of the interactive agent system according to this embodiment. In this embodiment, an acquisition unit 206a, a setting unit 206b, an execution unit 206c, etc. are realized by a calculation unit 206 included in an in-vehicle system (an example of a first computer) of a vehicle 2 executing a program stored in a memory 207. In this embodiment, the acquisition unit 206a, the setting unit 206b, the execution unit 206c, etc. are realized in the vehicle 2, but they may also be realized in cooperation with a user's information terminal 1 or the cloud 3.
[0049] The acquisition unit 206a (first computer or an authentication computer capable of communicating with the first computer) acquires authentication information from the user's information terminal 1 (an example of a communication terminal). Here, the authentication computer is a computer capable of communicating with the vehicle 2's on-board system, and may be, for example, a cloud 3 equipped with AI or a server of a vehicle dispatch service provider or the like. Here, the authentication information may also include information on an electronic key for the user to use the vehicle 2.
[0050] Furthermore, after determining that the authentication information is valid, the acquisition unit 206a acquires access information for accessing the agent used by the user from the information terminal 1. That is, once authentication between the information terminal 1 and the vehicle 2 is obtained, the acquisition unit 206a receives access information for the agent (for example, connection destination information such as an API end point) from the information terminal 1 and connects. The acquisition of access information may be executed when it is determined, based on the authentication information acquired from the information terminal 1, that information on an agent available to the user is not registered in the memory 207 of the in-vehicle system.
[0051] Furthermore, after determining that the authentication information is valid, the acquisition unit 206a may establish a connection to the information terminal 1 (or the cloud 3) on which the multimodal agent is running. In this case, the acquisition unit 206a acquires sensing data acquired via one or more sensor units 101 provided in the vehicle 2 while the user is using the vehicle 2. Here, the sensing data may be data including at least one of the user's facial expressions, gestures, emotions, and biometric information, the interior environment of the vehicle 2, the surrounding conditions around the vehicle 2, and the driving state of the vehicle 2. Here, the output data may include the type of data from which the sensing data was extracted, an identification code for identifying the state or event of the vehicle 2 or the user, and a timestamp indicating the time when the user's state or the event occurred. Here, the multimodal agent is not limited to a partner-type agent and may be, for example, a cloud agent or a vehicle agent.
[0052] Then, the acquiring unit 206a converts the acquired sensing data into output data of a format and type that the agent can handle, for each unit of data acquired during a predetermined period. After determining that the authentication information is valid, the acquiring unit 206a may acquire data format information indicating a data format and data type that the agent can handle from the information terminal 1. In this case, the acquiring unit 206a may convert the sensing data into output data in accordance with the acquired data format information.
[0053] Here, if the response generation part of the agent (such as the inference processing of the AI model) is implemented in the information terminal 1, the access information may include an address for accessing the agent in the information terminal 1. Also, here, if the agent is implemented in a server that can communicate with the in-vehicle system of the vehicle 2 via a wide area communication network, the access information may include an address (such as an API end point) for accessing the agent in the server via the wide area communication network.
[0054] Here, the address included in the access information may be a global address for accessing an agent implemented in a server that can communicate with the in-vehicle system via a wide-area communication network, or a local address for accessing an agent implemented in an information terminal 1 that can communicate with the in-vehicle system without going through a wide-area communication network. Connection to the information terminal 1 may be performed via a common API, regardless of whether the address included in the access information is a global address or a local address. This allows the in-vehicle system to connect to both the agent on the server side and the agent on the edge side (information terminal 1, such as a smartphone) using the common API.
[0055] Furthermore, based on the access information, the acquisition unit 206a acquires agent attribute information (an example of agent attribute information) for representing the agent's character from a cloud 3 that can communicate with the calculation unit 206 (an agent database realized by the cloud 3 where the AI inference process is executed, an information terminal 1 such as a smartphone, or another server (an example of a third computer)). That is, the acquisition unit 206a acquires agent attribute information including information on the agent's UI from the information terminal 1 or the cloud 3. Here, the agent database may be located in the cloud 3 or the information terminal 1, or a third computer such as another server with which the in-vehicle system of the vehicle 2 can communicate via a network such as a wide area communication network.
[0056] The setting unit 206b sets at least one of a GUI and a VUI representing the character of the agent in the calculation unit 206 based on the agent attribute information.
[0057] Based on the access information, the execution unit 206c connects to the cloud 3 on which the agent is implemented or the information terminal 1 (an example of a second computer or a communication terminal), and executes an interaction process with the user using at least one of the GUI and the VUI and a response generation process by the agent. Here, the GUI may be displayed on a display provided in the UI unit 204. The VUI may be input and output via a speaker and a microphone provided in the UI unit 204. In addition, in the interaction process with the user inside the vehicle 2, the execution unit 206c may refer to a usage log (or user attribute information) including the content of the conversation between the agent and the user before the user used the vehicle 2. This allows the user to have a conversation with the agent after getting into the vehicle 2 using the conversation log from before getting into the vehicle 2.
[0058] Furthermore, the agent that is automatically activated when the user gets in or out of the vehicle is selected from one or more agents available to the user based on the usage log stored in the information terminal 1. In this case, the agent may be the agent that the user last used before getting in or out of the vehicle. In other words, even if the device the user is using changes, the agent with which the user most recently interacted may be called up on the device to be used next. Alternatively, if the authentication information at the time of getting in the vehicle includes information on an electronic key for the user to use the vehicle 2, the agent may be an agent selected in the information terminal 1 from one or more agents associated with the electronic key. In other words, an agent linked to the electronic key used to unlock or start the vehicle 2 in which the user is riding may be automatically called up on the in-vehicle system of the vehicle 2 when the user gets in the vehicle.
[0059] Alternatively, an agent may be selected from one or more agents by the user selecting an electronic key from a plurality of electronic keys available to the user via the UI unit 103 of the information terminal 1. Here, the plurality of electronic keys may include a corporate electronic key associated with the company to which the user belongs and the user's personal electronic key. If the selected electronic key is a corporate electronic key, a corporate agent linked to the corporate electronic key may be selected as the agent. On the other hand, if the selected electronic key is a personal electronic key, a private agent linked to the personal electronic key may be selected as the agent.
[0060] Furthermore, if a multimodal agent is implemented in the information terminal 1, the execution unit 206c may periodically or intermittently transmit output data converted from the sensing data by the acquisition unit 206a to the agent in the information terminal 1, and cause the agent to execute dialogue processing reflecting the latest sensing data regarding the user and / or the vehicle 2. In this case, if the agent determines, based on the output data, that the user is in a situation where they should concentrate on driving the vehicle 2, it may change the timing or amount of information presented to the user. Furthermore, if the agent determines, based on the output data, that the user has indicated their intention to respond to the agent, but has not received a response from the user via the agent's VUI, it may execute dialogue processing in accordance with the inference based on the output data. This allows the camera on the vehicle 2 to determine that the user is reacting in some way, but allows the user to confirm if there is no voice input or the voice input is unclear.
[0061] Furthermore, when the calculation unit 206 is equipped with a vehicle agent different from the agent, the execution unit 206c may cause the vehicle agent to generate request data indicating constraints on the dialogue processing executed by the agent based on the output data. In this case, the execution unit 206c may send the request data in addition to the output data to the agent in the information terminal 1 and cause the agent to execute the dialogue processing based on the output data and the request data. Here, the request data is a request made by the vehicle agent to the agent, such as "Please refrain from unnecessary conversation for now," and the dialogue processing of the agent based on this request data may be, for example, an affirmative response or may be immediately reflected in the dialogue processing with the user.
[0062] In this embodiment, the calculation unit 104 of the information terminal 1 executes a program stored in the memory 105 to realize the communication unit 104b, the execution unit 104a, etc. In this embodiment, the communication unit 104b, the execution unit 104a, etc. are realized in the information terminal 1, but they may also be realized in cooperation with the vehicle 2 or the cloud 3.
[0063] The communication unit 104b transmits authentication information for the user to use the vehicle 2 to the in-vehicle system of the vehicle 2. After determining that the authentication information is valid, the communication unit 104b transmits access information for accessing the agent to the in-vehicle system of the vehicle 2. Furthermore, the communication unit 104b transmits agent attribute information for representing the agent's character to the in-vehicle system of the vehicle 2, and causes the in-vehicle system to set at least one of a GUI and a VUI based on the agent attribute information. The communication unit 104b may transmit access information for accessing the agent selected by the execution unit 104a (described later) to the in-vehicle system of the vehicle 2 simultaneously with the authentication information or after determining that the authentication information is valid.
[0064] The execution unit 104a executes an interactive process with the user using at least one of the GUI and the VUI and a response generation process by the agent while connected to the in-vehicle system of the vehicle 2. In the interactive process with the user inside the vehicle 2, the execution unit 104a may also refer to a usage log including the content of the conversation between the agent and the user before the user uses the vehicle 2. Furthermore, the execution unit 104a may select, as the agent, one of one or more agents that the user last used based on the usage log, either simultaneously with the authentication information or after it is determined that the authentication information is valid.
[0065] The execution unit 104a receives an input operation from the user via the UI unit 103 regarding which of multiple electronic keys to select. The execution unit 104a may then select an agent associated with the electronic key selected by the user. Specifically, if the selected electronic key is a company electronic key, the execution unit 104a may select a company agent managed by the company as the agent. If the selected electronic key is a personal electronic key, the execution unit 104a may select the user's private agent as the agent. Note that although the description here assumes that the user selects an electronic key to use, the present disclosure is not limited to this. An application or agent on the information terminal may automatically select an electronic key and / or an agent suited to the vehicle's intended use by referring to past conversations with the user or a schedule.
[0066] 6 is a sequence diagram showing an example of the process flow for continuing interaction with an agent even when switching clients in the interactive agent system according to this embodiment (summoning an agent linked to an electronic key to a new client). In this example, the process shows a flow in which a user gets into a vehicle while having a conversation with an agent executed on information terminal 1, and the conversation with the agent is automatically continued with the in-vehicle system as the UI unit.
[0067] First, an agent (e.g., a partner-type agent) implemented by a client application on the information terminal 1 receives a request from the user via the UI unit 103 (step S601) and transmits the received request to a server application (or the agent's API endpoint) on the information terminal 1 that serves as the agent's brain (step S602). The server application executes data processing in response to the received request (step S603), transmits a response to the client application (step S605), and updates the usage log (step S604). The client application on the information terminal 1 responds to the user using the agent on the client application that serves as the body in accordance with the received response (step S606). In this step, the client application transmits a weather-related user question, "What's the weather in Osaka?" to the server application via a predetermined API / protocol (step S602), and the server application returns a reply, "It's sunny, then it's going to rain," to the client application via a predetermined API / protocol (step S605). All of this processing is performed by the execution unit 104a of the information terminal 1.
[0068] When the user approaches vehicle 2 or performs an unlocking operation, vehicle 2 requests electronic key authentication via short-range wireless communication to information terminal 1 (step S607). The electronic key app of information terminal 1 authenticates the key based on the electronic key selected by the user (step S608), calculates a response value indicating the authentication result, and transmits the response value to vehicle 2 (step S609). The on-board system of vehicle 2 verifies the response value, and if the electronic key authentication is successful (step S610), unlocks the vehicle and / or starts vehicle 2 (step S611).
[0069] The acquisition unit 206a of the vehicle 2 acquires, from the memory 207, access information (e.g., API endpoint) of the agent linked to the electronic key selected by the user (step S612). This may be acquired by the acquisition unit 206a by starting a client application installed in the vehicle 2. If the memory 207 does not contain access information linked to the electronic key or if access is not possible, the acquisition unit 206a acquires, from the information terminal 1, the access information of the agent linked to the electronic key selected by the user (steps S613 and S614).
[0070] Next, the acquisition unit 206a uses the acquired access information to request an access, agent attribute information, a usage log, etc. from the agent for which inference processing of an AI model is being executed on the information terminal 1 (step S615). Upon receiving the access request from the agent, the server application of the information terminal 1 transmits the agent attribute information and the most recent usage log of the agent to the vehicle 2 (step S616).
[0071] The setting unit 206b of the vehicle 2 sets the attributes of the agent that will communicate with the user based on the received agent attribute information, and the execution unit 206c starts the agent (step S617). Furthermore, the execution unit 206c displays the most recent interaction between the user and the agent based on the received usage log (step S618). Thereafter, when the agent receives a request from the user (step S619), the execution unit 206c transmits the received request to the server application of the information terminal 1 using a predetermined API / protocol (step S620).
[0072] The server application of the information terminal 1 executes data processing including inference processing of the AI model in response to the received request (step S621), and transmits a response to the vehicle 2 (step S622). Furthermore, the server application updates the usage log stored in the memory 105 (step S623). The execution unit 206c of the vehicle 2 responds to the user while controlling the expression / representation of the agent based on the agent attribute information in response to the response received from the information terminal 1 (step S624).
[0073] For example, the agent is not limited to an electronic key, but may be managed and automatically activated by linking it to one or more of the information terminal 1's identification ID, IMEI (International Mobile Equipment Identity), SIM (Subscriber Identity Module), the user's face, iris, retina, fingerprint, palm print, vein pattern, voice, PIN (Personal Identification Number), activation phrase (wake-up word), gesture, or the agent's nickname.As in the example shown in the figure, the server (server application) that generates the agent's response may send data from the most recent interaction to the currently used client based on the user account and synchronize them, so that communication with the agent continues even when the client changes from information terminal 1 to the in-vehicle system of vehicle 2.
[0074] In other words, by logging in (authenticating) with a user account from a new client, the server can obtain the latest transactions (e.g., usage logs) for that user's account and display / notify them via the UI of the new client. This can be achieved by continuously updating the usage logs on the server side. This section describes a process in which a user in vehicle 2 continuously interacts with an agent represented / depicted by the in-vehicle system, rather than with information terminal 1. The client app sends a follow-up question from the user, "What time will it start raining?" to the server app via a predetermined API / protocol (step S620), and the server app replies to the client app via a predetermined API / protocol, "I'll get off around 3 o'clock" (step S622). In this process, the server app is running on the execution unit 104a of information terminal 1, and the client app is running on the execution unit 206c of vehicle 2. The execution unit 104a of information terminal 1 and the execution unit 206c of vehicle 2 are connected via short-range wireless communication (wired communication is also acceptable) via their respective communication units.
[0075] 7-1 and 7-2 are sequence diagrams showing an example of the process flow for continuing interaction with an agent even when the client is switched in the interactive agent system according to this embodiment (summoning the agent that was used immediately before to a new client). In the following explanation, the same process as that shown in FIG. 6 will not be explained.
[0076] If the memory 207 of the vehicle 2 does not contain access information for the agent that the user used most recently, or if the memory 207 is inaccessible, the acquisition unit 206a requests the access information for the agent that the user used most recently from a server application (the access method is known) that manages the agent use of the information terminal 1 (step S701). The server application of the information terminal 1 may identify the agent that the user used most recently based on the last use log information for each agent (step S702), and may transmit the access information for the identified agent to the vehicle 2 (step S703).
[0077] Currently, when using an AI model equipped with LLM and capable of chatting, a typical usage scenario involves a user accessing a server on which the AI model is available from a client's web browser and logging in with a user account, where they can view past usage logs and individually configure settings for response policies and the scope of data referenced for chatting with the AI model. Therefore, currently, when continuing to use an agent while switching clients, as in the present disclosure, the user must manually access the agent's server and log in again, which is time-consuming. A challenge with the current usage scenario is that it is not possible to provide a continuous experience where the user continues to use a specific agent regardless of the client. This challenge is likely to become more pronounced when users want to switch clients when getting in / out of vehicle 2 or entering / exiting a home or office building.
[0078] This disclosure provides several solutions to this problem. In addition to the above, the following methods can be considered for identifying the agent last used by a user when switching clients, as shown in Figure 7. [Variation 1] Realizing agent communication functions on the client with an app A user can register one or more agents to be used by the user in an app used by the user on a client, and manage the usage history of each agent using a server (app integration server) used by the app. If this app is pre-installed on the information terminal 1, vehicle 2, and the client used by the user, when the app is launched, the app can log in with the user's account if necessary and obtain the usage history and access information of the last agent used in the app integration server. Furthermore, the latest usage log can be obtained by logging in to the agent. In the present disclosure, it is also assumed that an agent is represented on the client as a communication partner of the user, and therefore the app may not only function seamlessly to communicate with the agent, but also have a function to represent the agent. [Variation 2] Realized by a website that aggregates the agents used by users You could also offer a web service that allows users to use one or more agents through a one-stop web service that can be accessed by users or clients at a specific URL or API endpoint. In this case, users can log in to this one-stop web service to retrieve the agent usage history stored on the server. This allows the web service to identify the agent that the user last used and retrieve the latest usage log. [Variation 3] The server app and the electronic key app share information about the last agent used. If the server application shares the last use date and time information of the agent with the electronic key application (or another application), when the electronic key application performs key authentication with vehicle 2, it can transmit the access information of the agent last used by the user to vehicle 2, and vehicle 2, which receives the information, can record it in memory 207. In this way, it is possible to obtain the access information of the last used agent in step S612. Once the access information is known, it can be accessed to obtain the latest usage log, allowing communication with the agent to continue seamlessly.
[0079] The above variants 1 to 3 are merely examples, and there are other methods for the client to identify the agent last used by a user and obtain its latest usage log. For example, one or more users can log in to all servers of agents used by the client, obtain communication session information and last update date and time information (timestamp), and compare them. The method for identifying the agent last used by a user may be any of the above, an improved method of the above, or another method.
[0080] In the case of the above-mentioned modified example 1, the processing from step S701 to step S703 described in FIG. 7-1 changes as shown in FIG. 7-2. FIG. 7-2 shows the changed parts from FIG. 7-1. The acquisition unit 206a starts an application that expresses / represents an agent installed in the vehicle 2 (step S710). The started application requests access information of the agent last used by the user from the cooperating server (the application cooperation server in FIG. 7-2) (step S711). The server cooperating with the application identifies the agent last used by the user (step S712). Then, it sends the access information of the identified agent to the client application (step S713). Note that this "server cooperating with the application" may be accessed from an application installed on any client, so it is desirable that it be a server that is always accessible via the Internet.
[0081] 8 is a sequence diagram showing an example of a processing flow in which a client enhances interaction with an agent in the interactive agent system according to this embodiment. The acquisition unit 206a of the vehicle 2 periodically (or intermittently, the same applies below) acquires (collects) sensing data (additional data) using the sensor unit 203 (step S801), converts the acquired additional data into output data of a format and type that the agent can handle, and periodically transmits the output data to the information terminal 1 (step S802). Every time the calculation unit 104 of the information terminal 1 receives additional data from the vehicle 2, it updates the additional data stored in the memory 105 based on the received additional data (step S803).
[0082] When the execution unit 206c acquires a request from the user using the agent (step S804), it transmits the acquired request to the information terminal 1 (step S805). When the request is received from the vehicle 2, the calculation unit 104 of the information terminal 1 executes data processing based on the latest added data (step S806) and transmits a response to the vehicle 2 (step S807). Furthermore, the calculation unit 104 of the information terminal 1 updates the usage log stored in the memory 105 (step S808). The execution unit 206c of the vehicle 2 responds to the user using the agent represented by the client in accordance with the response received from the information terminal 1 (step S809).
[0083] When the user makes a disgusted expression, the acquisition unit 206a of the vehicle 2 analyzes the user's facial expression from the video captured by the camera inside the vehicle 2 and transmits the result of the analysis to the information terminal 1 as additional data (step S810). This may be performed periodically (or intermittently) as step S802, or may be performed separately from step S802 each time the agent responds to the user. By performing step S810 each time the agent responds, the vehicle 2 can notify the information terminal 1 of the user's reaction more immediately and at a more appropriate timing, thereby improving the quality of communication and the value of the experience. The calculation unit 104 of the information terminal 1 updates the additional data stored in the memory 105 to the latest state based on the additional data received from the vehicle 2 (step S811). Furthermore, the calculation unit 104 executes data processing related to communication to be conducted with the user based on the latest additional data stored in the memory 105 (step S812). If it determines that further communication should be conducted, it transmits the response generated in step S812 to the vehicle 2 (step S813). If not, the calculation unit 104 waits to receive an explicit request from the user. Thereafter, the calculation unit 104 updates the usage log stored in the memory 105 (step S814). The execution unit 206c of the vehicle 2 responds to the user by an agent in accordance with the response received from the information terminal 1 (step S815). The vehicle 2 and the information terminal 1 similarly exchange requests and responses (steps S816 to S821).
[0084] Here, we will explain an example of an interaction between a user and an agent shown in Figure 8. The user asks the agent, "How far is it to our destination?" (verbal communication) (step S804), and the agent responds, "It will take about an hour and a half" (step S809). Although the user does not respond clearly using words, he / she shows a facial expression (non-verbal communication) that suggests disgust or that he / she perceives the situation as undesirable. This is detected by the vehicle's sensors as the user's reaction to the response in step S809, and this reaction is transmitted to the server in step S810. The server, having detected this reaction, makes a suggestion via the agent, "Would you like to take a short break?" (step S815). The user's silent nod (non-verbal communication) is detected as the user's reaction, and the agent proposes a specific break suggestion, "How about a coffee shop 3 km down the road?"
[0085] In communication between the user and the agent, even if a clear response from the user (for example, text data obtained by voice recognition of oral utterances and including a clear expression of intent, as in verbal communication) is not obtained, at least one of the following non-verbal communication of the user detected by the sensor unit 101 equipped in the interior of the vehicle 2, such as facial expressions, presence or absence of interjections, head movements indicating affirmation or denial, estimated emotions, and biological information (changes in pupils, fluctuations in heart rate, changes in sweating, etc.); or vehicle status data, such as vehicle position, vehicle speed, accelerator and brake operation status, indication of intention to turn right or left (presence or absence of turn signal), indication of intention to make an emergency stop (presence or absence of hazard lamp), and road classification on which the vehicle is located (highway, general road, school zone (road with many children), within an intersection, stop or slow down area, private property (home property, home parking lot), etc.) can be added to the input data of the agent as additional data expressed in a predetermined format. This way, even if the user does not or cannot respond clearly verbally, the system can understand the user's intention (e.g., positive or negative), emotion (e.g., joy, anger, sadness, or happiness), and situation (e.g., concentrating on driving at an intersection), and continue the interaction in a way that is appropriate, less irritating, and makes driving safer.
[0086] 9 is a flowchart showing an example of the flow of a process for querying and transmitting the type of multimodal data input to the interactive agent system according to this embodiment. First, the client side (vehicle 2 or information terminal 1), which mediates the UI with the user, queries the server side (vehicle 2, information terminal 1, or cloud 3) that executes the agent about additional data that the server side can handle, and obtains a list of such data (steps S901 and S902). From the types or formats of additional data that the client side can handle, the client side identifies additional data of a type or format that can be obtained from a connected sensor (for example, the sensor unit 101 of the information terminal 1 or the sensor unit 203 of the vehicle 2) (step S903).
[0087] The client side converts the data acquired from the sensor (sensing data) into output data in a format that can be accepted by the specified server side (step S904). The client side then checks whether there is at least one of the following: additional data due to a status change, additional data due to an update period, or additional data for which an update is requested (step S905). If there is at least one of the following additional data due to a status change, additional data due to an update period, or additional data for which an update is requested (step S905: Yes), the client side transmits the corresponding additional data to the server side (step S906, which corresponds to steps S802, S810, and S816 in FIG. 8). By repeating this process, the agent (data processing for response generation, including inference processing by an AI model) running on the server side can accurately detect or recognize the user's status and changes in the surrounding environment when the status changes, periodically or intermittently, or at appropriate times depending on the user's status and reaction.
[0088] Here, for example, the additional data may include raw data of the user's state or reaction acquired from an image sensor inside the vehicle (e.g., in-vehicle camera image), or results expressed as lightweight analysis data showing the results of analyzing the raw data. The additional data may include at least one of the following information measuring the user's state or reaction: intonation of speech, pause before utterance, tone of voice, heart rate (heart rate variability), blood pressure, body temperature, breathing, pupils, sweating, facial expression, gaze, posture, and gestures.
[0089] Furthermore, when the user is inside the vehicle while driving, etc., and the on-board system of vehicle 2 is the client, the additional data may include at least one of the following acquired from the sensor unit 203: vehicle position, vehicle speed, accelerator and brake operation status, indication of intention to turn right or left (whether or not blinkers are on), indication of intention to make an emergency stop (whether or not hazard lights are on), road category on which the vehicle is located (highway, general road, school zone (road with many children), intersection, stop or slow down area, private property (home premises, home parking lot), etc.), sensor recognition information of the environment around the vehicle, information about the currently set route, and autonomous driving response level information.
[0090] In this way, depending on the client's detection capabilities, additional data can be sent periodically or in synchronization with requests, making it possible to make the information multimodal. If the server side has a wealth of input information for the agent, it is expected that the agent's responses will be timely and appropriate, taking into account the situation of the user and vehicle. Because the in-vehicle system has the ability to sense various conditions related to the user and vehicle 2, when the agent is run with the in-vehicle system as the client and information terminal 1, vehicle 2, or cloud 3 as the server, it is expected that the agent's responses will be more natural, more comfortable, and safer.
[0091] For example, if vehicle 2 is currently driving, the quality, quantity, and timing of communication from the agent could be controlled to ensure safe driving. Another use case could be to detect or estimate that the user's concentration is declining or that they are yawning after a long drive, and then suggest that the agent stop at a rest stop along the route. Furthermore, if there are changes in the user's biometric information, the agent could ask if they are feeling unwell, guide vehicle 2 to a suitable parking spot, or, if vehicle 2 is capable of autonomous driving, switch to autonomous driving mode and move the vehicle to a suitable parking spot or to the nearest hospital. Furthermore, if the user's status or vehicle status data detects that the vehicle is not currently driving, the agent could generate a response that includes more detailed information or use visual information to aid understanding. Such capabilities would be difficult or insufficient if only text recognized from linguistic utterances were sent to the agent; this can only be achieved through multimodal input, as described above. Furthermore, as described above, by arbitrating in advance which multimodal data will be shared and sharing it instantly when needed in a lightweight data format, the processing load on the inter-systems can be reduced, leading to lower costs and faster responses.
[0092] 10 is a diagram illustrating an example of the structure of additional data converted from multimodal raw data into text data in the interactive agent system according to this embodiment. Raw data obtained by a wide variety of sensors, which represents state quantities that fluctuate over time, is large in size and places a heavy burden on the network and calculations. Therefore, it is possible to periodically transmit the raw data as text data (additional data) representing the state or level of each type of raw data (for example, in JSON format) to the agent (server side). By transmitting the additional data as independent text data (additional data) periodically, the server side can easily update the additional data for each type of modal to the latest one and refer to it when generating a response, even if an arbitrary number of modals (additional data) are updated at any time.
[0093] The text shown as an example of "Vehicle Status 1" is an example of a vehicle status detected by an in-vehicle system (also known as a Cockpit Domain Controller: CDC). Here, it is stated that the in-vehicle system (CDC) detected that the vehicle was in operation, the engine was running, and the vehicle speed was 60 km / h at the measurement time of 11:46:50 on April 25, 2024. As a client, the in-vehicle system periodically or intermittently sends this additional vehicle status data to a server that generates responses through inference processing by an AI model.
[0094] Similarly, the texts listed as "User Status 2" and "User Status 3" are examples of user statuses converted from video images of the user captured by an image sensor inside the vehicle. "User Status 2" indicates that the analysis of the user's emotion (emotion_scores) based on the in-cabin monitor (ICM) on the dashboard at 11:47:00 AM on April 25, 2024, showed that 80% of the user's emotions were negative, indicating a negative reaction. This may be the result of an emotion analysis of the user interacting with the agent, the user sitting in the driver's seat, or all passengers in the vehicle. Information identifying the user being observed may also be included (not shown). Similarly, "User Status 3" indicates the result of an analysis of the user's state based on gestures (gesture scores), indicating that the user is responding positively.
[0095] Furthermore, the periodic transmission of this additional data may be realized by including the additional data in an HTTPS request (e.g., a POST request) between the client and the server, by scheduling the HTTPS request to be sent periodically on the client side using JavaScript (registered trademark) or the like, by establishing a WebSocket connection between the client and the server that enables two-way, real-time communication and transmitting the additional data through that connection, or by an app installed on the client (which may be an app that controls the expression / representation of an agent) notifying the server using a predetermined API. In the present disclosure, the mechanism for periodically transmitting such multimodal additional data from the client to the server is not limited in its implementation form as long as it is realized.
[0096] This diagram also shows, using Figure 8 as an example, the time relationship between the interaction between the user and the agent, and the relationship with the additional data that was periodically updated at that time from the client (vehicle 2) to the server. As shown here, when the user asks the agent, "How far is it to the destination?", vehicle 2 obtains information about the user's status, "User Status 1," information about the vehicle's status, "Vehicle Status 1," and information about the situation around the vehicle, "Vehicle Surroundings Situation 1," using the sensor unit 203, and processes each of these into specified text data and sends it to the server.
[0097] When the agent responds, "It will take about an hour and a half," vehicle 2 acquires information about the user's emotions ("User Status 2") and information about the user's biological reactions ("User Biometric Information 1"), and sends these to the server in the same way. Here, "User Status 2" indicates that the user has reacted negatively through non-verbal communication (emotion estimation through facial expressions or other means), as described above.
[0098] Based on this additional data, the agent detects that the user is tired from driving and suggests taking a break, asking "Would you like to take a short break?" to provide the user with a more comfortable travel experience. Here, vehicle 2 obtains information about the user's state at the time the agent suggested the break, "User State 3," and information about the vehicle's state, "Vehicle State 2," and sends these to the server. Here, "User State 3" indicates that the user has responded positively through non-verbal communication (gestures), as described above.
[0099] If the agent detects a positive response from the user from the additional data, it will present specific break suggestions to users who responded positively to a break suggestion such as "Is there a coffee shop 3km away?" The user's state at this time is also recorded as "User State 4," and the vehicle interior environment is recorded as "Vehicle Interior Temperature and Humidity 1," which are then converted into text data and sent to the server.
[0100] Alternatively, the user's facial expressions can be captured by an in-vehicle camera (ICM) and the video data can be sent to the server either directly or after partial trimming and video processing. This approach increases the network and server processing load. To reduce the server's processing load for analyzing video and audio, the time and content of the agent's communication with the user can be embedded as metadata in the video, or the data can be sent to the server via a separate data file or API. This information can also include the agent's appearance, facial expressions, and voice characteristics on the client side. Because the server may not have a complete picture of how the agent communicated with the user on the client side, information about the agent's actual communication style can be used to analyze the user's state and reactions. This approach is expected to enable accurate understanding and analysis of communication with the user.
[0101] FIG. 11 is a diagram illustrating an example of the structure of data converted from multimodal raw data into text data in the interactive agent system according to this embodiment. It is conceivable that an agent (a function that performs data processing including inference processing of an AI model) running on a client that was not involved in the interaction with the user may make a suggestion to the user via an agent (an agent that generates a response to the user on the server). In FIG. 11, a vehicle agent that was not involved in the interaction with the user makes a suggestion to the agent, "Vehicle Agent Proposal 1," based on the previous interaction between the user and the agent.
[0102] "Vehicle Agent Proposal 1" suggests a coffee shop 3 km away as a place to rest. The advantages of this proposal include that the coffee shop is famous for its latte art and that it has a fast charger. It also shows that this proposal was made at 11:47:20 AM on April 25, 2024, by a vehicle agent, an AI agent built into the vehicle system, which can be accessed using access information specified by a URL (e.g., https: / / 192.168.1.100:443 / VehicleAgent). Proposals from agents other than those directly interacting with the user (vehicle agents in Figures 10 and 11) are made when the opportunity and conditions are right, and are therefore sent to the agent as irregular additional data.
[0103] When the agent receives this additional data "Vehicle Agent Proposal 1" from the client vehicle 2, it decides whether to make the proposal to the user. If the proposal is not made, the vehicle agent's proposal is rejected and is not used in interactions with the user. If a proposal is made, as shown in FIG. 11, the agent proposes a specific break idea to the user, such as "Is there a coffee shop 3 km away?". Here, to simplify the interaction between the user and the agent, the proposal is treated as if it were from the agent, but the present disclosure is not limited to this. The agent may also reveal the name of the agent that made the proposal and convey the contents of the proposal to the user, noting that it is from that agent.
[0104] One of the key points of this disclosed technology is that when there are multiple agents available to a user (functions that perform data processing, including inference processing for AI models), it is possible to smoothly and lightly implement suggestions and information from an agent different from the agent that is communicating with the user, without confusing the user. When various specialized agents appear, input unique to the specialized agent can be sent to the partner agent in the form of a proposed text via a partner agent that communicates directly with the user and has a deep understanding of the specialized agent, and if it is determined to be useful, the partner agent can convey the information to the user in a form and granularity that is easy to understand.
[0105] In the above description, the proposal from the vehicle agent, "Vehicle Agent Proposal 1," is sent to the server-side agent for processing, but the present disclosure is not limited to this. The client app may directly and self-containedly handle the proposal from the vehicle agent and make the proposal from the vehicle agent using an agent expressed / represented on the client, or the vehicle agent may appear separately as a separate communication partner and make the proposal to the user from that vehicle agent, or a message (nudge) indicating that there is a coffee shop suitable for a break 3 km away may be notified to the user via a screen on the UI unit, without mentioning the source of the proposal.
[0106] Fig. 12 is a diagram for explaining an example of a dialogue agent system according to this embodiment. Specifically, Fig. 12 is a diagram showing how the agent interacting with the user in Fig. 11 (a partner-type agent dedicated to the user) is connected to other devices or other agents.
[0107] The user interacts with the agent via a client terminal (or an app running there; the same applies below) that handles UI operations. In the example shown in Figure 11, the client terminal is an in-vehicle system. As explained in Figure 11, proposals from the vehicle agent, which is an agent built into the in-vehicle system, to the agent are sent from the client terminal's built-in agent (vehicle agent) to the client terminal (its app), and are communicated from the client terminal to the partner agent (the agent running on the information terminal 1 or cloud 3 server that interacts with the user) via a server-client API. Management by a communication session may be effective between the server and client, and the connection may be via HTTPS protocol or WebSocket connection. Furthermore, if the client terminal's built-in agent knows the access information for the partner agent, it may directly connect to the partner agent and provide information.
[0108] On the other hand, agents that interact with users also connect to the Internet and act as a bridge, using external knowledge and services to notify or provide them to the user. This may be a wide variety of publicly available information on the Internet that can be accessed via an Internet connection, or a service provider site that offers specific services such as an e-commerce site or search site. Furthermore, agents that are highly capable of performing specific tasks and can receive those services via APIs, etc., are envisioned; these are referred to as specialized agents in Figure 12.
[0109] Specialized agents are agents that specialize in carrying out knowledge or tasks in a specific field, such as a free agent that suggests the most CO2-efficient route, or a paid agent provided by an entertainment company that provides conversation companionship during long-distance drives. Agents that interact with users can make their interactions with users richer and more valuable by utilizing the resources of information and services on the Internet.
[0110] FIG. 13 is a sequence diagram for explaining an example of a dialogue agent system according to this embodiment. In this diagram, a user is connected via a client application (body) that represents the agent's appearance and voice to a server application (brain) that generates responses from a partner-type agent that communicates directly with the user. The partner-type agent obtains answers from a trusted specialized agent in a specific field, customizes the answers for the user, and responds. The processing will be explained below with reference to the diagram.
[0111] The client application running on the in-vehicle system of vehicle 2 communicates with the user to obtain a user request (step S1301) and sends it to a server that executes data processing, including inference processing for the AI model (step S1302). In this figure, the server is information terminal 1, and in reality, the request is sent to a partner-type agent running on information terminal 1. Here, the client application obtains a question from the user, "How far is it to the destination?", and requests a response from the partner-type agent.
[0112] Upon receiving this request, the partner-type agent searches for candidate specialized-type agents that are likely to be able to generate answers that are more accurate or more useful to the user than the partner-type agent, and if one or more specialized-type agents are found, it decides to use the answer from the specialized agent (step S1303).
[0113] The partner agent notifies the user of this (step S1304). In this example, the user is notified, "I'll ask the navigation agent." Upon receiving this, the client application uses a GUI or VUI to express the agent's response and responds to the user (step S1305).
[0114] The partner-type agent sends an access request to a specialized-type agent with which it can communicate via the network (step S1306). The specialized-type agent, upon receiving this, grants access (step S1307). At this time, the conditions of use may also be determined. The specialized-type agent responds by notifying the access permission and, if any, indicating the conditions of use of the specialized-type agent (step S1308).
[0115] When the partner-type agent receives this, if the terms of use are included, it checks the terms of use and accepts them if there are no problems (step S1309). If it does not accept them, it requests access from another specialized-type agent and performs the above processing. If necessary, the partner-type agent provides the specialized-type agent with information in accordance with the terms of use, or provides monetary or economically valuable points or cryptocurrency (step S1310).
[0116] The partner-type agent requests the specialized-type agent to estimate the time required to reach the destination, along with the current vehicle position, the currently set route information, etc. (Step S1311). This request is an instruction to the specialized-type agent that is generated autonomously by the partner-type agent, and may take the form of a prompt describing the information processing content expected of the specialized-type agent.
[0117] The specialized agent that receives this estimate estimates the required time to the destination based on the latest traffic information database (step S1312) and sends the result to the partner agent (step S1313). The partner agent that receives this estimate processes the estimated required time obtained from the specialized agent for the user based on at least one of the above-mentioned user attribute information, user status, and vehicle status, such as summarizing or supplementing it (step S1314), and generates a response to the user and sends it to the client application (step S1315). In this case, the response to the user is "It's congested, so it will take about an hour and a half," as a result of adapting it to the situation and the user's preferences. Alternatively, the partner agent may send the specialized agent's estimate directly to the client application without processing the response in step S1314.
[0118] The client application that receives this response answers the user's question based on the received response (step S1316). At the same time, the client application acquires the user's state and reaction at the time the response is notified via one or more sensors connected to the in-vehicle system (step S1317). This may be part of periodic or intermittent detection of the user's state.
[0119] The client application either receives the sensing data detected via the sensors regarding the user's state and reaction as raw data without processing, or analyzes and converts it into the aforementioned JSON-formatted text data classified into categories and levels, and sends it to the partner-type agent as the aforementioned additional data (step S1318). Here, the user's reaction may be clearly expressed in words (for example, "OK," "Got it," "Huh?", "Will it take that long?"), or it may be the emotion analysis results obtained from the user's facial expressions and posture changes through the aforementioned ICM video analysis (for example, information on the strongest detected emotion, information on the degree to which each emotion is detected, etc.).
[0120] Upon receiving the user's reaction indicating dislike, the partner-type agent creates and requests the specialized-type agent to propose a new route based on at least one additional piece of data on the user, vehicle condition, and external environment (in this case, traffic conditions along the route, weather, event information near the route, etc.), and / or user attribute information including information on points that may be of interest to the user and ways to rest (steps S1319, S1320).
[0121] The specialized agent then generates a response (new route proposal) based on the received instructions (step S1321). At this time, the specialized agent receives advertising and promotional expenses from multiple commercial facilities and services. It may adjust the appeal and priority of the facilities and services included in the new proposal based on the monetary value of the remaining advertising and promotional expenses. Alternatively, the specialized agent may enter into a contract to charge a fee for inclusion in the proposal, and the specialized agent may charge the facilities and service providers included in the proposal it generated for their advertising (step S1322). This approach allows for market-based advertising and promotion, and allows users to use the specialized agent free of charge or at a very low cost. Furthermore, instead of advertising within a limited, narrow area (e.g., prefecture-level), the specialized agent can match facilities and services available along the currently set route with pinpoint accuracy, thereby introducing facilities and services that are highly accessible to the user. This is likely to be beneficial for both the user and the service provider.
[0122] The specialized agent that generated the new route proposal sends the proposal to the partner agent (step S1323). Upon receiving the proposal, the partner agent, as described above, processes the proposal for the user by adding at least one of the additional data related to the user's attribute information, user status, and vehicle status (step S1324) and sends it to the client application (step S1325). In this case, based on the fact that the vehicle's remaining battery power is low and that the user often stops at coffee shops, the specialized agent responds with a suggestion such as, "You'll also need to charge your car. How about taking a break and charging at a coffee shop 3 km away?" (step S1326). The client application that receives this response notifies the user via the agent's appearance (GUI) or voice (VUI). At the same time, as described above, the partner agent detects the status and reaction of the user who has received the response, thereby realizing multimodal communication (step S1327).
[0123] Fig. 14 is a diagram for explaining an example of a dialogue agent system according to this embodiment. In Fig. 10 and Fig. 11, communication progressed through interactions between the user and the agent, but Fig. 14 explains an expression form in which the agent that was interacting with the user calls the vehicle agent in Fig. 11, causes it to participate in the interaction, and then causes it to withdraw from the interaction once the task is completed.
[0124] If, during an interaction with a user, the agent determines that there is a candidate agent who can provide a better response than the agent itself, it will call that agent who is able to communicate via the network. In this case, the agent that was conversing with the user determines that it would be most appropriate to ask the vehicle agent about the time and distance to the destination, and calls the vehicle agent into the conversation. When calling the vehicle agent, the agent asks the vehicle agent a question about the time required to reach the destination, and the vehicle agent responds to the question based on the current situation and provides a response to the user and the agent.
[0125] Apart from the original communication between the user and the agent, a separate server-client type communication may be established between the agent and the vehicle agent. For example, communication may be performed using the HTTPS protocol or the vehicle agent's API. Even in this case, based on the response obtained from the vehicle agent, the client (application) may use the UI section to present the response to the user as if the vehicle agent were saying "It's congested, so it will take about an hour and a half."
[0126] The agent can individually connect with the vehicle agent (another agent other than itself) to obtain information necessary for interaction with the user, and then use or quote the answer in the interaction with the user while claiming it is the answer from the vehicle agent. This is thought to increase the possibility of the agent interacting appropriately and to expand the possibilities for story development.
[0127] Once the necessary information has been obtained from the vehicle agent and confirmed with the user, the agent disconnects the vehicle agent from the interaction with the user. In Figure 13, the communication is disconnected as if the agent said "Thank you, see you later" to the vehicle agent. From then on, the conversation returns to the traditional server-client type connection with the user's client and continues.
[0128] The interaction between this agent and the vehicle agent can be realized as an individual connection between the partner-type agent and the agent built into the client terminal in Fig. 12. If it is determined that it would be better to ask the specialized agent rather than the vehicle agent, an individual connection may be established between the partner-type agent in Fig. 12 and the specialized agent on the Internet to obtain an appropriate answer or recommendation and share it with the user. In this case, it is possible to have the user experience a three-way interaction including the specialized agent by operating the UI unit 204 of the client, or to have the user experience a two-way interaction between the partner-type agent and the user by operating the UI unit 204, taking into account information obtained from the specialized agent, without directly showing the interaction between the specialized agent and the partner-type agent to the user.
[0129] As shown in this diagram, when an agent calls a specialized agent, it is desirable to have the specialized agent appear in the conversation with the user by asking the specialized agent who can respond more accurately to the current topic, and then have the specialized agent exit the conversation after expressing gratitude once the requirement is met. This makes it easy for the user to start and end communication using a specialized agent.
[0130] Furthermore, the specialized agent does not necessarily have to be displayed in the same form as the agent on the client's UI unit 204; the specialized agent's name and appearance may be displayed in a predetermined frame, like a participant participating in an online conference. In this case, the agent can act as a moderator or be given a certain degree of progress management authority to promote communication with the user, making it easier for the user to understand the progress of the conversation. When inviting / exiting a specialized agent, the agent can do so in an easy-to-understand manner, allowing the user to utilize the specialized agent's knowledge and skills without any additional effort.
[0131] FIG. 15 is a diagram illustrating an example of multimodal data in the interactive agent system according to this embodiment. Specifically, it shows the types of data related to the state of the user and the vehicle 2 sensed on the client side. By inputting these data as multimodal data to an agent, which is an AI model that generates a response, it is expected that the agent's response will be appropriate, accurate, and quick in accordance with the user's non-verbal communication (facial expressions, gestures, etc.), the state of the vehicle 2, and changes in the user's surrounding environment, even if an explicit verbal response (such as speech or a text version of speech) is not obtained from the user.
[0132] A wide variety of time-series data can be obtained from various sensors, but sending this data to the server where the agent is running can put a strain on the communication bandwidth, increase the load on the server's data processing, cause delays, and increase costs.For this reason, it is possible to treat the data from each sensor as separate, independent additional data, such as small amounts of text data written in a specified format (for example, JSON format data) that indicates the status and level on the sensing client side, and update it periodically.
[0133] The client manages this multimodal additional data along with its type and the date and time of its measurement, and generates responses based on the latest information. It may stop referencing old additional data that has passed a certain amount of time, or additional data that is no longer updated because the client has lost connection, or delete that data type from the latest information dataset. This allows the agent to generate responses based on additional data that is updated at an appropriate interval (i.e., the latest information and situation about the user and the user's surrounding environment).
[0134] The following are the types of sensors and systems used for sensing on the client side and the types of data that can be obtained from them. This data can be used to read a wide variety of information about users, vehicles, and the situation around the vehicle.
[0135] Sensors worn by users include smart watches, smart rings, smart glasses, etc. These are accessories that are embedded with sensors that can measure the user's biological responses, such as heart rate (heart rate, heart rate variability), electrocardiogram, blood pressure, number of steps, calories burned, calories ingested, blood glucose level, sleep, stress, electrodermal activity, breathing, voice analysis, electroencephalogram, electromyogram, skin gas, blood oxygen saturation, chewing, swallowing, gaze, and blinking.
[0136] The information devices carried by users also have many sensors that can measure information about users, such as the information they search and view online, the information they send or receive via email, chat, or social networking sites, their current location, or its history.
[0137] In-vehicle sensors include the ICM, seats, and seatbelts, which can measure the user's facial expression, posture, responses, gaze, brainwaves, surface temperature, seatbelt status, passenger identification and identification, and object identification and identification.
[0138] Vehicle control sensors include data sent from each ECU in the vehicle to the network (CAN) and data used by ADAS (Advanced Driver Assistance Systems).For example, they can measure vehicle status such as the state of driving, parking, or stopping, whether the engine or motor power source is on or off, the remaining gasoline level, the remaining battery level, the state and level of application of autonomous driving or driver assistance systems, vehicle speed, operation of the steering wheel, accelerator, and brake, road classification (highway, general road, inside or outside an intersection, inside or outside a slow-down area, on private property (within a home), etc.), driving conditions such as whether or not turn signals are on and whether or not hazard lights are on, route information, and navigation system setting information such as traffic congestion information.
[0139] Sensors around the vehicle include LiDAR and radar, which can measure surrounding conditions such as sensing data of the environment around the vehicle, its position relative to roads, intersections, and traffic infrastructure, the types of objects around the vehicle and their positions, and lane information about the vehicle.
[0140] The client processes at least one of these pieces of data into a predetermined format periodically, intermittently, or at a predetermined timing, and sends it to the server, which performs inference processing for the AI model, thereby enabling the server to generate more accurate responses when generating communications.
[0141] In this way, according to the interactive agent system of this embodiment, when the user gets into the vehicle 2, it is possible to display and continue to use an agent that is linked to the user or the electronic key being used, or an agent that the user was using on the information terminal 1 immediately before the user got into the vehicle 2. Furthermore, by using a mechanism similar to that used when getting into the vehicle 2, it is possible to display and continue to use the agent that the user was using on the vehicle 2 immediately before the user got out of the vehicle 2 on the information terminal 1.
[0142] In this embodiment, an example has been described in which an agent capable of dialogue linked to a user is realized in a vehicle 2, but the present invention is not limited to this. Even in spaces such as a home, office, or store, it is possible to obtain an ID that identifies an information terminal 1 such as a smartphone, obtain agent access information from an application installed on the information terminal 1, or obtain agent access information linked to the user's entry authentication information for their home, office, store, etc., and thereby summon a partner-type agent that the user has set or the agent that the user was using immediately before onto the computer system of that space when the user enters that space.
[0143] For example, just as a conversation with a partner-type agent can be seamlessly switched from information terminal 1 to the vehicle's in-vehicle system (client terminal) when getting into vehicle 2 and continued, a conversation with a partner-type agent can be seamlessly switched from information terminal 1 to a client terminal provided in the home or office when entering the home or office and continued.
[0144] The programs executed by the information terminal 1 and the vehicle 2 of this embodiment are provided by being pre-installed in a ROM (Read Only Memory) or the like. The programs executed by the information terminal 1 and the vehicle 2 of this embodiment may be provided by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM, a flexible disk (FD), a CD-R, a DVD (Digital Versatile Disk), or an SD card.
[0145] Furthermore, the programs executed by the information terminal 1 and the vehicle 2 of this embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the programs executed by the information terminal 1 and the vehicle 2 of this embodiment may be provided or distributed via a network such as the Internet.
[0146] In addition, in this embodiment, in the information terminal 1, an example of a processor (calculation unit 104) such as a CPU (Central Processing Unit) uses a RAM (Random Access Memory) or the like as a working area and executes various programs stored in a ROM (memory 105) or RAM, thereby realizing a communication unit 104b and an execution unit 104a.
[0147] In addition, in this embodiment, in the in-vehicle system, an example of a processor (calculation unit 206) such as a CPU (Central Processing Unit) uses a RAM (Random Access Memory) or the like as a working area and executes various programs stored in a ROM (memory 207) or the RAM or the like, thereby realizing an acquisition unit 206a, a setting unit 206b, and an execution unit 206c.
[0148] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents.
[0149] Fig. 16 is a diagram showing an example of the functional configuration of the interactive agent system according to this embodiment. Here, the processing executed by the aforementioned acquisition unit 206a, setting unit 206b, and execution unit 206c (see Fig. 5) has, in block units, a communication unit 1601, an identification unit 1602, a summoning unit 1603, and an agent providing unit 1604. Other embodiments will be described below in [Item 1] to [Item 7].
[0150] [Item 1] An in-vehicle device that communicates with a cloud and a user terminal via a network, The vehicle and cloud can summon one or more AI agents to At least one of the one or more AI agents is a partner AI agent that accesses information about the corresponding user; The in-vehicle equipment is A communication unit 1601, a summoning unit 1603 that communicates with the cloud and summons an AI agent; an agent providing unit 1604 for giving a prompt from the user to the summoned AI agent and receiving a response from the AI agent; an identification unit 1602 for identifying a vehicle user who uses the vehicle; The summoning unit 1603 summons a partner-type AI agent corresponding to the identified vehicle user. In-vehicle equipment.
[0151] [Item 2] Support for users of device owners The in-vehicle device according to [Item 1] above, wherein the identification unit 1602 identifies the user of the user terminal with which the communication unit 1601 communicates as a vehicle user.
[0152] [Item 3] Supports users who own smart keys and their devices The user terminal has at least a smart key that allows entry into the vehicle compartment, operation of the vehicle, and / or use of the vehicle equipment; The in-vehicle device according to [Item 2] above, wherein the identification unit 1602 identifies a user of a user terminal that has a smart key as a vehicle user.
[0153] [Item 4] Supports partner-type AI agents on user devices The user device and the cloud summon one or more AI agents, At least one of the one or more AI agents is a partner AI agent that accesses information about the corresponding user; The in-vehicle device according to [Item 2] above, wherein the identification unit 1602 identifies the user corresponding to the partner-type AI agent summoned by the user terminal with which the communication unit 1601 communicates as the vehicle user.
[0154] [Item 5] Supports vehicle wallet owners In-vehicle equipment also includes: It has a wallet unit that performs settlement processing between external equipment and the wallet owner, The identification unit 1602 identifies the wallet owner, The summoning unit 1603 summons a partner-type AI agent corresponding to the wallet owner identified by the identification unit 1602. The in-vehicle device described in [Item 1] above.
[0155] [Item 6] The summoning unit 1603 is an in-vehicle device described in [Item 2] above, which notifies the summoning AI agent of information on whether the vehicle user is in the driver's seat or the passenger seat.
[0156] [Item 7] The summoning unit 1603 notifies the AI agent of the vehicle's location information, and the location information indicates at least one of the home, a parking lot, and a road.
[0157] [Item 8] Methods corresponding to [Item 1] to [Item 7] above. [Explanation of symbols]
[0158] 1. Information terminal 2 vehicles 3. Cloud 101,203 Sensor section 103,204 UI department 104,206,303 Arithmetic unit 104a Executive Department 105,207,302 memory 104b, 106, 208, 301 Communications Department 206a Acquisition Department 206b Setting section 206c Executive Department
Claims
1. 1. An information processing method executed by a first computer mounted on a vehicle in an interactive agent system capable of interacting with a user, comprising: The first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the communication terminal of the user and determines that the authentication information is valid, and then acquires access information for accessing the AI agent used by the user from the communication terminal; acquiring, based on the access information, agent attribute information for representing a character of the AI agent from an agent database that can communicate with the first computer; setting at least one of a GUI and a VUI representing a character of the AI agent in the first computer based on the agent attribute information; Based on the access information, while connecting to a second computer on which the AI agent is implemented, execute an interactive process with the user using at least one of the GUI and the VUI and a process generated by the AI agent. Information processing methods.
2. the GUI is displayed on a display of the first computer; The information processing method according to claim 1 , wherein the VUI is input and output via a speaker and a microphone of the first computer.
3. 2. The information processing method according to claim 1, wherein the agent database is located in the second computer or a third computer with which the first computer can communicate via a network.
4. the first computer is an in-vehicle system; the second computer is the communication terminal, The information processing method according to claim 1 , wherein the access information includes an address for accessing the AI agent in the communication terminal.
5. the first computer is an in-vehicle system; the second computer is a server capable of communicating with the in-vehicle system via a network, The information processing method according to claim 1 , wherein the access information includes an address for accessing the AI agent in the server via the network.
6. The information processing method according to claim 1, wherein a usage log including the content of a conversation between the AI agent and the user before the user uses the vehicle is referenced in the interactive processing with the user inside the vehicle.
7. the first computer is an in-vehicle system; the access information includes an address for accessing the AI agent; The address is a global address for accessing the AI agent implemented in a server that can communicate with the first computer via the Internet, or a local address for accessing the AI agent implemented in the communication terminal that can communicate with the first computer without going through the Internet, The information processing method according to claim 1 , wherein the connection to the second computer is executed via a common API regardless of whether the address is the global address or the local address.
8. The information processing method according to claim 1, wherein the AI agent is the AI agent last used by the user, selected from one or more AI agents available to the user based on a usage log stored in the communication terminal.
9. The information processing method described in claim 1, wherein the acquisition of the access information is performed when it is determined based on the authentication information acquired from the communication terminal that information regarding an AI agent available to the user is not registered in the memory of the first computer.
10. the authentication information includes information on an electronic key for the user to use the vehicle; The information processing method according to claim 1 , wherein the AI agent is an AI agent selected from one or more AI agents associated with the electronic key in the communication terminal.
11. The information processing method according to claim 1, wherein the AI agent is selected from the one or more AI agents by the user selecting the electronic key from a plurality of electronic keys available to the user via a UI of the communication terminal.
12. The plurality of electronic keys include a company electronic key associated with a company to which the user belongs and a personal electronic key of the user; If the selected electronic key is the company electronic key, a company agent managed by the company is selected as the AI agent; The information processing method according to claim 11 , wherein if the selected electronic key is the personal electronic key, a private agent of the user is selected as the AI agent.
13. A computer installed in a vehicle, a processor; A computer comprising: a memory storing a program for causing the processor to execute the information processing method according to any one of claims 1 to 12 as the first computer.
14. A program for causing the first computer to execute the information processing method according to any one of claims 1 to 12.
15. An information processing method for a dialogue agent system capable of dialogue with a user, the method being executed by a communication terminal capable of communicating with a first computer mounted on a vehicle and having an AI agent used by the user implemented therein, transmitting authentication information for the user to use the vehicle to the first computer; After determining that the authentication information is valid, sending access information to the first computer for accessing the AI agent; transmitting agent attribute information for representing a character of the AI agent to the first computer, and causing the first computer to set at least one of a GUI and a VUI based on the agent attribute information; While connected to the first computer, executing an interactive process with the user using at least one of the GUI and the VUI and a process generated by the AI agent. Information processing methods.
16. the first computer includes a display for displaying the GUI, and a speaker and a microphone for inputting and outputting the VUI; The information processing method according to claim 15, wherein the communication terminal comprises an AI processor for causing the AI agent to execute the generation process, and a memory for storing the access information including the address of the AI agent, and the agent attribute information relating to the AI agent.
17. the communication terminal includes a memory that stores a usage log of the AI agent by the user; The information processing method according to claim 15, wherein the interaction processing with the user in the vehicle refers to the usage log including the content of the conversation between the AI agent and the user before the user uses the vehicle.
18. the communication terminal includes a memory that stores a usage log relating to one or more AI agents available to the user; The information processing method of claim 15, wherein after it is determined that the authentication information is valid, one of the one or more AI agents last used by the user is selected as the AI agent based on the usage log.
19. the communication terminal manages information about a plurality of electronic keys available to the user; receiving an input operation from the user regarding which of the plurality of electronic keys to select; Selecting the AI agent associated with the electronic key according to the electronic key selected by the user; The information processing method according to claim 15 , further comprising transmitting the access information for accessing the selected AI agent to the first computer after determining that the authentication information is valid.
20. The plurality of electronic keys include a company electronic key associated with a company to which the user belongs and a personal electronic key of the user; If the selected electronic key is the company electronic key, a company agent managed by the company is selected as the AI agent; 20. The information processing method according to claim 19, wherein if the selected electronic key is the personal electronic key, a private agent of the user is selected as the AI agent.
21. a processor; A communication terminal comprising: a memory storing a program for causing the processor to execute the information processing method according to any one of claims 15 to 20.
22. A program for causing the communication terminal to execute the information processing method according to any one of claims 15 to 20.
23. An interactive agent system capable of interacting with a user, comprising: a first computer mounted on the vehicle; a second computer on which an AI agent is implemented; the first computer includes a processor and a memory storing a program for causing the processor to execute predetermined information processing; The predetermined information processing includes: The first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the communication terminal of the user and determines that the authentication information is valid, and then acquires access information for accessing the AI agent used by the user from the communication terminal; acquiring, based on the access information, agent attribute information for representing a character of the AI agent from an agent database that can communicate with the first computer; setting at least one of a GUI and a VUI representing the character of the AI agent in the first computer based on the agent attribute information; and performing an interactive process with the user using at least one of the GUI and the VUI and a process generated by the AI agent while connecting to a second computer having a processor on which the AI agent is implemented based on the access information. Dialogue agent system.
24. 1. An information processing method executed by a first computer mounted on a vehicle in an interactive agent system capable of interacting with a user, comprising: The first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, and then establishes a connection to a communication terminal on which a multimodal AI agent is implemented; While the user is using the vehicle, sensing data acquired via one or more sensors installed in the vehicle is converted into output data of a format and type that can be handled by the AI agent for each unit of data acquired over a predetermined period of time; An information processing method that periodically or intermittently transmits the output data to the AI agent in the communication terminal, and causes the AI agent to execute dialogue processing that reflects the latest output data.
25. After determining that the authentication information is valid, obtain data format information from the communication terminal that indicates a data format and a data type that the AI agent can handle; 25. The information processing method according to claim 24, wherein the conversion to the output data is performed in accordance with the data format information.
26. the sensing data includes at least one of facial expressions, gestures, emotions, and biological information of the user, an interior environment of the vehicle, a surrounding situation around the vehicle, and a running state of the vehicle; 25. The information processing method according to claim 24, wherein the output data includes a type of data from which the output data was extracted, an identification code for identifying a state or event of the vehicle or the user, and a timestamp indicating a time when the state or the event occurred.
27. 25. The information processing method according to claim 24, wherein the AI agent changes the timing or amount of information to be presented to the user when it determines, based on the output data from the first computer, that the user is in a situation where he or she should concentrate on driving the vehicle.
28. 25. The information processing method of claim 24, wherein when the AI agent infers based on the output data from the first computer that the user has indicated an intention to respond to the AI agent, and when the AI agent is unable to obtain a response from the user via the VUI of the AI agent, the AI agent performs interactive processing in accordance with the inference based on the output data.
29. the first computer includes a processor in which an in-vehicle AI agent different from the AI agent is implemented; causing the in-vehicle AI agent to generate request data indicating constraints required for the dialogue processing executed by the AI agent based on the output data; 26. The information processing method according to claim 25, further comprising transmitting the request data to the AI agent in the communication terminal in addition to the output data, and causing the AI agent to execute interactive processing based on the output data and the request data.
30. A computer installed in a vehicle, a processor; 30. A computer comprising: a memory storing a program for causing the processor to execute the information processing method according to any one of claims 24 to 29 as the first computer.
31. 30. A program for causing the first computer to execute the information processing method according to any one of claims 24 to 29.
Citation Information
Patent Citations
Agent system, agent server, and agent program
JP2021117302A