Information processing method, computer, program, communication terminal, and interactive agent system

The method allows vehicles to acquire and set AI agent characteristics from user terminals, ensuring seamless interaction by setting GUI and VUI, addressing the challenge of transitioning AI agent use from communication terminals to in-vehicle systems.

WO2026004526A1PCT designated stage Publication Date: 2026-01-02PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/020449
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-06-05
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing systems face challenges in seamlessly continuing personalized interactions with AI agents when users transition from their communication terminals to in-vehicle systems, particularly in vehicles they are using for the first time or with multiple users.

Method used

An information processing method that enables a vehicle-mounted computer to acquire and set the character of a user's AI agent from their communication terminal, allowing seamless interaction by setting a GUI and/or VUI based on agent attribute information, and facilitating dialogue processing using the AI agent.

Benefits of technology

Enables the continuation of personalized interactions with AI agents within vehicles, allowing users to access and interact with their preferred agents effortlessly, even in new or shared vehicles, through a simple authentication process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025020449_02012026_PF_FP_ABST
    Figure JP2025020449_02012026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing method according to one embodiment of the present disclosure is executed by a first computer mounted on a vehicle in an interactive agent system which can interact with a user. The information processing method involves: after the first computer or an authentication computer capable of communicating with the first computer has acquired authentication information from a communication terminal of a user and has determined that said authentication information is valid, acquiring, from the communication terminal, access information for accessing an AI agent for use by the user; acquiring, on the basis of the access information and from an agent database capable of communicating with the first computer, agent attribute information for representing a character of the AI agent; setting, on the basis of the agent attribute information and in the first computer, a GUI and / or a VUI representing the character of the AI agent; and executing, on the basis of the access information, an interaction process with the user by using the GUI and / or the VUI and a generation process by the AI agent while connecting to a second computer on which the AI agent is implemented.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, computer, program, communication terminal, and interactive agent system

[0001] The present disclosure relates to an information processing method, a computer, a program, a communication terminal, and a dialogue agent system.

[0002] Patent Document 1 discloses an in-vehicle system in which the presence of an agent is planned.

[0003] Japanese Patent Application Laid-Open No. 2021-117302

[0004] An information processing method according to one aspect of the present disclosure is an information processing method executed by a first computer installed in a vehicle in an interactive agent system capable of interacting with a user, wherein the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, and then acquires access information for accessing the AI ​​agent used by the user from the communication terminal, acquires agent attribute information for representing the AI ​​agent's character from an agent database capable of communicating with the first computer based on the access information, sets at least one of a GUI and a VUI representing the AI ​​agent's character on the first computer based on the agent attribute information, and, while connecting to a second computer on which the AI ​​agent is implemented based on the access information, performs interactive processing with the user using at least one of the GUI and VUI and generation processing by the AI ​​agent.

[0005] FIG. 1 is a diagram showing an example of the overall configuration of a dialogue agent system according to this embodiment. FIG. 2 is a diagram showing an example of the hardware configuration of the dialogue agent system according to this embodiment. FIG. 3 is a sequence diagram showing an example of API exchange via a wide area communication network between an information terminal and a cloud in the dialogue agent system according to this embodiment. FIG. 4 is a sequence diagram showing an example of a process flow in which a client downloads and uses agent attribute information in the dialogue agent system according to this embodiment. FIG. 5 is a diagram showing an example of the functional configuration of the dialogue agent system according to this embodiment. FIG. 6 is a sequence diagram showing an example of a process flow in which interaction with an agent is continued even when the client is switched in the dialogue agent system according to this embodiment. FIG. 7-1 is a sequence diagram showing an example of a process flow in which interaction with an agent is continued even when the client is switched in the dialogue agent system according to this embodiment. FIG. 7-2 is a sequence diagram showing an example of a process flow in which interaction with an agent is continued even when the client is switched in the dialogue agent system according to this embodiment. FIG. 8 is a sequence diagram showing an example of a process flow in which a client enhances interaction with an agent in the dialogue agent system according to this embodiment. FIG. 9 is a flowchart showing an example of a process flow in which a client inquires about the type of multimodal data input to the dialogue agent system according to this embodiment and sends the data. Fig. 10 is a diagram for explaining an example of the structure of additional data converted into text data from multimodal raw data in the dialogue agent system according to this embodiment. Fig. 11 is a diagram for explaining an example of the structure of data converted into text data from multimodal raw data in the dialogue agent system according to this embodiment. Fig. 12 is a diagram for explaining an example of a dialogue agent system according to this embodiment. Fig. 13 is a diagram for explaining an example of a dialogue agent system according to this embodiment. Fig. 14 is a diagram for explaining an example of a dialogue agent system according to this embodiment.Fig. 15 is a diagram for explaining an example of multimodal data in the dialogue agent system according to this embodiment. Fig. 16 is a diagram showing an example of the functional configuration of the dialogue agent system according to this embodiment.

[0006] [Findings underlying the present disclosure] In recent years, technological development of AI agents using large-scale language models (LLMs) has progressed. AI agents have short-term and long-term memories (user usage logs or portions of their contents) and can autonomously communicate with external applications and web services via a network, as well as launch and operate other applications and web services. As a result, AI agents are computer systems or software programs that set or update goals through text or voice communication with users (instructions to the AI ​​are also called prompts), autonomously generate a set of tasks necessary to achieve the goal, and sequentially execute the information processing of the generated tasks autonomously or while communicating with the user, thereby achieving the final goal. AI that can process input or output information not only in a single modal (data format), such as text, but also in a combination of multiple different modalities, such as voice and images, is called multimodal AI.

[0007] AI agents can be given specializations and characteristics depending on the databases they refer to when processing information and the algorithms they use to generate tasks. This allows them to be implemented as highly specialized agents specialized in specific functions, for example. On the other hand, AI agents can also be implemented as personalized AI agents that are closely aligned with individual users by learning the preferences, biometric information, and past behavioral history of the users they communicate with, and by accessing databases that store such personal user data (hereinafter referred to as user attribute information). The former AI agents are sometimes called specialized agents, and the latter AI agents are sometimes called partner agents.

[0008] Partner-type agents are thought to be particularly effective partners in mobility spaces. This is because when a user travels by vehicle to a place outside their usual range of activity and has a new (or unusual) experience, a partner-type agent can act as an appropriate navigator that is attuned to the user's individuality. For example, when navigating a vehicle's route, a partner-type agent can make selections or suggestions that reflect the user's preferences. Examples of user preferences include whether the user prefers the shortest route, roads that are easy to drive on with separated sidewalks, or whether the user likes to stop by tourist spots.

[0009] The present inventors have considered a series of user experiences related to the use of a vehicle and an AI agent. The initial scenario envisioned is a scenario in which a user uses an AI agent using a smartphone or the like before getting into the vehicle, and then continues to use the AI ​​agent after getting into the vehicle using an information device (e.g., an in-vehicle infotainment system; IVI). In such a scenario, there is a need for seamless continuation of interactions with a partner-type agent that is personalized to the user, even after getting into the vehicle. However, there are challenges in terms of customer value and implementation means regarding how to appropriately set up a personal partner-type agent in a vehicle that a user is using for the first time or in a vehicle that multiple users use.

[0010] The following aspects of the present disclosure are based on the above findings, but the claimed invention is not limited to the above findings.

[0011] [Summary of the Embodiment] An information processing method according to one aspect of the present disclosure is an information processing method executed by a first computer mounted on a vehicle in an interactive agent system capable of interacting with a user. The information processing method includes the steps of: the first computer or an authentication computer capable of communicating with the first computer acquiring authentication information from the user's communication terminal and determining that the authentication information is valid, and then acquiring, from the communication terminal, access information for accessing an AI agent used by the user; acquiring, based on the access information, agent attribute information for representing a character of the AI ​​agent from an agent database capable of communicating with the first computer; setting, based on the agent attribute information, at least one of a GUI and a VUI representing the character of the AI ​​agent on the first computer; and, based on the access information, performing an interactive process with the user using at least one of the GUI and the VUI and a generation process by the AI ​​agent while connecting to a second computer on which the AI ​​agent is implemented.

[0012] According to this information processing method, a first computer installed in a vehicle can acquire agent attribute information related to the character of an AI agent used by a user via the user's communication terminal and set the agent attribute information in the first computer's GUI (Graphical User Interface) and / or VUI (Voice User Interface). For example, if a second computer on which an AI agent is implemented is a communication terminal, the communication terminal can function as the AI ​​agent's brain, and the first computer can function as the AI ​​agent's body (face and / or voice). For example, if the second computer on which an AI agent is implemented is a server, the server can function as the AI ​​agent's brain, and the first computer can function as the AI ​​agent's body (face and / or voice). This allows an AI agent that a user regularly uses via a communication terminal to be represented on the first computer installed in the vehicle and used within the vehicle.

[0013] The process from obtaining access information to setting the AI ​​agent's character may be triggered by determining that the authentication information is valid. This allows the process from the authentication process to the setting process for the AI ​​agent's character to be executed smoothly. For example, even in a vehicle the user is riding in for the first time or in a vehicle that requires the use of another AI agent's character, the user can call and use their own partner AI agent in the vehicle through a simple procedure, automatically, or in a short time. When the vehicle unlocking process is executed as a trigger of the authentication process, the setting of the AI ​​agent can be completed before or immediately after the user sits in the seat.

[0014] An information processing method according to one aspect of the present disclosure is an information processing method executed by a communication terminal capable of communicating with a first computer mounted on a vehicle and implementing an AI agent used by the user in an interactive agent system capable of interacting with a user, the information processing method including: transmitting authentication information for the user to use the vehicle to the first computer; after determining that the authentication information is valid, transmitting access information for accessing the AI ​​agent to the first computer; transmitting agent attribute information for representing a character of the AI ​​agent to the first computer and causing the first computer to set at least one of a GUI and a VUI based on the agent attribute information; and, while connected to the first computer, executing an interactive process with the user using at least one of the GUI and the VUI and a generation process by the AI ​​agent.

[0015] According to this information processing method, the AI ​​agent in the communication terminal that the user uses on a daily basis can be used while being represented on the first computer installed in the vehicle.

[0016] According to one aspect of the present disclosure, there is provided an information processing method executed by a first computer mounted on a vehicle in a dialogue agent system capable of dialogue with a user. The information processing method includes: establishing a connection to a communication terminal implementing a multimodal AI agent after the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid; converting sensing data acquired via one or more sensors installed in the vehicle while the user is using the vehicle into output data in a format and type compatible with the AI ​​agent for each unit of data acquired over a predetermined period; and periodically or intermittently transmitting the output data to the AI ​​agent in the communication terminal and causing the AI ​​agent to execute dialogue processing reflecting the latest output data.

[0017] According to this information processing method, an AI agent installed in a communication terminal can receive information sensed by a vehicle and carry out dialogue with the user. In this case, output data downsized to a format and type that the AI ​​agent can handle is transmitted to the communication terminal, rather than the actual sensing data (e.g., video and audio) detected by the vehicle. This allows the AI ​​agent in the communication terminal to carry out dialogue while appropriately incorporating information from the vehicle.

[0018] [Embodiments] Hereinafter, exemplary embodiments of an information processing method, a computer, a program, a communication terminal, and a dialogue agent system according to the present disclosure will be described with reference to the drawings.

[0019] First, the definition of an AI agent (Artificial Intelligence Agent) will be explained. An AI agent is software or a mechanism for achieving a predefined goal. An AI agent is designed to autonomously generate and select actions to achieve a goal and execute those actions to achieve the goal based on communication with a user, the situation of the user's surrounding space, external information regarding dialogue with the user acquired via a network, and the like. Communication with a user includes all means for conveying the user's emotions, will, and thoughts. For example, this includes one or more of the following means: conversation, timing before speaking, tone of voice, expression of intention via visual information such as letters and symbols (GUI), expression of intention via physical operations such as buttons and switches, expression of intention via facial expressions, gaze, posture, or bodily actions such as gestures. In this embodiment, an AI agent is also simply referred to as an agent. An agent may be represented as a unique character to facilitate communication with a user, in which case it has attribute information of that character. The agent's attribute information includes, for example, information about one or more of the character's appearance, body, clothing, accessories, gestures, facial expressions, voice, personality, preferences, habits, knowledge, or experience (records of past interactions with users).

[0020] FIG. 1 is a diagram showing an example of the overall configuration of a dialogue agent system according to the present embodiment (in this disclosure, an agent system that communicates with a user as described above will be referred to as a "dialogue agent," "dialogue agent system," or simply an "agent"). The dialogue agent system according to the present embodiment is an example of a dialogue agent system that can dialogue with a user (the communication between a user and an agent system is not limited to dialogue, and the above-mentioned communication may also take place; however, for ease of reading, the phrase "capable of dialogue" will be used in this disclosure rather than "capable of communication"). As shown in FIG. 1, the dialogue agent system according to the present embodiment includes an information terminal 1, a vehicle 2, a cloud 3, etc. The information terminal 1, the vehicle 2, and the cloud 3 are connected to each other so as to be able to communicate with each other via a network (for example, a wide area communication network such as the Internet).

[0021] The user's information terminal 1 stores information about an electronic key for unlocking and starting a vehicle 2. The agent operates by AI inference processing in devices such as the information terminal 1, the vehicle 2, and the cloud 3.

[0022] The cloud agent is an agent that performs processing in response to a request on the cloud 3 and generates a response. The corporate agent and private agent are partner-type agents in which AI inference processing runs on an information terminal 1 such as a smartphone. Specifically, the corporate agent is a work partner of the user, and is an agent that is managed in association with a corporate electronic key for using the vehicle 2 for work purposes. The private agent is a private partner, and is an agent that is managed in association with a personal electronic key for using the vehicle 2 for private purposes.

[0023] A vehicle agent is an agent running on an in-vehicle system (a computer system installed in vehicle 2). A cloud agent is an agent that operates in a server-client format via the Internet. A corporate agent, a private agent, and a vehicle agent are agents that can operate using only the computer resources and data of the device, even when there is no Internet connection. When any agent has an Internet connection, it may perform a data search on the Internet to obtain information useful for generating a response and generate the response.

[0024] Cloud agents, corporate agents, and private agents are agents that can be used by users from any terminal (client) and can interact with the user after gaining a deep understanding of the user's actions, thoughts, and experiences by working together with the user while completing daily tasks. Vehicle agents are agents that can be used in in-vehicle systems and are specialized for vehicle-related tasks with unique functions such as supporting safe driving, setting and updating routes according to traffic conditions, and answering questions and setting instructions about vehicle 2.

[0025] 2 is a diagram showing an example of the hardware configuration of the interactive agent system according to this embodiment. The information terminal 1 includes a sensor unit 101 for acquiring video information, audio information, and / or physical quantities of the surrounding environment, a UI unit 103 for providing the user with video and audio information and accepting button presses, touch operations, etc., a calculation unit 104 for performing various calculations including learning and inference processing of an AI model performed within the information terminal 1 and information processing such as information drawing, a memory 105 for storing data and files used by the calculation unit 104, and a communication unit 106 for communicating with other computers on a communication network.

[0026] The UI unit 103 has a display that displays a GUI (Graphical User Interface) and a speaker and microphone that input and output a VUI (Voice User Interface). The calculation unit 104 is an example of an AI processor that causes an agent (e.g., a partner-type agent) to execute a generation process. The memory 105 is an example of a memory that stores access information including the address of the agent and agent attribute information related to the agent. The memory 105 is also an example of a memory that stores a usage log of the agent by the user. The memory 105 is also an example of a memory that stores a usage log related to one or more agents available to the user. The memory 105 also manages (stores) information related to multiple electronic keys available to the user. Here, the multiple electronic keys include a company electronic key associated with the company to which the user belongs and the user's personal electronic key.

[0027] When an application for managing keys is installed in the information terminal 1, the program and necessary data are recorded in the memory 105 of the information terminal 1, and the program is executed by the calculation unit 104.

[0028] In this embodiment, the information terminal 1 is described as a smartphone, but is not limited to this. The information terminal 1 may be in the form of a wristwatch-type smart watch, smart glasses-type eyeglasses, smart earphones worn on the ears, a smart ring-type finger ring, a smart speaker operated by voice, or a robot with movable parts.

[0029] The cloud 3 includes a communication unit 301 for communicating with other computers on a communication network (wide area communication network), a memory 302 that stores information about the vehicle 2 and the user and its management program, and a calculation unit 303 that performs various data processing.

[0030] Vehicle 2 has a movable unit 201 for moving vehicle 2 and operating equipment (seats, etc.) within the vehicle cabin, a lighting unit 202 for illuminating the area around vehicle 2, a sensor unit 203 for detecting the position and status of people and cars around vehicle 2, as well as people and objects within the vehicle cabin, a UI unit 204 for providing passengers with various video and audio information and accepting input from passengers such as touch operations and voice operations, a key control unit 205 for authenticating the key to be unlocked and controlling the locking / unlocking of vehicle 2's doors, a calculation unit 206 for executing various processes related to the vehicle core system and vehicle functions, a memory 207 for recording various data including the vehicle core system's program and key management database, and a communication unit 208 for wireless communication with external devices.

[0031] In this embodiment, the UI unit 204, the key control unit 205, the calculation unit 206, the memory 207, and the communication unit 208 are realized by an in-vehicle system mounted on the vehicle 2. Here, the calculation unit 206 is an example of a processor, and may be a processor that executes a vehicle agent different from the agent. Also, here, the memory 207 is an example of a memory that stores a program for causing the calculation unit 206 to execute predetermined information processing. Here, the predetermined information processing includes processing executed by the acquisition unit 206a, the setting unit 206b, and the execution unit 206c (see FIG. 5), which will be described later.

[0032] The vehicle 2, the information terminal 1, and the cloud 3 may communicate with each other via a communication means other than the wide area communication network Internet. For example, the unlocking authentication process performed between the vehicle 2 and the information terminal 1 may use short-range wireless communication.

[0033] FIG. 3 is a sequence diagram showing an example of API exchanges between an information terminal and a cloud via a wide area communication network in a dialogue agent system according to this embodiment. First, the premise of the dialogue agent system according to this embodiment will be described. Due to the limited computational resources required to run large-scale language models in current AI, in many cases, an information terminal 1 (client) such as a smartphone is used as a UI terminal. A user inputs questions or requests in text or voice into an app or web browser on the information terminal 1. The app or web browser on the information terminal 1 transmits the input data to the cloud 3 (such as an address indicated by a URL (API endpoint) specified for using the AI ​​model) via an API or HTTP / HTTPS protocol.

[0034] The API (Application Programming Interface) and HTTP / HTTPS protocol are rules for communication between the app or web browser of the information terminal 1 and the web app of the cloud 3, and are responsible for exchanging data sent and received between them in a predetermined format (e.g., HTTPS request / response). Upon receiving a request, the cloud 3 returns the response to the information terminal 1 in a predetermined format (e.g., HTTPS response). The API and HTTP / HTTPS protocol are responsible for authentication to ensure security, sending and receiving requests and responses, etc. The cloud 3 (server) processes data and generates a response. The cloud 3 processes requests using vast computing resources and large amounts of data. Inference processing using an AI model is executed on the cloud 3, and image processing, audio processing, data processing, natural language processing, etc. are performed as appropriate in response to the request.

[0035] For example, an information terminal 1 such as a smartphone requests user authentication from the cloud 3 (step S301). If the cloud 3 successfully authenticates the user (step S302), it notifies the information terminal 1 of the success of user authentication (step S303). If the authentication fails, the process ends here. Upon being notified of the success of user authentication, the information terminal 1 acquires a user request (step S304) and transmits the acquired request to the cloud 3 (step S305). The cloud 3 executes data processing in response to the request received from the information terminal 1 (step S306) and notifies the information terminal 1 of the result as a response (step S307). The information terminal 1 outputs the notified response (step S308). Thereafter, the information terminal 1 and the cloud 3 repeatedly exchange requests and responses (steps S309 to S313).

[0036] Looking at the overall system configuration of the interactive agent system according to this embodiment, it can also be seen as a single device with the information terminal 1 as the UI unit, the cloud 3 as the calculation unit that generates responses, and a data communication path including the Internet as the data bus. The data exchanged here can use the same API / HTTPS protocol whether the UI unit and the calculation unit are contained in a single device or whether they are composed of two physically different devices, and does not need to depend on the implementation form.

[0037] In other words, if the UI unit and the processing unit are in different terminals, they communicate with each other via a network using a predetermined API / protocol to carry out processing. Even if the UI unit and the processing unit are in a single terminal, they can communicate with each other via a bus within the terminal using the same predetermined API / protocol to carry out processing, allowing for a high degree of freedom in the combination of embodiments of the interactive agent system.

[0038] For example, the processing described in FIG. 3 as being executed by the information terminal 1 may be executed by an application acting as a client running on the in-vehicle system of the vehicle 2. Alternatively, the processing described as being executed by a client and a server may be executed by a UI unit and a calculation unit, respectively, within a single information terminal 1. As long as the same predetermined API and protocol are used, the function of the UI unit (or client) in the interactive agent system may be implemented or executed in the information terminal 1 or the vehicle 2, and the function of the calculation unit (or server) that performs inference processing of the AI ​​model may be implemented or executed in the cloud 3, the information terminal 1, or the vehicle 2. Of course, these may be realized by software running on the information terminal 1, the vehicle 2, or the cloud 3. In other words, all of the disclosures in this embodiment can be realized in any of these configurations, and may be realized in any system configuration.

[0039] 4 is a sequence diagram showing an example of the process flow in which a client downloads and uses agent attribute information in the interactive agent system according to this embodiment. First, information terminal 1 requests user authentication from cloud 3 (step S401). If user authentication is successful (step S402), cloud 3 obtains agent attribute information of the agent specified (or default) by the user from memory 302 (step S403), and transmits the obtained agent attribute information to information terminal 1 (step S404).

[0040] Upon receiving the agent attribute information, the information terminal 1 sets the agent attribute information (step S405) and activates the agent based on the set agent attribute information (step S406). This allows the user to recognize the specified agent and understand that they can begin communication, such as talking to the agent, based on the agent's status displayed on the UI unit 103. The activated agent then acquires a request based on the user's speech (step S407) and transmits the acquired request to the cloud 3 using a predetermined protocol / API (step S408). The cloud 3 executes data processing in response to the request received from the information terminal 1 (step S409) and notifies the information terminal 1 of the result as a response using the predetermined protocol / API (step S410). The cloud 3 updates the user attribute information based on the interaction with the user (step S411) and updates the agent usage log (a database recording the date and time of interactions between the user and the agent, the content of the conversation, etc.) (step S412). The information terminal 1 then responds to the user via the agent displayed on the UI unit 103 in accordance with the received response (step S413). Thereafter, the information terminal 1 and the cloud 3 repeatedly exchange requests and responses.

[0041] User attribute information is a database that includes, for example, one or more of the user's name (nickname), age, gender, career history, interests, preferences, past conversation history, thoughts, experiences, schedule information, frequently used or subscribed external information / services, incomplete or unresolved tasks, unique information (identification information, access information) of the device used by the user, biometric information, medical history, or behavioral history (movement history).

[0042] By storing agent attribute information on the cloud 3 (server side), the agent attribute information can be represented on any information terminal 1 (client) when accessed from that information terminal 1. Agent attribute information is a data set used to represent an agent, including, for example, one or more of 3D model data (data defining the agent's appearance and 3D physique, including texture, clothing, etc.), animation data (data on facial expressions, mouth, and gestures for reproducing natural movements), voice data (data for reproducing the characteristics of an agent's vocalizations), emotion data (data for reproducing specific behavioral patterns based on emotions), or control scripts (control codes for ensuring consistency of the agent's overall behavior when moving and responding).

[0043] In the cloud 3, a usage log (user attribute information) recording interactions between the agent and the user is stored and managed while being updated continuously in memory 302. This makes it possible to search for past events with the user and provide answers or suggestions. For example, if a user asks the agent about the final status of a specific matter, the agent can refer to the conversation history in the user attribute information to find the final status of that matter and generate an answer. This type of technology is called RAG (Retrieval-Augmented Generation), and is known as a natural language processing technology that combines information retrieval and generative modeling.

[0044] The cloud 3 may store user attribute information (interests, preferences, thoughts, etc.) obtained through interactions between the agent and the user, and update it in memory 302. Furthermore, by recording, managing, and updating information related to the user's knowledge system and experience as user attribute information, it becomes possible to provide replies and suggestions based on the user's knowledge system (which allows interactions based on the user's knowledge in areas in which the user is knowledgeable) and experience (recorded data such as the user's past experiences and events that have occurred).

[0045] The contents of the request and response may be text data such as text chat, or may be data such as images, videos, and audio. If the information terminal 1 (client) on which the agent that interacts with the user operates is equipped with various sensors, modal data other than the above-mentioned text, images, and audio may be included.

[0046] 5 is a diagram showing an example of the functional configuration of the interactive agent system according to this embodiment. In this embodiment, an acquisition unit 206a, a setting unit 206b, an execution unit 206c, etc. are realized by a calculation unit 206 of an in-vehicle system (an example of a first computer) of the vehicle 2 executing a program stored in a memory 207. In this embodiment, the acquisition unit 206a, the setting unit 206b, the execution unit 206c, etc. are realized in the vehicle 2, but they may also be realized in cooperation with the user's information terminal 1 or the cloud 3.

[0047] The acquisition unit 206a (a first computer or an authentication computer capable of communicating with the first computer) acquires authentication information from the user's information terminal 1 (an example of a communication terminal). Here, the authentication computer is a computer capable of communicating with the in-vehicle system of the vehicle 2, and may be, for example, a cloud 3 equipped with AI or a server of a vehicle dispatching service provider or the like. Here, the authentication information may include information on an electronic key for the user to use the vehicle 2.

[0048] Furthermore, after determining that the authentication information is valid, the acquisition unit 206a acquires access information for accessing the agent used by the user from the information terminal 1. That is, once authentication between the information terminal 1 and the vehicle 2 is obtained, the acquisition unit 206a receives access information for the agent (for example, connection destination information such as an API end point) from the information terminal 1 and establishes a connection. The acquisition of access information may be executed when it is determined, based on the authentication information acquired from the information terminal 1, that information about an agent available to the user is not registered in the memory 207 of the in-vehicle system.

[0049] Furthermore, after determining that the authentication information is valid, the acquisition unit 206a may establish a connection to the information terminal 1 (or the cloud 3) on which a multimodal agent is running. In this case, the acquisition unit 206a acquires sensing data acquired via one or more sensor units 101 provided in the vehicle 2 while the user is using the vehicle 2. Here, the sensing data may include at least one of the user's facial expressions, gestures, emotions, and biometric information, the interior environment of the vehicle 2, the surrounding conditions around the vehicle 2, and the driving state of the vehicle 2. Here, the output data may include the type of data from which the sensing data was extracted, an identification code for identifying the state or event of the vehicle 2 or the user, and a timestamp indicating the time when the user's state or the event occurred. Here, the multimodal agent is not limited to a partner-type agent and may be, for example, a cloud agent or a vehicle agent.

[0050] Then, the acquiring unit 206a converts the acquired sensing data into output data of a format and type that the agent can handle, for each unit of data acquired during a predetermined period. Furthermore, after determining that the authentication information is valid, the acquiring unit 206a may acquire data format information that indicates a data format and data type that the agent can handle from the information terminal 1. In this case, the acquiring unit 206a may convert the sensing data into output data in accordance with the acquired data format information.

[0051] Here, if the response generation part of the agent (such as the inference processing of the AI ​​model) is implemented in the information terminal 1, the access information may include an address for accessing the agent in the information terminal 1. Also, here, if the agent is implemented in a server that can communicate with the in-vehicle system of the vehicle 2 via a wide area communication network, the access information may include an address (such as an API endpoint) for accessing the agent in the server via the wide area communication network.

[0052] Here, the address included in the access information may be a global address for accessing an agent implemented on a server that can communicate with the in-vehicle system via a wide-area communication network, or a local address for accessing an agent implemented on an information terminal 1 that can communicate with the in-vehicle system without going through a wide-area communication network. Connection to the information terminal 1 may be performed via a common API, regardless of whether the address included in the access information is a global address or a local address. This allows the in-vehicle system to connect to both the agent on the server side and the agent on the edge side (information terminal 1, such as a smartphone) using the common API.

[0053] Furthermore, based on the access information, the acquisition unit 206a acquires agent attribute information (an example of agent attribute information) for representing the agent character from a cloud 3 that can communicate with the calculation unit 206 (an agent database realized by the cloud 3 on which AI inference processing is executed, an information terminal 1 such as a smartphone, or another server (an example of a third computer)). That is, the acquisition unit 206a acquires agent attribute information including information on the agent's UI, etc. from the information terminal 1 or the cloud 3. Here, the agent database may be located in the cloud 3 or the information terminal 1, or a third computer such as another server with which the in-vehicle system of the vehicle 2 can communicate via a network such as a wide area communication network.

[0054] The setting unit 206b sets at least one of a GUI and a VUI representing the character of the agent in the calculation unit 206 based on the agent attribute information.

[0055] Based on the access information, the execution unit 206c connects to the cloud 3 on which the agent is implemented or the information terminal 1 (an example of a second computer or a communication terminal), and executes an interaction process with the user using at least one of a GUI and a VUI and a response generation process by the agent. Here, the GUI may be displayed on a display of the UI unit 204. The VUI may be input and output via a speaker and a microphone of the UI unit 204. In addition, in the interaction process with the user inside the vehicle 2, the execution unit 206c may refer to a usage log (or user attribute information) including the content of the conversation between the agent and the user before the user used the vehicle 2. This allows the user to have a conversation with the agent after getting into the vehicle 2 using the conversation log from before getting into the vehicle 2.

[0056] Furthermore, the agent that is automatically activated when the user gets in or out of the vehicle is selected from one or more agents available to the user based on the usage log stored in the information terminal 1. In this case, the agent may be the agent that the user last used before getting in or out of the vehicle. In other words, even if the device the user is using changes, the agent with which the user most recently interacted may be called up on the device to be used next. Alternatively, if the authentication information at the time of getting in the vehicle includes information about an electronic key for the user to use the vehicle 2, the agent may be an agent selected in the information terminal 1 from one or more agents associated with the electronic key. In other words, an agent linked to the electronic key used to unlock or start the vehicle 2 in which the user is riding may be automatically called up on the in-vehicle system of the vehicle 2 when the user gets in the vehicle.

[0057] Alternatively, an agent may be selected from one or more agents by the user selecting an electronic key from a plurality of electronic keys available to the user via the UI unit 103 of the information terminal 1. Here, the plurality of electronic keys may include a corporate electronic key associated with the user's company and the user's personal electronic key. If the selected electronic key is a corporate electronic key, a corporate agent linked to the corporate electronic key may be selected as the agent. On the other hand, if the selected electronic key is a personal electronic key, a private agent linked to the personal electronic key may be selected as the agent.

[0058] Furthermore, if a multimodal agent is implemented in the information terminal 1, the execution unit 206c may periodically or intermittently transmit output data converted from the sensing data by the acquisition unit 206a to the agent in the information terminal 1, and cause the agent to execute dialogue processing reflecting the latest sensing data regarding the user and / or the vehicle 2. In this case, if the agent determines, based on the output data, that the user is in a situation where they should concentrate on driving the vehicle 2, it may change the timing or amount of information presented to the user. Furthermore, if the agent determines, based on the output data, that the user has indicated their intention to respond to the agent, but has not yet obtained a response from the user via the agent's VUI, it may execute dialogue processing in accordance with the inference based on the output data. This allows the camera on the vehicle 2 to determine that the user is reacting in some way, but allows the user to confirm this if there is no voice input or the voice input is unclear.

[0059] Furthermore, when the calculation unit 206 is equipped with a vehicle agent different from the agent, the execution unit 206c may cause the vehicle agent to generate request data indicating constraints on the dialogue processing executed by the agent based on the output data. In this case, the execution unit 206c may send the request data in addition to the output data to the agent in the information terminal 1 and cause the agent to execute the dialogue processing based on the output data and the request data. Here, the request data is a request made by the vehicle agent to the agent, such as "Please refrain from unnecessary conversation for now," and the dialogue processing of the agent based on this request data may be, for example, an affirmative response or may be immediately reflected in the dialogue processing with the user.

[0060] In this embodiment, the calculation unit 104 of the information terminal 1 executes a program stored in the memory 105 to realize the communication unit 104b, the execution unit 104a, etc. In this embodiment, the communication unit 104b, the execution unit 104a, etc. are realized in the information terminal 1, but they may also be realized in cooperation with the vehicle 2 or the cloud 3.

[0061] The communication unit 104b transmits authentication information for the user to use the vehicle 2 to the in-vehicle system of the vehicle 2. After it is determined that the authentication information is valid, the communication unit 104b transmits access information for accessing the agent to the in-vehicle system of the vehicle 2. Furthermore, the communication unit 104b transmits agent attribute information for representing the agent's character to the in-vehicle system of the vehicle 2, and causes the in-vehicle system to set at least one of a GUI and a VUI based on the agent attribute information. The communication unit 104b may transmit access information for accessing the agent selected by the execution unit 104a (described later) to the in-vehicle system of the vehicle 2 simultaneously with the authentication information, or after it is determined that the authentication information is valid.

[0062] The execution unit 104a executes an interactive process with the user using at least one of a GUI and a VUI and a response generation process by an agent while connected to the in-vehicle system of the vehicle 2. In the interactive process with the user in the vehicle 2, the execution unit 104a may also refer to a usage log including the content of a conversation between the agent and the user before the user uses the vehicle 2. Furthermore, the execution unit 104a may select, as the agent, one of one or more agents that the user last used based on the usage log, simultaneously with the authentication information or after it is determined that the authentication information is valid.

[0063] The execution unit 104a accepts an input operation from the user via the UI unit 103 regarding which of multiple electronic keys to select. The execution unit 104a may then select an agent associated with the electronic key selected by the user. Specifically, if the selected electronic key is a company electronic key, the execution unit 104a may select a company agent managed by the company as the agent. If the selected electronic key is a personal electronic key, the execution unit 104a may select the user's private agent as the agent. Note that although the description here assumes that the user selects an electronic key to be used by the user, the present disclosure is not limited to this. An app or agent on the information terminal may automatically select an electronic key and / or an agent suited to the vehicle's intended use by referring to past conversations with the user or a schedule.

[0064] 6 is a sequence diagram showing an example of the process flow for continuing interaction with an agent even when switching clients in the interactive agent system according to this embodiment (summoning an agent linked to an electronic key to a new client). In this example, the process shows a flow in which a user gets into a vehicle while having a conversation with an agent running on information terminal 1, and the conversation with the agent is automatically continued using the in-vehicle system as a UI unit.

[0065] First, an agent (e.g., a partner-type agent) implemented by a client application on the information terminal 1 receives a user request via the UI unit 103 (step S601) and transmits the request to a server application (or the agent's API endpoint) on the information terminal 1 that functions as the agent's brain (step S602). The server application executes data processing in response to the received request (step S603), transmits a response to the client application (step S605), and updates the usage log (step S604). The client application on the information terminal 1 responds to the user using the agent on the client application that functions as the body in accordance with the received response (step S606). In this process, the client application transmits the user's weather-related question, "What's the weather in Osaka?" to the server application via a predetermined API / protocol (step S602), and the server application returns the answer, "It's sunny, then it's going to rain," to the client application via a predetermined API / protocol (step S605). All of this processing is performed by the execution unit 104a of the information terminal 1.

[0066] When the user approaches the vehicle 2 or performs an unlocking operation, the vehicle 2 requests authentication of the electronic key via short-range wireless communication to the information terminal 1 (step S607). The electronic key application of the information terminal 1 authenticates the key based on the electronic key selected by the user (step S608), calculates a response value indicating the authentication result, and transmits the response value to the vehicle 2 (step S609). The on-board system of the vehicle 2 verifies the response value, and if the electronic key authentication is successful (step S610), unlocks the vehicle and / or starts the vehicle 2 (step S611).

[0067] The acquisition unit 206a of the vehicle 2 acquires, from the memory 207, access information (e.g., an API endpoint) of the agent associated with the electronic key selected by the user (step S612). The acquisition unit 206a may acquire this information by starting a client application installed in the vehicle 2. If the memory 207 does not contain access information associated with the electronic key or if the memory 207 is inaccessible, the acquisition unit 206a acquires, from the information terminal 1, the access information of the agent associated with the electronic key selected by the user (steps S613 and S614).

[0068] Next, the acquisition unit 206a uses the acquired access information to request an access, agent attribute information, a usage log, etc. from the agent for which the inference process of the AI ​​model is being executed on the information terminal 1 (step S615). When the server application of the information terminal 1 receives the access request from the agent, it transmits the agent attribute information and the most recent usage log of the agent to the vehicle 2 (step S616).

[0069] The setting unit 206b of the vehicle 2 sets the attributes of the agent that will communicate with the user based on the received agent attribute information, and the execution unit 206c starts the agent (step S617). Furthermore, the execution unit 206c displays the most recent interaction between the user and the agent based on the received usage log (step S618). After that, when the agent receives a request from the user (step S619), the execution unit 206c transmits the received request to the server application of the information terminal 1 using a predetermined API / protocol (step S620).

[0070] The server application of the information terminal 1 executes data processing including inference processing of the AI ​​model in response to the received request (step S621), and transmits a response to the vehicle 2 (step S622). Furthermore, the server application updates the usage log stored in the memory 105 (step S623). The execution unit 206c of the vehicle 2 responds to the user while controlling the expression / representation of the agent based on the agent attribute information in response to the response received from the information terminal 1 (step S624).

[0071] For example, the agent is not limited to an electronic key, but may be managed and automatically activated in association with one or more of the identification ID of the information terminal 1, the International Mobile Equipment Identity (IMEI), the Subscriber Identity Module (SIM), the user's face, iris, retina, fingerprint, palm print, vein pattern, voice, a Personal Identification Number (PIN), a wake-up phrase (wake-up word), a gesture, or the name of the agent. As shown in the example of the figure, the server (server application) that generates the agent's response may transmit data from the most recent interaction to the currently used client based on the user account and synchronize the data, so that communication with the agent continues even when the client changes from the information terminal 1 to the in-vehicle system of the vehicle 2.

[0072] In other words, by logging in (authentication) with a user account from a new client, the server may obtain the latest transactions (e.g., usage logs) for that user's account and display / notify them via the UI of the new client. This can be achieved by continuously updating the usage logs on the server side. This section describes a process in which a user in vehicle 2 continuously interacts with an agent represented / depicted by the in-vehicle system, rather than with information terminal 1. The client app sends a follow-up question from the user, "What time will it start raining?" to the server app via a predetermined API / protocol (step S620), and the server app replies to the client app via a predetermined API / protocol, "I'll get off around 3 o'clock" (step S622). In this process, the server app is running on the execution unit 104a of information terminal 1, and the client app is running on the execution unit 206c of vehicle 2. The execution unit 104a of information terminal 1 and the execution unit 206c of vehicle 2 are connected via short-range wireless communication (wired communication is also acceptable) via their respective communication units.

[0073] 7A and 7B are sequence diagrams showing an example of the process flow for continuing interaction with an agent even when switching clients in the interactive agent system according to this embodiment (summoning the agent that was used immediately before to a new client). In the following explanation, the same process as that shown in FIG. 6 will not be described.

[0074] If the memory 207 of the vehicle 2 does not contain access information for the agent most recently used by the user, or if the memory 207 is inaccessible, the acquisition unit 206a requests access information for the agent most recently used by the user from a server application (the access method is assumed to be known) that manages agent use in the information terminal 1 (step S701). The server application of the information terminal 1 may identify the agent most recently used by the user based on the last use log information for each agent (step S702), and may transmit the access information for the identified agent to the vehicle 2 (step S703).

[0075] Currently, when using an AI model equipped with an LLM and capable of chatting, a typical usage scenario involves a user accessing a server on which the AI ​​model is available from a client's web browser and logging in with a user account, where they can view past usage logs and individually configure settings for response policies and the scope of data referenced for chatting with the AI ​​model. Therefore, currently, when continuing to use an agent while switching clients, as in the present disclosure, the user must manually access the agent's server and log in again, which is time-consuming. A challenge with the current usage scenario is that it is not possible to provide a continuous experience where the user continues to use a specific agent regardless of the client. This challenge is likely to become more pronounced when users want to switch clients when getting in / out of a vehicle 2 or entering / exiting a home or office building.

[0076] This disclosure discloses multiple solutions to this problem. As shown in FIG. 7 , other methods for identifying the agent last used by the user when switching clients include the following: [Variation 1] Implementing Agent Communication Functions on a Client Using an App A user can register one or more agents in an app used by the user on the client, and manage their usage history using a server (app integration server) used by the app. If this app is pre-installed on the information terminal 1, the vehicle 2, and the client used by the user, when the app is launched, the app can log in with a user account if necessary to obtain the usage history and access information of the last used agent stored on the app integration server. Furthermore, the app can also obtain the latest usage log by logging in to the agent. This disclosure also assumes that an agent can be represented as a user's communication partner on the client. Therefore, the app may not only seamlessly communicate with the agent but also have a function for representing the agent. [Variation 2] Realization via a website that aggregates the agents used by the user. It is also possible to provide a web service that allows users to use one or more agents via a one-stop web service that can be accessed by users or clients via a specific URL or API endpoint. In this case, users can log in to this one-stop web service to obtain the agent usage history stored on the server. This allows the web service to identify the agent last used by the user and obtain the latest usage log. [Variation 3] Realization via sharing information about the last agent used between the server app and the electronic key app. If the server app shares the last usage date and time of that agent with the electronic key app (or another app), when the electronic key app performs key authentication with the vehicle 2, it can send access information for the agent last used by the user to the vehicle 2, and the vehicle 2, which receives this information, can record it in memory 207.In this way, it is possible to obtain the access information of the agent that was last used in step S612. Once the access information is known, it can be accessed to obtain the latest usage log, allowing communication with the agent to continue seamlessly.

[0077] The above variants 1 to 3 are merely examples, and there are other methods for the client to identify the agent last used by a user and obtain its latest usage log. For example, one or more users can log in to all of the agent servers used by the users, obtain communication session information and last update date and time information (timestamps), and compare them. The method for identifying the agent last used by a user may be any of the above, an improved version of the above, or another method.

[0078] In the case of the above-described modified example 1, the processing from step S701 to step S703 described in FIG. 7-1 changes as shown in FIG. 7-2. FIG. 7-2 illustrates the changes from FIG. 7-1. The acquisition unit 206a launches an app that expresses / represents an agent installed in the vehicle 2 (step S710). The launched app requests access information for the agent last used by the user from the linked server (the application linking server in FIG. 7-2) (step S711). The server linked to the app identifies the agent last used by the user (step S712). Then, it sends the access information for the identified agent to the client app (step S713). Note that this "server linked to the app" may be accessed by an app installed on any client, so it is desirable for it to be a server that is always accessible via the Internet.

[0079] 8 is a sequence diagram showing an example of a process flow for enhancing interaction between a client and an agent in the interactive agent system according to this embodiment. The acquisition unit 206a of the vehicle 2 periodically (or intermittently, the same applies below) acquires (collects) sensing data (additional data) using the sensor unit 203 (step S801), converts the acquired additional data into output data in a format and type that the agent can handle, and periodically transmits the output data to the information terminal 1 (step S802). Every time the calculation unit 104 of the information terminal 1 receives additional data from the vehicle 2, it updates the additional data stored in the memory 105 based on the received additional data (step S803).

[0080] When the execution unit 206c acquires a request from the user using the agent (step S804), it transmits the acquired request to the information terminal 1 (step S805). When the request is received from the vehicle 2, the calculation unit 104 of the information terminal 1 executes data processing based on the latest added data (step S806) and transmits a response to the vehicle 2 (step S807). Furthermore, the calculation unit 104 of the information terminal 1 updates the usage log stored in the memory 105 (step S808). The execution unit 206c of the vehicle 2 responds to the user using the agent represented by the client in accordance with the response received from the information terminal 1 (step S809).

[0081] When the user makes a disgusted expression, the acquisition unit 206a of the vehicle 2 analyzes the user's facial expression from the video captured by the camera inside the vehicle 2 and transmits the results of the analysis to the information terminal 1 as additional data (step S810). This analysis may be performed periodically (or intermittently) as step S802, or may be performed separately from step S802 each time the agent responds to the user. By performing step S810 each time the agent responds, the vehicle 2 can notify the information terminal 1 of the user's reaction more immediately and at a more appropriate timing, thereby improving the quality of communication and the value of the experience. The calculation unit 104 of the information terminal 1 updates the additional data stored in the memory 105 to the latest version based on the additional data received from the vehicle 2 (step S811). Furthermore, the calculation unit 104 performs data processing related to communication to be conducted with the user based on the latest additional data stored in the memory 105 (step S812). If it determines that further communication is necessary, the calculation unit 104 transmits the response generated in step S812 to the vehicle 2 (step S813). If not, the calculation unit 104 waits for an explicit request from the user. The calculation unit 104 then updates the usage log stored in the memory 105 (step S814). The execution unit 206c of the vehicle 2 responds to the user via an agent in accordance with the response received from the information terminal 1 (step S815). The vehicle 2 and the information terminal 1 similarly exchange requests and responses (steps S816 to S821).

[0082] Here, an example of an interaction between a user and an agent is shown in Figure 8. The user asks the agent, "How far is it to your destination?" (verbal communication) (step S804), and the agent responds, "It will take about an hour and a half" (step S809). Although the user does not respond clearly in words, he shows a facial expression (non-verbal communication) that suggests disgust or that he perceives the situation as undesirable. This is detected by the vehicle's sensors as the user's reaction to the response in step S809, and this reaction is transmitted to the server in step S810. The server, having detected this reaction, makes a suggestion via the agent, "Shall we take a short break?" (step S815). The user's silent nod (non-verbal communication) is detected as the user's reaction, and the agent proposes a specific break suggestion, "How about a coffee shop 3 km down the road?"

[0083] In communication between the user and the agent, even if a clear response from the user (for example, text data obtained by voice recognition of oral utterances and including a clear expression of intent, as in verbal communication) is not obtained, at least one of the following non-verbal communication of the user detected by the sensor unit 101 equipped in the interior of the vehicle 2, such as facial expressions, presence or absence of interjections, head movements indicating affirmation or denial, estimated emotions, biometric information (changes in pupils, fluctuations in heart rate, changes in sweating, etc.), or vehicle status data such as vehicle position, vehicle speed, accelerator and brake operation status, indication of intention to turn right or left (presence or absence of turn signal), indication of intention to make an emergency stop (presence or absence of hazard lights), and the road category on which the vehicle is located (highway, general road, school zone (road with many children), intersection, stop or slow down area, private property (home property, home parking lot), etc.) can be added to the agent's input data as additional data expressed in a predetermined format. This way, even if the user does not or cannot respond clearly verbally, the system can understand the user's intention (e.g., positive or negative), emotions (e.g., joy, anger, sadness, or happiness), and situation (e.g., concentrating on driving at an intersection), and continue the interaction in a way that is appropriate, less irritating, and makes driving safer.

[0084] 9 is a flowchart showing an example of the flow of a process for querying and transmitting multimodal data types input to the interactive agent system according to this embodiment. First, the client side (vehicle 2 or information terminal 1), which acts as a UI for the user, queries the server side (vehicle 2, information terminal 1, or cloud 3) that runs the agent about additional data that the server side can handle, and obtains a list of such data (steps S901 and S902). The client side then identifies, from among the types or formats of additional data that it can handle, additional data of a type or format that can be acquired from a connected sensor (e.g., the sensor unit 101 of the information terminal 1 or the sensor unit 203 of the vehicle 2) (step S903).

[0085] The client side converts the data acquired from the sensor (sensing data) into output data in a format that can be received by the specified server side (step S904). The client side then checks whether there is at least one of the following: additional data due to a status change, additional data due to an update period, or additional data for which an update is requested (step S905). If there is at least one of the following additional data due to a status change, additional data due to an update period, or additional data for which an update is requested (step S905: Yes), the client side transmits the corresponding additional data to the server side (step S906, which corresponds to steps S802, S810, and S816 in FIG. 8). By repeating this process, the agent (data processing for response generation, including inference processing by an AI model) running on the server side can accurately detect or recognize the user's status and changes in the surrounding environment when the status changes, periodically or intermittently, or at an appropriate timing depending on the user's status and reaction.

[0086] Here, for example, the additional data may include raw data of the user's state or reaction acquired from an image sensor inside the vehicle (e.g., an image from an in-vehicle camera), or results expressed by lightweight analysis data showing the results of analyzing the raw data. The additional data may include at least one of the following information measuring the user's state or reaction: intonation of speech, pause before utterance, tone of voice, heart rate (heart rate variability), blood pressure, body temperature, breathing, pupils, sweating, facial expression, gaze, posture, and gestures.

[0087] Furthermore, when the user is inside the vehicle while driving, etc., and the on-board system of vehicle 2 is the client, the additional data may include at least one of the following: vehicle position, vehicle speed, accelerator / brake operation status, indication of intention to turn right or left (whether or not turn signals are on), indication of intention to make an emergency stop (whether or not hazard lights are on), road category on which the vehicle is located (highway, general road, school zone (road with many children), intersection, stop or slow down area, private property (home premises, home parking lot), etc.), sensor recognition information of the environment around the vehicle, information about the currently set route, and autonomous driving response level information obtained from the sensor unit 203.

[0088] In this way, depending on the client's detection capabilities, additional data information can be sent periodically or in synchronization with requests, making it possible to make the information multimodal. If the server side has a wealth of input information for the agent, it is expected that the agent's responses will be appropriate and timely, taking into account the situation of the user and vehicle. Because the in-vehicle system has the ability to sense various conditions related to the user and vehicle 2, when the agent is executed with the in-vehicle system as the client and the information terminal 1, vehicle 2, or cloud 3 as the server, it is expected that the agent's responses will be more natural, more convenient, and safer.

[0089] For example, if vehicle 2 is currently driving, the quality, quantity, and timing of communication from the agent can be controlled to ensure safe driving. Another use case is when the agent detects or estimates that the user's concentration is declining or that they are yawning after a long drive, suggesting that the user stop at a rest stop along the route. Furthermore, if there are changes in the user's biometric information, the agent can ask if they are feeling unwell, guide vehicle 2 to a suitable parking spot, or, if vehicle 2 is capable of autonomous driving, switch to autonomous driving mode and move the vehicle to a suitable parking spot or to the nearest hospital. Furthermore, if the agent can detect from data on the user's status or vehicle status that the vehicle is not currently driving, the agent's response can include more detailed information or use visual information to aid understanding. Such capabilities would be difficult or insufficient if only text recognized from linguistic utterances were sent to the agent; they can only be achieved through multimodal input, as described above. Furthermore, as described above, by arbitrating in advance which multimodal data will be shared and sharing it instantly when needed in a lightweight data format, the processing load on the inter-systems can be reduced, leading to cost reductions and faster response times.

[0090] 10 is a diagram illustrating an example of the structure of additional data converted from multimodal raw data into text data in the interactive agent system according to this embodiment. Raw data acquired by a wide variety of sensors, which represents state quantities that fluctuate over time, is large in size and places a heavy burden on the network and calculations. Therefore, it is conceivable to periodically transmit the additional data (text data) representing the state or level of each type of raw data to the agent (server side) (e.g., in JSON format). By transmitting the additional data as independent text data (additional data) periodically, the server side can easily update the additional data for each type of modal to the latest one and refer to it when generating a response, even if any number of modals (additional data) are updated at any time.

[0091] The text shown as an example of "Vehicle Status 1" is an example of a vehicle status detected by an in-vehicle system (CDC: also known as Cockpit Domain Controller, etc.). Here, it is stated that the in-vehicle system (CDC) detected that the vehicle was in operation, the engine was running, and the vehicle speed was 60 km / h at the measurement time of 11:46:50 on April 25, 2024. The in-vehicle system, as a client, periodically or intermittently transmits this additional vehicle status data to a server that generates responses through inference processing of an AI model.

[0092] Similarly, examples of texts describing the user's state from video images of the user captured by an image sensor inside the vehicle are shown as "User State 2" and "User State 3." "User State 2" describes that, at the measurement time of 11:47:00 AM on April 25, 2024, the user's emotion analysis results (emotion_scores) were analyzed using the in-cabin monitor (ICM) on the dashboard, showing that 80% of the user's emotions were negative, indicating a negative reaction. This may be for the user interacting with the agent, for the user sitting in the driver's seat, or for all passengers in the vehicle. Information identifying the user being observed may also be included (not shown). Similarly, "User State 3" shows the results of analyzing the user's state based on gestures (gesture scores), indicating that the user is responding positively.

[0093] Furthermore, the periodic transmission of this additional data may be realized by including the additional data in an HTTPS request (e.g., a POST request) between the client and the server, by scheduling the HTTPS request to be sent periodically on the client side using JavaScript (registered trademark) or the like, by establishing a WebSocket connection between the client and the server that enables two-way, real-time communication and transmitting the additional data through that connection, or by an application installed on the client (which may be an application that controls the expression / representation of an agent) notifying the server using a predetermined API. In the present disclosure, the mechanism for periodically transmitting such multimodal additional data from the client to the server is not limited in its implementation form as long as it is realized.

[0094] This diagram also shows, using Figure 8 as an example, the relationship between the time relationship of the interaction between the user and the agent and the additional data that is periodically updated at that time from the client (vehicle 2) to the server. As shown here, when the user asks the agent, "How far is it to the destination?", vehicle 2 uses the sensor unit 203 to acquire information about the user's status "User Status 1," information about the vehicle's status "Vehicle Status 1," and information about the situation around the vehicle "Vehicle Surroundings Situation 1," which are all processed into predetermined text data and sent to the server.

[0095] When the agent responds, "It will take about an hour and a half," vehicle 2 acquires information about the user's emotions, "User Status 2," and information about the user's biological reactions, "User Biological Information 1," and transmits these to the server in the same way. Here, "User Status 2" indicates that the user has reacted negatively through non-verbal communication (emotion estimation by facial expressions or other means), as described above.

[0096] Based on this additional data, the agent detects that the user is tired from driving and suggests taking a break, asking "Would you like to take a break?" to provide the user with a more comfortable travel experience. Here, vehicle 2 acquires information about the user's state at the time the agent suggested the break, "User State 3," and information about the vehicle's state, "Vehicle State 2," and sends these to the server. Here, "User State 3" indicates that the user has responded positively through non-verbal communication (gestures), as described above.

[0097] If the agent detects a positive response from the user from the additional data, it will present specific break suggestions to users who responded positively to a break suggestion such as "Is there a coffee shop 3 km away?" The user's state at this time is also recorded as "User State 4," and the vehicle interior environment is also recorded as "Vehicle Interior Temperature and Humidity 1," which are then converted into text data and sent to the server.

[0098] Alternatively, the user's facial expressions may be captured by an in-vehicle camera (ICM) and the video data may be sent to the server either directly or after partial trimming and video processing. In this case, it is important to note that this increases the processing load on the network and server. To reduce the processing load of video and audio analysis on the server, the time and content of the agent's communication with the user may be embedded as metadata in the video, or the data may be sent to the server via a separate data file or API. This may also include the agent's appearance, facial expressions, and voice characteristics on the client side. Because the server may not have a complete picture of how the agent communicated with the user on the client side, information about the agent's actual communication style may be used when analyzing the user's state and reactions. This is expected to enable accurate understanding and analysis of communication with the user.

[0099] FIG. 11 is a diagram illustrating an example of the structure of data converted from multimodal raw data into text data in the interactive agent system according to this embodiment. It is conceivable that an agent (a function that performs data processing, including inference processing of an AI model) running on a client that was not involved in the interaction with the user may make a suggestion to the user via an agent (an agent that generates a response to the user on the server). In FIG. 11, a vehicle agent that was not involved in the interaction with the user makes a suggestion, "Vehicle Agent Proposal 1," to the agent based on the previous interaction between the user and the agent.

[0100] "Vehicle Agent Proposal 1" suggests a coffee shop 3 km away as a rest stop. The advantages of the proposal include the coffee shop being famous for its latte art and having a quick charger installed. It also shows that this proposal was made at 11:47:20 AM on April 25, 2024, by a vehicle agent, an AI agent built into the vehicle system that can be accessed using access information specified by a URL (e.g., https: / / 192.168.1.100:443 / VehicleAgent). Proposals from agents other than those directly interacting with the user (such as the vehicle agent in Figures 10 and 11) are sent to the agent as irregular additional data when the opportunity and conditions are right.

[0101] When the agent receives this additional data "Vehicle Agent Proposal 1" from the client vehicle 2, it determines whether to make the proposal to the user. If the proposal is not made, the vehicle agent's proposal is rejected and is not used in interactions with the user. If a proposal is made, as shown in FIG. 11 , the agent proposes a specific break idea to the user, such as "Is there a coffee shop 3 km away?". Here, to simplify the interaction between the user and the agent, the proposal is treated as if it were from the agent, but the present disclosure is not limited to this. The agent may also disclose the name of the agent that made the proposal and convey the contents of the proposal to the user, noting that it is from that agent.

[0102] One of the key points of the disclosed technology is that when there are multiple agents available to a user (functions that perform data processing, including inference processing of AI models), it is possible to smoothly and lightly implement suggestions and information from an agent different from the agent that is communicating with the user, without confusing the user. When various specialized agents appear, input unique to the specialized agent can be sent to the partner agent in the form of a proposed text via a partner agent that communicates directly with the user and has a deep understanding of the specialized agent, and if it is determined to be useful, the partner agent can convey the information to the user in an easy-to-understand format and at a granularity level.

[0103] In the above description, the proposal from the vehicle agent, "Vehicle Agent Proposal 1," is sent to the server-side agent for processing, but the present disclosure is not limited to this. The client application may directly and self-containedly handle the proposal from the vehicle agent and make the proposal from the vehicle agent using an agent expressed / represented on the client, or the vehicle agent may appear separately as a separate communication partner and make the proposal to the user from that vehicle agent, or a message (nudge) indicating that there is a coffee shop suitable for a break 3 km away may be notified to the user via a UI screen, without mentioning the source of the proposal.

[0104] Fig. 12 is a diagram for explaining an example of a dialogue agent system according to this embodiment. Specifically, Fig. 12 is a diagram showing how the agent interacting with the user in Fig. 11 (a partner-type agent dedicated to the user) is connected to other devices or other agents.

[0105] The user interacts with the agent via a client terminal (or an application running there; the same applies below) that handles UI operations. In the example of FIG. 11 , the client terminal is an in-vehicle system. As described in FIG. 11 , proposals from the vehicle agent, which is an agent built into the in-vehicle system, to the agent are sent from the client terminal built-in agent (vehicle agent) to the client terminal (application), and are communicated from the client terminal to the partner-type agent (an agent running on the information terminal 1 or cloud 3 server that interacts with the user) via a server-client API. Management by a communication session may be effective between the server and the client, and the connection may be via an HTTPS protocol, a WebSocket connection, or the like. Furthermore, if the client terminal built-in agent knows access information to the partner-type agent, it may directly connect to the partner-type agent to provide information.

[0106] On the other hand, the agent that interacts with the user also connects to the Internet and acts as a bridge, utilizing external knowledge and services to notify or provide them to the user. This may be a wide variety of publicly available information on the Internet that can be accessed via an Internet connection, or a service provider site that offers specific services, such as an e-commerce site or search site. Furthermore, agents that are highly capable of executing specific tasks and can receive those services via APIs, etc., are envisioned; these are referred to as specialized agents in Figure 12.

[0107] A specialized agent may, for example, 2These agents may be free agents that suggest efficient routes, paid agents provided by entertainment companies that provide conversation companions for long-distance drives, or agents that specialize in carrying out specific tasks or areas of knowledge. Agents that interact with users can utilize the resources of information and services on the Internet to enrich interactions with users and increase customer value.

[0108] FIG. 13 is a sequence diagram illustrating an example of a dialogue agent system according to this embodiment. In this diagram, a user is connected via a client application (body) that represents the agent's appearance and voice to a server application (brain) that generates responses from a partner-type agent that communicates directly with the user. The partner-type agent obtains answers from a trusted specialized agent in a specific field, customizes the answers for the user, and responds. The processing will be explained below with reference to the diagram.

[0109] The client application running on the in-vehicle system of vehicle 2 communicates with the user to obtain a user request (step S1301) and sends it to a server that executes data processing including inference processing of an AI model (step S1302). In this figure, the server is information terminal 1, and in reality, the request is sent to a partner-type agent running on information terminal 1. Here, the client application obtains a question from the user, such as "How far is it to the destination?", and requests a response from the partner-type agent.

[0110] The partner-type agent that receives this request searches for candidates for specialized-type agents that are likely to be able to generate answers that are more accurate or more useful to the user than the partner-type agent, and if one or more specialized-type agents are found, it decides to use the answer from the specialized-type agent (step S1303).

[0111] The partner agent notifies the user of this (step S1304). In this example, the user is notified, "I'll ask the navigation agent." The client application that receives this notification uses a GUI or VUI to represent the agent's response and responds to the user (step S1305).

[0112] The partner-type agent sends an access request to a specialized-type agent with which it can communicate via the network (step S1306). The specialized-type agent, upon receiving this request, grants access (step S1307). At this time, the conditions of use may also be determined. The specialized-type agent responds by notifying the access permission and indicating the conditions of use of the specialized-type agent, if any (step S1308).

[0113] When the partner-type agent receives this, if the terms of use are included, it checks the terms of use and accepts them if there are no problems (step S1309). If it does not accept them, it requests access from another specialized-type agent and performs the above processing. If necessary, the partner-type agent provides the specialized-type agent with information in accordance with the terms of use, or provides monetary or economically valuable points or cryptocurrency (step S1310).

[0114] The partner agent requests the specialized agent to estimate the time required to reach the destination, along with the current vehicle position, the currently set route information, etc. (step S1311). This request is an instruction to the specialized agent that is generated autonomously by the partner agent, and may take the form of a prompt describing the information processing content expected of the specialized agent.

[0115] The specialized agent, upon receiving this, estimates the required time to the destination based on the latest traffic information database (step S1312) and sends the result to the partner agent (step S1313). The partner agent, upon receiving this, processes the estimated time obtained from the specialized agent for the user based on at least one of the user attribute information, user status, and vehicle status (step S1314), such as summarizing or supplementing it, and generates a response to the user and sends it to the client application (step S1315). In this case, the response to the user is "it's congested, so it will take about an hour and a half," based on the situation and the user's preferences. Alternatively, the partner agent may send the specialized agent's estimate directly to the client application without processing the response in step S1314.

[0116] The client application that receives the response answers the user's question based on the received response (step S1316). At the same time, the client application acquires the user's state and reaction at the time when the response is notified via one or more sensors connected to the in-vehicle system (step S1317). This may be part of periodic or intermittent detection of the user's state.

[0117] The client application transmits the sensing data regarding the user's state and reaction detected via the sensors as raw data without processing, or analyzes and converts it into the aforementioned JSON format text data classified into categories and levels, and transmits it to the partner-type agent as the aforementioned additional data (step S1318). Here, the user's reaction may be clearly expressed in words (e.g., "OK," "Got it," "Huh?", "It's taking that long," etc.), or it may be the emotion analysis results obtained from the user's facial expressions and posture changes through the aforementioned ICM video analysis (e.g., information about the strongest detected emotion, information about the degree to which each emotion is detected, etc.).

[0118] Upon receiving the user's reaction indicating dislike, the partner-type agent creates and requests the specialized-type agent to propose a new route based on at least one additional data item regarding the user, vehicle condition, and external environment (in this case, traffic conditions along the route, weather, event information near the route, etc.), and / or user attribute information including information about locations that may be of interest to the user and ways to rest (steps S1319, S1320).

[0119] The specialized agent then generates a response (new route proposal) based on the received instructions (step S1321). The specialized agent, which receives advertising and promotional expenses from multiple commercial facilities and services, may adjust the appeal and priority of the facilities and services included in the new proposal based on the monetary value of the remaining advertising and promotional expenses. Alternatively, the specialized agent may enter into a contract to charge a fee for inclusion in the proposal, and the specialized agent may charge the facilities and service providers included in the proposal it generated for their advertising (step S1322). This approach allows for market-based advertising and promotion, and allows users to use the specialized agent for free or at a very low cost. Furthermore, instead of advertising within a limited, narrow area (e.g., prefecture-level), the specialized agent can match facilities and services available along the currently set route with pinpoint accuracy, thereby introducing facilities and services that are highly accessible to the user. This is likely to be beneficial for both the user and the service provider.

[0120] The specialized agent that generated the new route proposal sends the proposal to the partner agent (step S1323). Upon receiving the proposal, the partner agent, as described above, processes the proposal for the user by adding at least one of the additional data related to user attribute information, user status, and vehicle status (step S1324) and sends it to the client application (step S1325). In this example, based on the fact that the vehicle's remaining battery power is low and that the user frequently stops at coffee shops, the specialized agent responds with, "You'll also need to charge your car. How about taking a break and charging at a coffee shop 3 km away?" (step S1326). The client application that receives this response notifies the user via the agent's appearance (GUI) or voice (VUI). At the same time, as described above, the partner agent detects the user's status and reaction to receive the response, enabling multimodal communication (step S1327).

[0121] Fig. 14 is a diagram for explaining an example of a dialogue agent system according to this embodiment. In Fig. 10 and Fig. 11, communication progressed through interactions between the user and the agent, but Fig. 14 explains an expression form in which the agent that was interacting with the user calls the vehicle agent in Fig. 11, causes it to participate in the interaction, and then causes it to withdraw from the interaction once the task is completed.

[0122] If, during an interaction with a user, the agent determines that there is a candidate agent who can provide a better response than the agent itself, it will call that agent who is able to communicate via the network. In this case, the agent that was conversing with the user determines that it would be most appropriate to ask the vehicle agent about the time and distance to the destination, and calls the vehicle agent into the conversation. When calling the vehicle agent, the agent asks the vehicle agent a question about the time required to reach the destination, and the vehicle agent responds to the question based on the current situation and provides a response to the user and the agent.

[0123] Apart from the original communication between the user and the agent, a separate server-client type communication may be established between the agent and the vehicle agent. For example, the communication may be performed using the HTTPS protocol or the API of the vehicle agent. Even in this case, based on the response obtained from the vehicle agent, the client (application) may use a UI unit to present the response to the user as if the vehicle agent were to say, "It's congested, so it will take about an hour and a half."

[0124] The agent can individually connect with the vehicle agent (another agent other than itself) to obtain information necessary for interaction with the user, and then use or quote the answer in the interaction with the user while claiming it is the answer from the vehicle agent. This is thought to increase the possibility of the agent interacting appropriately and to expand the possibilities for story development.

[0125] Once the necessary information has been obtained from the vehicle agent and confirmed with the user, the agent disconnects the vehicle agent from the interaction with the user. In Figure 13, the communication is disconnected as if the agent said "Thank you, see you later" to the vehicle agent. From then on, the conversation returns to the conventional server-client type connection with the user's client, and continues.

[0126] The interaction between this agent and the vehicle agent can be realized as an individual connection between the partner-type agent and the agent built into the client terminal in FIG. 12. If it is determined that it would be better to ask the specialized agent rather than the vehicle agent, an individual connection may be established between the partner-type agent in FIG. 12 and the specialized agent on the Internet to obtain appropriate answers or recommendations and share them with the user. In this case, it is possible to have the user experience a three-way interaction including the specialized agent by operating the UI unit 204 of the client, or to have the user experience a two-way interaction between the partner-type agent and the user by operating the UI unit 204, taking into account information obtained from the specialized agent, without directly showing the interaction between the specialized agent and the partner-type agent to the user.

[0127] As shown in this diagram, when an agent calls a specialized agent, it is desirable to have the specialized agent appear in the conversation with the user by asking the specialized agent who can respond more accurately to the current topic, and then have the specialized agent exit the conversation after expressing gratitude once the topic has been addressed. This makes it easy for the user to start and end communication using a specialized agent.

[0128] Furthermore, the specialized agent does not necessarily have to be displayed in the same form as the agent on the client's UI unit 204; the specialized agent's name and appearance may be displayed in a predetermined frame, like a participant participating in an online conference. In this case, the agent can act as a moderator or be given a certain degree of progress management authority to promote communication with the user, making it easier for the user to understand the progress of the conversation. When inviting / exiting a specialized agent, the agent can do so in an easy-to-understand manner, allowing the user to utilize the specialized agent's knowledge and skills without any additional effort.

[0129] FIG. 15 is a diagram illustrating an example of multimodal data in the interactive agent system according to this embodiment. Specifically, it shows data types related to the state of the user and the vehicle 2 sensed on the client side. By inputting these data as multimodal data to an agent, which is an AI model that generates a response, it is expected that the agent's response will be appropriate, accurate, and quick in accordance with the user's non-verbal communication (facial expressions, gestures, etc.), the state of the vehicle 2, and changes in the user's surrounding environment, even if an explicit verbal response (such as speech or a text version of speech) is not received from the user.

[0130] A wide variety of time-series data can be obtained from various sensors, but sending this data to the server where the agent is running can put a strain on the communication bandwidth and increase the load on the server's data processing, resulting in delays and increased costs.For this reason, it is possible to treat the data from each sensor as individual, independent additional data, such as small amounts of text data written in a specified format (for example, data in JSON format) that indicates the status or level on the sensing client side, and update it periodically.

[0131] The client manages this multimodal additional data along with its type and measurement date and time information, and generates responses based on the latest information. It may stop referencing old additional data that has passed a certain amount of time, or additional data that is no longer updated because the client connection has been lost, or delete that data type from the latest information dataset. This allows the agent to generate responses based on additional data that is updated at an appropriate interval (i.e., the latest information and situation about the user and the user's surrounding environment).

[0132] The following are the types of sensors and systems used for sensing on the client side and the types of data that can be obtained from them. This data can be used to read a wide variety of information about users, vehicles, and the situation around the vehicle.

[0133] Sensors worn by users include smart watches, smart rings, smart glasses, etc. These are accessories that are embedded with sensors that can measure the user's heart rate (heart rate, heart rate variability), electrocardiogram, blood pressure, number of steps, calories burned, calories ingested, blood glucose level, sleep, stress, electrodermal activity, breathing, voice analysis, electroencephalogram, electromyogram, skin gas, blood oxygen saturation, chewing, swallowing, gaze, blinking, and other biological responses.

[0134] The information devices carried by users also have many sensors that can measure information about users, such as the information they search and view online, the information they send or receive via email, chat, or social media, their current location, or its history.

[0135] In-vehicle sensors include the ICM, seats, and seatbelts, which can measure the user's facial expression, posture, responses, gaze, brainwaves, surface temperature, seatbelt status, passenger identification and identification, and object identification and identification.

[0136] Vehicle control sensors include data sent to a network (CAN) by each ECU in the vehicle, data used by ADAS (Advanced Driver Assistance Systems), etc. For example, they can measure vehicle conditions such as the driving, parking, and stopping state, whether the engine or motor power source is on or off, the remaining gasoline level, the remaining battery level, the application state and level of an autonomous driving or driver assistance system, vehicle speed, operation of the steering wheel, accelerator, and brake, road classification (highway, general road, inside or outside an intersection, inside or outside a slow-down area, inside private property (inside a home), etc.), driving conditions such as whether or not a turn signal is on, whether or not a hazard lamp is on, route information, and navigation system setting information such as traffic congestion information.

[0137] Sensors around a vehicle include LiDAR and radar, which can measure surrounding conditions such as sensing data of the environment around the vehicle, its position relative to roads, intersections, and traffic infrastructure, the types of objects around the vehicle and their positions, and lane information the vehicle is traveling in.

[0138] The client processes at least one of these pieces of data into a predetermined format periodically, intermittently, or at a predetermined timing, and sends it to the server, which performs inference processing on the AI ​​model. This allows the server to generate more accurate responses when generating communications.

[0139] In this way, according to the interactive agent system of this embodiment, when the user gets into the vehicle 2, it is possible to display and continue to use an agent that is linked to the user or the electronic key being used, or an agent that the user was using on the information terminal 1 immediately before the user got into the vehicle 2. Furthermore, by using a mechanism similar to that used when getting into the vehicle 2, it is possible to display and continue to use the agent that the user was using on the vehicle 2 immediately before the user got out of the vehicle 2 on the information terminal 1.

[0140] In this embodiment, an example has been described in which an agent capable of dialogue linked to a user is realized in a vehicle 2, but the present invention is not limited to this. Even in spaces such as a home, office, or store, it is possible to obtain an ID that identifies an information terminal 1 such as a smartphone, obtain agent access information from an application installed on the information terminal 1, or obtain agent access information linked to the user's entry authentication information for their home, office, store, etc., and thereby summon a partner-type agent that the user has set or the agent that the user was using immediately before onto the computer system of that space when the user enters that space.

[0141] For example, just as when getting into vehicle 2, the conversation with the partner-type agent can be seamlessly switched from information terminal 1 to the vehicle's in-vehicle system (client terminal) and continued, when entering home or office, the conversation with the partner-type agent can be seamlessly switched from information terminal 1 to a client terminal provided in the home or office and continued.

[0142] The programs executed by the information terminal 1 and the vehicle 2 of this embodiment are provided by being pre-installed in a ROM (Read Only Memory) or the like. The programs executed by the information terminal 1 and the vehicle 2 of this embodiment may be provided by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM, a flexible disk (FD), a CD-R, a DVD (Digital Versatile Disk), or an SD card.

[0143] Furthermore, the programs executed by the information terminal 1 and the vehicle 2 of this embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the programs executed by the information terminal 1 and the vehicle 2 of this embodiment may be provided or distributed via a network such as the Internet.

[0144] In addition, in the present embodiment, in the information terminal 1, an example of a processor (calculation unit 104) such as a CPU (Central Processing Unit) uses a RAM (Random Access Memory) or the like as a working area and executes various programs stored in a ROM (memory 105) or the RAM, thereby realizing the communication unit 104b and the execution unit 104a.

[0145] In addition, in the present embodiment, in the in-vehicle system, an example of a processor (calculation unit 206) such as a CPU (Central Processing Unit) uses a RAM (Random Access Memory) or the like as a working area and executes various programs stored in a ROM (memory 207) or the RAM, thereby realizing an acquisition unit 206a, a setting unit 206b, and an execution unit 206c.

[0146] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents.

[0147] 16 is a diagram showing an example of the functional configuration of the interactive agent system according to this embodiment. Here, the processing executed by the aforementioned acquisition unit 206a, setting unit 206b, and execution unit 206c (see FIG. 5) has, in block units, a communication unit 1601, an identification unit 1602, a summoning unit 1603, and an agent providing unit 1604. Other embodiments will be described below in [Item 1] to [Item 7].

[0148] [Item 1] An in-vehicle device that communicates with a cloud and a user terminal via a network, wherein the in-vehicle device and the cloud summon one or more AI agents, and at least one of the one or more AI agents is a partner-type AI agent that accesses information of a corresponding user, the in-vehicle device having: a communication unit 1601; a summoning unit 1603 that communicates with the cloud and summons the AI ​​agent; an agent providing unit 1604 that gives a prompt from the user to the summoned AI agent and obtains a response from the AI ​​agent; and an identification unit 1602 that identifies a vehicle user who uses the vehicle, and the summoning unit 1603 summons a partner-type AI agent corresponding to the identified vehicle user.

[0149] [Item 2] The in-vehicle device according to the above [Item 1], wherein the identification unit 1602 corresponding to the owner of the user terminal identifies the user of the user terminal with which the communication unit 1601 communicates as the vehicle user.

[0150] [Item 3] In-vehicle equipment according to the above [Item 2], which corresponds to the owner of a user terminal that possesses a smart key, the user terminal possesses at least a smart key that permits entry into the vehicle compartment, driving the vehicle, and / or use of on-vehicle equipment, and the identification unit 1602 identifies the user of the user terminal that possesses the smart key as a vehicle user.

[0151] [Item 4] Supports partner-type AI agents of user terminals The user terminal and the cloud summon one or more AI agents, at least one of which is a partner-type AI agent that accesses information about the corresponding user, and the identification unit 1602 identifies the user corresponding to the partner-type AI agent summoned by the user terminal with which the communication unit 1601 communicates as a vehicle user, in an in-vehicle device described in [Item 2] above.

[0152] [Item 5] The in-vehicle device according to [Item 1] above, which corresponds to the owner of the vehicle wallet, further includes a wallet unit that performs settlement processing between external equipment and the wallet owner, an identification unit 1602 that identifies the wallet owner, and a summoning unit 1603 that summons a partner-type AI agent that corresponds to the wallet owner identified by the identification unit 1602.

[0153] [Item 6] The in-vehicle device according to [Item 2] above, wherein the summoning unit 1603 notifies the summoning AI agent of information on whether the vehicle user is in the driver's seat or the passenger's seat.

[0154] [Item 7] The in-vehicle device according to [Item 1] above, wherein the summoning unit 1603 notifies the AI ​​agent of the vehicle's location information, and the location information indicates at least one of the home, a parking lot, and a road.

[0155] [Item 8] A method corresponding to each of [Item 1] to [Item 7] above.

[0156] REFERENCE SIGNS LIST 1 Information terminal 2 Vehicle 3 Cloud 101, 203 Sensor unit 103, 204 UI ​​unit 104, 206, 303 Calculation unit 104a Execution unit 105, 207, 302 Memory 104b, 106, 208, 301 Communication unit 206a Acquisition unit 206b Setting unit 206c Execution unit

Claims

1. An information processing method executed by a first computer mounted on a vehicle in an interactive agent system capable of interacting with a user, wherein the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, and then acquires access information for accessing an AI agent used by the user from the communication terminal, acquires agent attribute information for representing the AI ​​agent's character from an agent database capable of communicating with the first computer based on the access information, sets at least one of a GUI and a VUI representing the AI ​​agent's character on the first computer based on the agent attribute information, and performs interactive processing with the user using at least one of the GUI and the VUI and generation processing by the AI ​​agent while connecting to a second computer on which the AI ​​agent is implemented based on the access information.

2. The information processing method according to claim 1, wherein the GUI is displayed on a display of the first computer, and the VUI is input and output via a speaker and a microphone of the first computer.

3. The information processing method according to claim 1, wherein said agent database is located in said second computer or a third computer with which said first computer can communicate via a network.

4. The information processing method of claim 1, wherein the first computer is an in-vehicle system, the second computer is the communication terminal, and the access information includes an address for accessing the AI ​​agent in the communication terminal.

5. The information processing method of claim 1, wherein the first computer is an in-vehicle system, the second computer is a server capable of communicating with the in-vehicle system via a network, and the access information includes an address for accessing the AI ​​agent in the server via the network.

6. The information processing method according to claim 1, wherein in the interactive processing with the user inside the vehicle, a usage log containing the content of the conversation between the AI ​​agent and the user before the user uses the vehicle is referenced.

7. The information processing method of claim 1, wherein the first computer is an in-vehicle system, the access information includes an address for accessing the AI ​​agent, the address being a global address for accessing the AI ​​agent implemented on a server that can communicate with the first computer via the Internet, or a local address for accessing the AI ​​agent implemented on the communication terminal that can communicate with the first computer without going through the Internet, and the connection to the second computer is performed via a common API regardless of whether the address is the global address or the local address.

8. The information processing method of claim 1, wherein the AI ​​agent is the AI ​​agent last used by the user, selected from one or more AI agents available to the user based on a usage log stored in the communication terminal.

9. The information processing method described in claim 1, wherein the acquisition of the access information is performed when it is determined based on the authentication information acquired from the communication terminal that information regarding an AI agent available to the user is not registered in the memory of the first computer.

10. The information processing method described in claim 1, wherein the authentication information includes information of an electronic key for the user to use the vehicle, and the AI ​​agent is an AI agent selected from one or more AI agents associated with the electronic key in the communication terminal.

11. The information processing method of claim 1, wherein the AI ​​agent is selected from the one or more AI agents by the user selecting the electronic key from a plurality of electronic keys available to the user via a UI of the communication terminal.

12. The information processing method of claim 11, wherein the plurality of electronic keys include a company electronic key associated with the company to which the user belongs and a personal electronic key of the user, and when the selected electronic key is the company electronic key, a company agent managed by the company is selected as the AI ​​agent, and when the selected electronic key is the personal electronic key, a private agent of the user is selected as the AI ​​agent.

13. A computer mounted on a vehicle, comprising: a processor; and a memory storing a program for causing the processor to execute the information processing method described in any one of claims 1 to 12 as the first computer.

14. A program for causing the first computer to execute the information processing method according to any one of claims 1 to 12.

15. An information processing method for an interactive agent system capable of interacting with a user, executed on a communications terminal capable of communicating with a first computer mounted on a vehicle and implementing an AI agent used by the user, the information processing method comprising the steps of: sending authentication information for the user to use the vehicle to the first computer; after it is determined that the authentication information is valid, sending access information for accessing the AI ​​agent to the first computer; sending agent attribute information for representing the character of the AI ​​agent to the first computer, causing the first computer to set at least one of a GUI and a VUI based on the agent attribute information; and, while connected to the first computer, executing interactive processing with the user using at least one of the GUI and the VUI and generation processing by the AI ​​agent.

16. The information processing method of claim 15, wherein the first computer comprises a display for displaying the GUI, and a speaker and microphone for inputting and outputting the VUI, and the communication terminal comprises an AI processor for causing the AI ​​agent to execute the generation process, and a memory for storing the access information including the address of the AI ​​agent, and the agent attribute information relating to the AI ​​agent.

17. The information processing method described in claim 15, wherein the communication terminal has a memory that stores a usage log of the AI ​​agent by the user, and in the interactive processing with the user inside the vehicle, the usage log including the content of the conversation between the AI ​​agent and the user before the user used the vehicle is referenced.

18. The information processing method described in claim 15, wherein the communication terminal includes a memory that stores a usage log regarding one or more AI agents available to the user, and after it is determined that the authentication information is valid, the communication terminal selects, based on the usage log, one of the one or more AI agents that the user last used as the AI ​​agent.

19. The information processing method of claim 15, wherein the communication terminal manages information regarding a plurality of electronic keys available to the user, accepts an input operation from the user regarding which of the plurality of electronic keys to select, selects the AI ​​agent associated with the electronic key according to the electronic key selected by the user, and, after determining that the authentication information is valid, transmits the access information for accessing the selected AI agent to the first computer.

20. The information processing method of claim 19, wherein the plurality of electronic keys include a company electronic key associated with the company to which the user belongs and a personal electronic key of the user, and if the selected electronic key is the company electronic key, a company agent managed by the company is selected as the AI ​​agent, and if the selected electronic key is the personal electronic key, a private agent of the user is selected as the AI ​​agent.

21. A communication terminal comprising: a processor; and a memory storing a program for causing the processor to execute the information processing method according to any one of claims 15 to 20.

22. A program for causing the communication terminal to execute the information processing method according to any one of claims 15 to 20.

23. A dialogue agent system capable of dialogue with a user, comprising: a first computer mounted on a vehicle; and a second computer having an AI agent implemented therein, wherein the first computer includes a processor and a memory storing a program for causing the processor to execute predetermined information processing, and the predetermined information processing includes: after the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, acquiring access information for accessing the AI ​​agent used by the user from the communication terminal; acquiring agent attribute information for representing the AI ​​agent's character from an agent database capable of communicating with the first computer based on the access information; setting at least one of a GUI and a VUI representing the AI ​​agent's character in the first computer based on the agent attribute information; and executing dialogue processing with the user using at least one of the GUI and the VUI and generation processing by the AI ​​agent while connecting to a second computer having a processor having the AI ​​agent implemented therein based on the access information. Dialogue agent system.

24. An information processing method executed by a first computer installed in a vehicle in an interactive agent system capable of interacting with a user, wherein the first computer or an authentication computer capable of communicating with the first computer acquires authentication information from the user's communication terminal and determines that the authentication information is valid, and then establishes a connection to a communication terminal equipped with a multimodal AI agent; converts sensing data acquired via one or more sensors installed in the vehicle while the user is using the vehicle, for each unit of data acquired over a predetermined period, into output data in a format and type that can be handled by the AI ​​agent; periodically or intermittently transmits the output data to the AI ​​agent in the communication terminal, and causes the AI ​​agent to perform interactive processing reflecting the latest output data.

25. The information processing method according to claim 24, wherein, after determining that the authentication information is valid, data format information indicating the data format and data type that the AI ​​agent can handle is obtained from the communication terminal, and the conversion to the output data is performed in accordance with the data format information.

26. The information processing method of claim 24, wherein the sensing data includes at least one of the user's facial expressions, gestures, emotions, and biometric information, the interior environment of the vehicle's cabin, the surrounding conditions of the vehicle, and the driving state of the vehicle, and the output data includes the type of data from which the output data was extracted, an identification code for identifying the state or event of the vehicle or the user, and a timestamp indicating the time when the state or event occurred.

27. The information processing method described in claim 24, wherein the AI ​​agent changes the timing or amount of information to be presented to the user when it determines, based on the output data from the first computer, that the user is in a situation where they should concentrate on driving the vehicle.

28. An information processing method as described in claim 24, wherein, when the AI ​​agent infers based on the output data from the first computer that the user has indicated an intention to respond to the AI ​​agent, and when the AI ​​agent is unable to obtain a response from the user via the VUI of the AI ​​agent, the AI ​​agent performs interactive processing in accordance with the inference based on the output data.

29. The information processing method of claim 25, wherein the first computer is equipped with a processor on which an in-vehicle AI agent different from the AI ​​agent is implemented, and the first computer causes the in-vehicle AI agent to generate request data indicating constraints required for the dialogue processing performed by the AI ​​agent based on the output data, and sends the request data in addition to the output data to the AI ​​agent in the communication terminal, and causes the AI ​​agent to perform the dialogue processing based on the output data and the request data.

30. A computer mounted on a vehicle, comprising: a processor; and a memory storing a program for causing the processor to execute the information processing method described in any one of claims 24 to 29 as the first computer.

31. A program for causing the first computer to execute the information processing method set forth in any one of claims 24 to 29.

Citation Information

Patent Citations

  • Electronic equipment operation system using character display and electronic apparatuses

    JP2006154926A

  • Electronic equipment, assistant display method, assistant display program, and electronic equipment system

    JP2006285416A

  • Agent system and computer program

    JP2021144086A

  • Wireless connection establishment method, apparatus, device, and storage medium

    JP2022003772A

  • On-vehicle information device and linking method with mobile terminal

    WO2020026402A1