Information processing system, information processing method, and information processing program

JP2026131525APending Publication Date: 2026-08-14SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2026-08-14

Smart Images

  • Figure 2026131525000001_ABST
    Figure 2026131525000001_ABST
Patent Text Reader

Abstract

We propose an information processing system, information processing method, and information processing program that can respond according to the surrounding environment and the state of a moving object. [Solution] The information processing system comprises: an acquisition unit that acquires environmental information, which is information about the environment surrounding a mobile body, and state information, which is information about the state of the mobile body; a response control unit that generates response control information, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupants of the mobile body, based on the environmental information and the state information; and a response generation unit that generates response information based on the response control information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0006] , , , , , ,

[0005] , , ,

[0003] , , , , ,

[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program.

Background Art

[0002] Conventionally, technologies related to dialogue agents that communicate with passengers of a moving object in natural language are known. For example, a cognitive load imposed on a driver among the passengers of a vehicle is calculated from driving situation information indicating a situation where the driver is driving the vehicle, and when the cognitive load is large, it is determined whether the utterances of the driver and other passengers are dialogues related to the driving of the vehicle, and when the dialogue is a dialogue with low relevance to the driving of the vehicle, a technology for presenting a response for intervening in the dialogue is known.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the above prior art, when the cognitive load imposed on the driver is large, it only determines whether the utterances of the driver and other passengers are dialogues related to the driving of the vehicle, and when the dialogue is a dialogue with low relevance to the driving of the vehicle, it only presents a response for intervening in the dialogue, so it is not always possible to make a response according to the environment around the moving object.

[0005] Therefore, the present disclosure proposes an information processing system, an information processing method, and an information processing program capable of making a response according to the environment around the moving object.

Means for Solving the Problems

[0006] An information processing system according to one aspect of the present disclosure includes: an acquisition unit that acquires environmental information, which is information about the environment surrounding a mobile body, and state information, which is information about the state of the mobile body; a response control unit that generates response control information, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupant of the mobile body, based on the environmental information and the state information; and a response generation unit that generates the response information based on the response control information. [Brief explanation of the drawing]

[0007] [Figure 1] This figure shows an example of the configuration of an information processing system according to the first embodiment. [Figure 2] This is a block diagram showing an example of the functional configuration of a terminal device according to the first embodiment. [Figure 3] This figure shows an example of the data flow of the processing according to the first embodiment. [Figure 4] This figure shows an example of a dialogue in use case UC1. [Figure 5] This figure shows an example of a dialogue in use case UC2. [Figure 6] This figure shows an example of a dialogue in use case UC3. [Figure 7] This figure shows an example of a dialogue in use case UC4. [Figure 8] This is a flowchart showing an example of information processing according to the first embodiment. [Figure 9] This is a diagram illustrating the overview of the process according to the second embodiment. [Figure 10] This is a flowchart showing an example of information processing according to the second embodiment. [Figure 11] This is a diagram illustrating the overview of the process according to the third embodiment. [Figure 12] This is a diagram illustrating the overview of the process according to the third embodiment. [Figure 13] This is a flowchart showing an example of information processing according to the third embodiment. [Figure 14] It is a diagram for explaining a method for determining an interrupt utterance according to the third embodiment. [Figure 15] It is a diagram for explaining an outline of a process according to the fourth embodiment. [Figure 16] It is a diagram for explaining an outline of a process according to the fourth embodiment. [Figure 17] It is a flowchart showing an example of information processing according to the fourth embodiment. [Figure 18] It is a diagram for explaining an outline of a process according to the fifth embodiment. [Figure 19] It is a diagram for explaining an example of information processing according to the fifth embodiment. [Figure 20] It is a diagram for explaining an example of information processing according to the fifth embodiment. <{ [Figure 21] It is a diagram for explaining an example of information processing according to the fifth embodiment. [Figure 22] It is a flowchart showing an example of information processing according to the fifth embodiment. [Figure 23] It is a diagram for explaining an outline of a process according to the sixth embodiment. [Figure 24] It is a diagram for explaining an outline of a process according to the sixth embodiment. [[ID=*33]] [Figure 25] It is a flowchart showing an example of information processing according to the sixth embodiment. [Figure 26] It is a diagram for explaining an outline of a process according to the seventh embodiment. [Figure 27] It is a diagram for explaining an outline of a process according to the seventh embodiment. [Figure 28] It is a flowchart showing an example of information processing according to the seventh embodiment. [Figure 29] It is a hardware configuration diagram showing an example of a computer that realizes the functions of an information processing apparatus.

Embodiments for Carrying Out the Invention

[0008] Embodiments of this disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals to avoid redundant descriptions.

[0009] This disclosure will be explained in the order of the items shown below. 1. Embodiment 1-1. First Embodiment 1-1-1. Configuration of the Information Processing System 1-1-2. Configuration of Information Processing Device 1-1-3. Data flow in information processing 1-1-4. Use Cases 1-1-4-1. Use Case UC1 1-1-4-2. Use Case UC2 1-1-4-3. Use Case UC3 1-1-4-4. Use Case UC4 1-1-5. Information Processing Procedures 1-2. Second Embodiment 1-2-1. Overview of Information Processing 1-2-2. Information Processing Procedures 1-3. Third Embodiment 1-3-1. Overview of Information Processing 1-3-2. Information Processing Procedures 1-4. Fourth Embodiment 1-4-1. Overview of Information Processing 1-4-2. Information Processing Procedures 1-5. Fifth Embodiment 1-5-1. Overview of Information Processing 1-5-2. Information Processing Procedures 1-6. Sixth Embodiment 1-6-1. Overview of Information Processing 1-6-2. Information Processing Procedures 1-7. Seventh Embodiment 1-7-1. Overview of Information Processing 1-7-2. Information Processing Procedures 1-8. About Generative Models 2. Others 3. Hardware Configuration

[0010] <1. Embodiments> The information processing system related to this disclosure and the information processing performed by the information processing system will be described below in each embodiment, from the first embodiment to the seventh embodiment, and so on.

[0011] The information processing system described herein generates response information regarding the responses of a dialogue agent that engages in natural language dialogue with the occupants of a mobile vehicle. The information processing system also outputs the response information as audio.

[0012] The mobile object relating to this disclosure may be any type of mobile object, such as an automobile, electric vehicle, hybrid electric vehicle, motorcycle, bicycle, personal mobility device, airplane, drone, ship, robot, construction machinery, or agricultural machinery (tractor). The following description will focus on the case where the mobile object is an automobile (also referred to as a vehicle).

[0013] <1-1. First Embodiment> For example, when driving in a congested city or in a situation requiring emergency avoidance, the driver (or one of the vehicle's occupants) may not want to be interrupted by small talk from a conversational agent. On the other hand, on a straight, empty highway, the driver may want long conversations to stay awake. In response to this, the first embodiment describes a case where the information processing system grasps the surrounding environment and driving conditions of the vehicle and performs response control based on that information. This allows the information processing system to refrain from having the conversational agent speak when the vehicle occupant does not want to be spoken to. Furthermore, the information processing system can have the conversational agent speak more when the vehicle occupant wants it to.

[0014] <1-1-1. Configuration of the Information Processing System> Figure 1 shows an example of the configuration of an information processing system 1 according to the first embodiment. As shown in Figure 1, the information processing system 1 includes a terminal device 10 and an information processing device 100. The terminal device 10 and the information processing device 100 are connected, for example, via a network N, by wired or wireless means.

[0015] Terminal device 10 is an information processing device operated by the occupants of a vehicle. Terminal device 10 is an information processing device operated by the occupants of a vehicle that utilizes conversational AI (Artificial Intelligence) (hereinafter also referred to as a conversational agent). For example, terminal device 10 may be a smartphone used by the occupants, or it may be a navigation device installed in the vehicle. The occupants of a vehicle are, for example, the driver or passengers of the vehicle. Hereinafter, the occupants of a vehicle may be simply referred to as occupants. Terminal device 10 accepts voice input from the occupants through speech and input from the occupants through operation. Terminal device 10 also accepts multimodal input, such as input of images captured by a camera. Terminal device 10 obtains response sentences from information processing device 100. Terminal device 10 also provides the occupants with a conversation with the conversational agent by outputting the response sentences obtained from information processing device 100 as voice or screen display. The terminal device 10 may, for example, generate a response sentence using a Small Language Model (SLM) and output the generated response sentence via voice, screen display, or other means.

[0016] The terminal device 10 includes a display unit 101, an operation unit 102, a camera 103, a microphone 104, a speaker 105, a communication unit 11, a storage unit 12, and a control unit 13. Examples of terminal devices 10 include personal computers, smartphones, and in-vehicle terminals.

[0017] The display unit 101 is a display device for displaying various information. The display unit 101 can be implemented as a display device such as a liquid crystal display or an organic EL (Electro-Luminescence) display. The display unit 101 displays various screens such as dialogue between the crew and the dialogue agent, and search results. The dialogue includes utterances corresponding to speech by the crew or the dialogue agent, and response sentences corresponding to the crew or the dialogue agent's response to the speech by the crew or the dialogue agent.

[0018] The operation unit 102 is an input device that receives various operations from the crew member operating the terminal device 10. The operation unit 102 can be implemented as an input device such as a keyboard, mouse, or touch panel. The display device of the display unit 101 and the input device of the operation unit 102 may be integrated, such as a touch panel display.

[0019] Camera 103 is, for example, an external camera (also called an on-board camera) that captures images of the environment around the vehicle. Alternatively, camera 103 may be, for example, an internal camera that captures images of the inside of the vehicle. For example, camera 103 may be an internal camera that captures images of occupants operating the terminal device 10. Camera 103 captures images using, for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor as the image sensor. Camera 103 generates an image by photoelectric conversion of the light received by the image sensor and A / D conversion. Camera 103 outputs the captured image to the control unit 130.

[0020] Microphone 104 acquires, for example, the voice of an occupant operating the terminal device 10. Microphone 104 acquires the voice of the vehicle driver or passenger. Microphone 104 can use various types of microphones, such as an electret condenser microphone. Microphone 104 outputs the audio signal of the acquired voice to the control unit 130.

[0021] Speaker 105 outputs, for example, the content of the dialogue agent's speech. Speaker 105 can use various types of speakers, such as dynamic or condenser speakers. Speaker 105 outputs sound based on the audio signal input from the control unit 130.

[0022] The communication unit 11 is implemented by, for example, a NIC (Network Interface Card), a wireless communication network such as LTE (Long Term Evolution), 4G (4th Generation), 5G (5th Generation: 5th Generation Mobile Communication System), or a wireless LAN (Local Area Network) such as Bluetooth (registered trademark) or Wi-Fi (registered trademark). The communication unit 11 is connected to the information processing device 100 or an external information processing device (e.g., an external API) via the network N by wired or wireless connection, and is a communication interface that manages the communication of information between the information processing device 100 and the external information processing device.

[0023] The memory unit 12 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical discs.

[0024] The control unit 13 is implemented, for example, by a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing a program stored in its internal memory using RAM as the working area. Alternatively, the control unit 13 may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0025] The control unit 13 acquires arbitrary modal data related to the vehicle's surrounding environment and occupants via input devices of the terminal device 10, such as the operation unit 102, camera 103, and microphone 104. For example, the control unit 13 acquires text input from the occupants via the operation unit 102.

[0026] Furthermore, the control unit 13 acquires images of the environment surrounding the vehicle via the camera 103. For example, the control unit 13 acquires images of the environment surrounding the vehicle via an external camera. For example, the control unit 13 acquires images of the environment surrounding the vehicle that are equivalent to the view of the occupants. When the control unit 13 acquires images of the environment surrounding the vehicle, it transmits information about the images of the environment surrounding the vehicle to the information processing device 100 via the communication unit 11.

[0027] Furthermore, the control unit 13 acquires images of the interior of the vehicle via the camera 103. For example, the control unit 13 acquires images of the interior of the vehicle via the interior camera. For example, the control unit 13 acquires images of the occupants as images of the interior of the vehicle. When the control unit 13 acquires images of the interior of the vehicle, it transmits information about the images of the interior of the vehicle to the information processing device 100 via the communication unit 11.

[0028] Furthermore, the control unit 13 acquires the voice of the crew member's speech or ambient sounds via the microphone 104. The voice of the crew member's speech may be a conversation between crew members. When the control unit 13 acquires the voice of the crew member's speech, it performs speech recognition and acquires the recognized text. For example, the control unit 13 uses any technology such as Speech2Text as appropriate to convert the voice of the crew member's speech into text data, and acquires this text data as dialogue history information that shows the history of the conversation between the crew member and the dialogue agent. When the control unit 13 acquires the dialogue history information, it transmits information related to the dialogue history information to the information processing device 100 via the communication unit 11.

[0029] Furthermore, the control unit 13 acquires status information, which is information related to the state of the vehicle. For example, the control unit 13 acquires information related to the vehicle's driving state as status information. For example, the control unit 13 acquires information related to the vehicle's current location, direction of travel, speed, or driving conditions such as right or left turns, cruise control, or emergency avoidance as information related to the vehicle's driving state. Also, if a driving route is set for the vehicle, the control unit 13 acquires information related to the driving route as information related to the vehicle's driving state. When the control unit 13 acquires status information, it transmits the information related to the status information to the information processing device 100 via the communication unit 11.

[0030] The status information also includes information about the status of the conversational agent installed in the vehicle. The control unit 13 acquires response status information, which is information about the current response state of the conversational agent. For example, the control unit 13 acquires information as response status information indicating whether or not an audio corresponding to a response sentence indicating a response by the conversational agent is being output from the speaker. For example, if an audio corresponding to a response sentence indicating a response by the conversational agent is being output from the speaker, the control unit 13 acquires information as response status information indicating that the conversational agent is responding. Also, if an audio corresponding to a response sentence indicating a response by the conversational agent is not being output from the speaker, the control unit 13 acquires information as response status information indicating that the response by the conversational agent has been interrupted (or is not responding). The control unit 13 also acquires information as response status information indicating the content of the response by the conversational agent. For example, the control unit 13 acquires information about a response sentence indicating a response by the conversational agent. The control unit 13 may also acquire information as response status information indicating the type of content of the response by the conversational agent (e.g., casual conversation). When the control unit 13 acquires response status information, it transmits information related to the response status information to the information processing device 100 via the communication unit 11.

[0031] The information processing device 100 is, for example, a cloud server. The information processing device 100 acquires various information from the terminal device 10. The information processing device 100 also calls the API of a generative model and inputs the information and prompts acquired from the terminal device 10 into the generative model to generate a response sentence. For example, as an example of a generative model, the information processing device 100 calls the API of a multimodal large-scale language model and inputs the information and prompts acquired from the terminal device 10 into the multimodal large-scale language model to generate a response sentence. The information processing device 100 also outputs the generated response sentence to the terminal device 10. The terminal device 10 provides the crew with the opportunity to interact with the dialogue agent by outputting the response sentence acquired from the information processing device 100 as speech. The information processing device 100 also performs searches for text and images based on multimodal input. For example, the information processing device 100 determines whether or not a search needs to be performed based on the information acquired from the terminal device 10, and if it determines that a search needs to be performed, it performs the search. For example, the information processing device 100 inputs information obtained from the terminal device 10 into a generation model and has the generation model determine whether or not a search needs to be performed. If the generation model determines that a search needs to be performed, the information processing device 100 decides to perform the search. If the information processing device 100 decides to perform the search, it generates a query to be used for the search. For example, if the information processing device 100 decides to perform the search, it generates a query according to the search method, such as a web search, map search, or image search. The information processing device 100 also calls an API to perform a web search, map search, or image search, and performs the search. The information processing device 100 also obtains the search results and generates a response statement based on the obtained search results. In this way, the information processing system 1 performs a method of generating a response using retrieved external knowledge (RAG: Retrieval Augmented Generation).

[0032] <1-1-2. Configuration of the Information Processing Device> Figure 2 is a block diagram showing an example of the functional configuration of the information processing device 100 according to the first embodiment. As shown in Figure 2, the information processing device 100 has a communication unit 110, a storage unit 120, and a control unit 130. In this embodiment, the case in which the information processing device 100 has a storage unit 120 and a control unit 130 is described, but the storage unit 120 and the control unit 130 may be provided by an information processing device other than the information processing device 100. For example, a terminal device 10 may have a storage unit 120 and a control unit 130.

[0033] The communication unit 110 is implemented by, for example, a NIC (Network Interface Card), a wireless communication network such as LTE (Long Term Evolution), 4G (4th Generation), 5G (5th Generation: 5th Generation Mobile Communication System), or a wireless LAN (Local Area Network) such as Bluetooth (registered trademark) or Wi-Fi (registered trademark). The communication unit 110 is connected to the terminal device 10 or an external information processing device (e.g., an external API) via the network N by wired or wireless connection, and is a communication interface that manages the communication of information between the terminal device 10 and the external information processing device. For example, the communication unit 110 receives images captured by the camera 103 from the terminal device 10.

[0034] The storage unit 120 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical discs. The storage unit 120 includes an environmental information storage unit 121, a state information storage unit 122, and a history information storage unit 123. The storage unit 120 also stores information (programs and data) used for processing in the control unit 130.

[0035] The environmental information storage unit 121 stores environmental information, which is information about the environment surrounding the vehicle. For example, the environmental information storage unit 121 stores images captured by the vehicle's external cameras as environmental information. For example, the environmental information storage unit 121 stores information about images of the environment surrounding the vehicle as environmental information. For example, the environmental information storage unit 121 stores information about images equivalent to the occupant's field of view as images of the environment surrounding the vehicle. The environmental information storage unit 121 may also store information about ambient sounds or smells as environmental information.

[0036] The state information storage unit 122 stores state information, which is information relating to the state of the vehicle. For example, the state information storage unit 122 stores information relating to the vehicle's driving state. For example, the state information storage unit 122 stores information relating to the vehicle's driving state, such as the vehicle's current location, direction of travel, speed, or driving conditions such as right or left turns, cruise control, or emergency avoidance. In addition, if a driving route is set for the vehicle, the state information storage unit 122 stores information relating to the driving route as information relating to the vehicle's driving state.

[0037] Furthermore, the state information storage unit 122 stores sensor information detected inside the vehicle as state information. For example, the state information storage unit 122 stores images captured by the vehicle's internal camera as sensor information. For example, the state information storage unit 122 stores information about the occupants as sensor information. For example, the state information storage unit 122 stores information about voice acquired by a microphone as sensor information. For example, the state information storage unit 122 stores information about the occupants' speech as sensor information. The state information storage unit 122 may also store information about the dialogue agent's speech as sensor information. In the following, the dialogue agent's speech may be referred to as the dialogue agent's response.

[0038] Furthermore, the state information storage unit 122 stores response state information, which is information about the current response state of the dialogue agent. For example, the state information storage unit 122 stores information as response state information indicating whether or not audio corresponding to a response statement indicating a response by the dialogue agent is being output from the speaker. For example, the state information storage unit 122 may store information as response state information indicating that the dialogue agent is responding if audio corresponding to a response statement indicating a response by the dialogue agent is being output from the speaker. Also, the state information storage unit 122 may store information as response state information indicating that the dialogue agent's response has been interrupted (or is not responding) if audio corresponding to a response statement indicating a response by the dialogue agent is not being output from the speaker. Furthermore, the state information storage unit 122 stores information indicating the content of the dialogue agent's response as response state information. For example, the state information storage unit 122 stores information about a response statement indicating a response by the dialogue agent as response state information. Furthermore, the state information storage unit 122 may store information indicating the type of response from the dialogue agent (for example, casual conversation) as response state information.

[0039] The history information storage unit 123 stores various types of history information. Specifically, the history information storage unit 123 stores dialogue history information that shows the history of conversations between the crew and the dialogue agent. In addition, the history information storage unit 123 stores search history information that shows the history of searches performed by the dialogue agent.

[0040] The memory unit 120 may store model information, which is information relating to the generative model. For example, the memory unit 120 may store information relating to the language model as model information. For example, the memory unit 120 may store information relating to large language models (LLMs) as model information. For example, the memory unit 120 may store information relating to a multimodal large language model, which is a generative model that supports multimodal input, as model information. For example, the memory unit 120 may store information relating to a multimodal large language model as model information.

[0041] The control unit 130 is implemented, for example, by a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing a program stored in its internal memory using RAM as the working area. Alternatively, the control unit 130 may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0042] The control unit 130 comprises an acquisition unit 131, a response control unit 132, a response generation unit 133, and an output control unit 134, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Figure 2, and other configurations are also acceptable as long as they perform the information processing described later.

[0043] The acquisition unit 131 acquires information from the terminal device 10 via the input devices of the terminal device 10, such as the operation unit 102, camera 103, and microphone 104. For example, the acquisition unit 131 acquires arbitrary modal data from the terminal device 10 related to the environment around the vehicle and the occupants.

[0044] Furthermore, the acquisition unit 131 acquires images of the environment surrounding the vehicle, which are captured via the camera 103 from the terminal device 10. For example, the acquisition unit 131 acquires images of the environment surrounding the vehicle, which are captured via an external camera from the terminal device 10. For example, the acquisition unit 131 acquires images equivalent to the occupant's field of view as images of the environment surrounding the vehicle. When the acquisition unit 131 acquires images of the environment surrounding the vehicle, it stores information related to the images of the environment surrounding the vehicle in the environment information storage unit 121. The acquisition unit 131 may also acquire images of the environment surrounding the vehicle by referring to the environment information storage unit 121.

[0045] Furthermore, the acquisition unit 131 acquires images of the interior of the vehicle, which are acquired from the terminal device 10 via the camera 103. For example, the acquisition unit 131 acquires images of the interior of the vehicle, which are acquired from the terminal device 10 via the internal camera. For example, the acquisition unit 131 acquires images of the occupants as images of the interior of the vehicle. When the acquisition unit 131 acquires images of the interior of the vehicle, it stores information related to the images of the interior of the vehicle in the state information storage unit 122. The acquisition unit 131 may also acquire images of the interior of the vehicle by referring to the state information storage unit 122.

[0046] Furthermore, the acquisition unit 131 acquires the voice of the crew member's speech or ambient sounds from the terminal device 10 via the microphone 104. The voice of the crew member's speech may be a conversation between crew members. When the acquisition unit 131 acquires the voice of the crew member's speech, it performs speech recognition and acquires the recognized text. The acquisition unit 131 also performs speech recognition to recognize the speaker. For example, the acquisition unit 131 uses any technology such as Speech2Text as appropriate to convert the voice of the crew member's speech into text data, and acquires this text data as dialogue history information showing the history of the conversation between the crew member and the dialogue agent. When the acquisition unit 131 acquires the dialogue history information, it stores the information related to the dialogue history information in the history information storage unit 123. For example, the acquisition unit 131 also acquires the dialogue history information by referring to the history information storage unit 123. The response generation unit 133 may also acquire these texts as queries. For example, the response generation unit 133 performs morphological analysis on the recognized text and, if it determines that it is a question for the dialogue agent, it obtains the text as a query. When the response generation unit 133 obtains the text as a query, it performs a search. For example, the response generation unit 133 calls an API to perform a web search, map search, or image search, and performs the web search, map search, or image search. The response generation unit 133 also obtains the search results obtained by performing the web search, map search, or image search. The acquisition unit 131 also obtains information about the query and search results as search history information that shows the history of searches by the dialogue agent. When the acquisition unit 131 obtains search history information, it stores information about the search history information in the history information storage unit 123. For example, the acquisition unit 131 also obtains search history information by referring to the history information storage unit 123.

[0047] Furthermore, the acquisition unit 131 acquires status information, which is information relating to the state of the vehicle. For example, the acquisition unit 131 acquires status information from the terminal device 10. Alternatively, the acquisition unit 131 acquires status information by referring to the status information storage unit 122. For example, the acquisition unit 131 acquires information relating to the vehicle's driving state as status information. For example, the acquisition unit 131 acquires information relating to the vehicle's current location, direction of travel, speed, or driving conditions such as right or left turns, cruise control, or emergency avoidance as information relating to the vehicle's driving state. Also, if a driving route is set for the vehicle, the acquisition unit 131 acquires information relating to the driving route as information relating to the vehicle's driving state.

[0048] Furthermore, the acquisition unit 131 acquires response status information, which is information about the current response state of the dialogue agent, as status information. For example, the acquisition unit 131 acquires information as response status information indicating whether or not audio corresponding to a response sentence indicating a response by the dialogue agent is being output from the speaker. For example, if audio corresponding to a response sentence indicating a response by the dialogue agent is being output from the speaker, the acquisition unit 131 acquires information as response status information indicating that the dialogue agent is responding. Also, if audio corresponding to a response sentence indicating a response by the dialogue agent is not being output from the speaker, the acquisition unit 131 acquires information as response status information indicating that the dialogue agent's response has been interrupted (or is not responding). Furthermore, the acquisition unit 131 acquires information as response status information indicating the content of the dialogue agent's response. For example, the acquisition unit 131 acquires information regarding a response sentence indicating a response by the dialogue agent as response status information. Furthermore, the acquisition unit 131 may acquire information as response status information indicating the type of content of the dialogue agent's response (e.g., casual conversation).

[0049] The response control unit 132 generates response control information, which is information for generating response information related to the response of a dialogue agent that interacts with the occupants of a mobile vehicle in natural language, based on environmental information and state information. Specifically, when the response control unit 132 obtains environmental information and state information from the acquisition unit 131, it generates response control information based on the environmental information and state information. For example, the response control unit 132 generates a prompt as response control information that instructs the generation of a response statement corresponding to the environmental information and state information. For example, the response control unit 132 generates a prompt by inputting the environmental information and state information obtained by the acquisition unit 131 into a prompt template that includes the environmental information and state information as variables. The response control unit 132 obtains the generated prompt as response control information.

[0050] The response control unit 132 may generate response control information based on environmental information, status information, search history information, and dialogue history information. For example, the response control unit 132 may generate a prompt as response control information that instructs the system to generate a response statement corresponding to the environmental information, status information, search history information, and dialogue history information. For example, the response control unit 132 may generate a prompt by inputting the environmental information, status information, search history information, and dialogue history information acquired by the acquisition unit 131 into a prompt template that includes each of the environmental information, status information, search history information, and dialogue history information as variables.

[0051] Furthermore, the response control unit 132 generates information indicating the behavior of the dialogue agent according to the environmental information and state information as response control information. For example, the response control unit 132 obtains information about the multimodal large-scale language model by referring to the storage unit 120. The response control unit 132 also inputs the environmental information and state information into the multimodal large-scale language model to generate information indicating the behavior of the dialogue agent according to the environmental information and state information. Alternatively, instead of using the multimodal large-scale language model, the response control unit 132 may generate information indicating the behavior of the dialogue agent according to the environmental information and state information based on predetermined rules. For example, as information indicating the behavior of the dialogue agent, the response control unit 132 generates information indicating whether or not to interrupt the dialogue agent's response if the dialogue agent is responding. Also, as information indicating the behavior of the dialogue agent, the response control unit 132 generates information indicating whether or not to resume or start a new response if the dialogue agent's response is interrupted (or not responding). Furthermore, the response control unit 132 generates information indicating the type of response statement that indicates the dialogue agent's response in accordance with the dialogue agent's behavior, as information indicating the dialogue agent's behavior. For example, the response control unit 132 generates information indicating the content of a response statement that indicates the dialogue agent's response in accordance with the dialogue agent's behavior, as information indicating the dialogue agent's behavior.

[0052] The response control unit 132 may also generate information indicating the behavior of the dialogue agent in accordance with the environment information, state information, search history information, and dialogue history information as response control information. For example, the response control unit 132 may input the environment information, state information, search history information, and dialogue history information into a multimodal large-scale language model to generate information indicating the behavior of the dialogue agent in accordance with the environment information, state information, search history information, and dialogue history information.

[0053] The response generation unit 133 generates response information based on the response control information generated by the response control unit 132. For example, if the response control unit 132 generates response control information, the response generation unit 133 calls the API of the multimodal large-scale language model. The response generation unit 133 may also obtain information about the multimodal large-scale language model by referring to the storage unit 120. For example, the response generation unit 133 inputs the response control information into the multimodal large-scale language model to generate response information. For example, the response generation unit 133 inputs environment information, state information, and response control information into the multimodal large-scale language model to generate response information. For example, the response information is a response statement indicating the response of the dialogue agent. For example, the response generation unit 133 inputs response control information, which is a prompt, into the multimodal large-scale language model to generate response information. For example, the response generation unit 133 inputs response control information, which is a prompt instructing the multimodal large-scale language model to generate a response statement corresponding to the environment information, state information, and response control information, to generate response information. Furthermore, the response generation unit 133 inputs response control information, which is information indicating the behavior of the dialogue agent, into the multimodal large-scale language model to generate response information. For example, the response generation unit 133 inputs response control information, which is information indicating the behavior of the dialogue agent, into the multimodal large-scale language model to generate response information. However, if the response control unit 132 generates information indicating that the dialogue agent's response will be interrupted, the response generation unit 133 decides not to generate response information. For example, if the response control unit 132 generates information indicating that the dialogue agent's response will be interrupted, the response generation unit 133 may decide not to input response control information into the multimodal large-scale language model.

[0054] Furthermore, the response generation unit 133 may input environmental information, state information, search history information, dialogue history information, and response control information into a multimodal large-scale language model to generate response information. For example, the response generation unit 133 may input response control information, which is a prompt instructing the multimodal large-scale language model to generate a response sentence corresponding to the environmental information, state information, search history information, dialogue history information, and response control information, to generate response information.

[0055] Furthermore, the response generation unit 133 may determine whether or not it is possible to generate response information based on environmental information and state information. For example, the response generation unit 133 may have a multimodal large-scale language model determine whether or not it is possible to generate response information based on environmental information and state information. The response generation unit 133 may have a multimodal large-scale language model generate information indicating whether or not it is possible to generate response information based on environmental information and state information. The response generation unit 133 obtains information indicating whether or not it is possible to generate response information based on environmental information and state information. If the response generation unit 133 obtains information indicating that it is possible to generate response information based on environmental information and state information, it decides not to perform a search. On the other hand, if the response generation unit 133 obtains information indicating that it is impossible to generate response information based on environmental information and state information, it decides to perform a search. If the response generation unit 133 decides to perform a search, it performs the search. For example, the response generation unit 133 generates a query to perform the search. The response generation unit 133 also performs the search based on the generated query. For example, the response generation unit 133 calls an API to perform a web search, map search, or image search, and executes the web search, map search, or image search. The response generation unit 133 also obtains the search results obtained by performing the web search, map search, or image search. When the response generation unit 133 obtains search results, it generates response information based on the search results. For example, the response generation unit 133 inputs the search results into a multimodal large-scale language model to generate response information.

[0056] The response generation unit 133 may determine whether or not it is possible to generate response information based on environmental information, state information, search history information, and dialogue history information. For example, the response generation unit 133 may cause a multimodal large-scale language model to determine whether or not it is possible to generate response information based on environmental information, state information, search history information, and dialogue history information. The response generation unit 133 may cause a multimodal large-scale language model to generate information indicating whether or not it is possible to generate response information based on environmental information, state information, search history information, and dialogue history information.

[0057] The output control unit 134 outputs response information from the speaker when response information is generated by the response generation unit 133. For example, when a response statement is generated by the response generation unit 133, the output control unit 134 outputs the audio corresponding to the response statement from the speaker. Also, when the response control unit 132 generates information indicating that the dialogue agent's response should be interrupted, the output control unit 134 stops outputting the audio corresponding to the response statement from the speaker.

[0058] <1-1-3. Data flow in information processing> Figure 3 shows an example of the data flow of processing according to the first embodiment. In Figure 3, the acquisition unit 131 acquires information regarding the vehicle's current location, direction of travel, speed, or driving conditions such as right or left turns, cruise control, or emergency avoidance, as state information. The acquisition unit 131 also acquires images of the environment around the vehicle as environmental information. The acquisition unit 131 acquires images that include information indicating the congestion status around the vehicle. The acquisition unit 131 also stores sensor information detected inside the vehicle as state information. For example, the acquisition unit 131 acquires images captured by the vehicle's internal camera as sensor information. For example, the acquisition unit 131 acquires information regarding occupants as sensor information. For example, the acquisition unit 131 acquires information regarding voice acquired by a microphone as sensor information. For example, the acquisition unit 131 acquires information regarding the voice of occupants' speech as sensor information. The acquisition unit 131 may also acquire information regarding the voice of a dialogue agent's speech as sensor information. Furthermore, the acquisition unit 131 acquires dialogue history information that shows the history of conversations between the crew and the dialogue agent. In addition, the acquisition unit 131 acquires search history information that shows the history of searches performed by the dialogue agent.

[0059] Furthermore, in Figure 3, the response control unit 132 generates response control information, which is information for generating response information related to the response of a dialogue agent that interacts with the occupants of a mobile vehicle in natural language, based on the environmental information, state information, dialogue history information, and search history information acquired by the acquisition unit 131. The response generation unit 133 generates response information based on the response control information. If the response generation unit 133 determines that it is necessary to perform a search, it calls an API to perform the search and performs the search. The response generation unit 133 also acquires the search results obtained by performing the search. The response generation unit 133 also generates response information based on the search results. When response information is generated by the response generation unit 133, the output control unit 134 outputs the response information from the speaker.

[0060] <1-1-4. Use Cases> From here, we will describe the use cases UC1 to UC4 of the first embodiment using Figures 4 to 7.

[0061] <1-1-4-1. Use Case UC1> Figure 4 shows an example of dialogue in Use Case UC1. In Use Case UC1, Information Processing System 1 interrupts the conversational agent's responses to casual conversation when the driver needs to concentrate on driving operations because a lane change is imminent. This prevents the driver from being interrupted by the conversational agent's responses when the driver needs to concentrate on driving operations.

[0062] In Figure 4, the acquisition unit 131 acquires an image G11 from an external camera showing the Tokyo Monorail on the left side of the road as environmental information. The acquisition unit 131 also acquires information about the vehicle's travel path as state information. For example, the acquisition unit 131 acquires information indicating that a lane change is imminent as information about the vehicle's travel path. For example, the acquisition unit 131 acquires information indicating that a highway exit is approaching as information indicating that a lane change is imminent. The acquisition unit 131 also acquires information indicating that the dialogue agent is responding as state information. For example, the acquisition unit 131 acquires information indicating that audio corresponding to the conversational agent's casual conversation response R11 is being output from the speaker. The response control unit 132 inputs the environmental information and state information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate response control information C11. For example, the response control unit 132 generates response control information C11 indicating that the conversational agent's casual conversation response will be interrupted because a lane change is imminent. Furthermore, the response generation unit 133 decides not to generate a response sentence corresponding to the casual conversation response if the response control unit 132 has generated the response control information C11. Also, the output control unit 134 stops outputting the audio corresponding to the response sentence corresponding to the casual conversation response from the speaker if the response control unit 132 has generated the response control information C11.

[0063] Furthermore, the response generation unit 133 generates a response sentence based on the response control information C11. For example, the response generation unit 133 inputs the environmental information, state information, and response control information C11 acquired by the acquisition unit 131 into a multimodal large-scale language model to generate a response sentence corresponding to response R12, which indicates that a lane change is imminent (that the highway exit is approaching). The output control unit 134 also outputs audio corresponding to the response sentence generated by the response generation unit 133 from the speaker.

[0064] Furthermore, situations in which the driver must concentrate on driving operations include not only when a lane change is imminent, but also, for example, when a right or left turn is imminent, when traffic congestion begins, or when emergency maneuvers are required in response to a sudden cut-in. For example, if the response control unit 132 receives information from the acquisition unit 131 indicating that a right or left turn is imminent, information indicating a change in vehicle speed such as when traffic congestion begins, or information indicating an emergency maneuver in response to a sudden cut-in, it generates response control information indicating that the conversational agent's response to casual conversation should be interrupted. Alternatively, the response control unit 132 may generate response control information indicating that the conversational agent's response to casual conversation should be ended quickly, instead of generating response control information indicating that the conversational agent's response to casual conversation should be ended quickly. In addition, the information processing system 1 may interrupt the conversational agent's response to casual conversation if there is a response with a higher priority waiting from the conversational agent (in Figure 4, the response R12 indicating an imminent lane change has a higher priority than the casual conversation response R11).

[0065] <1-1-4-2. Use Case UC2> Figure 5 shows an example of a dialogue in Use Case UC2. Use Case UC2 corresponds to a situation in Use Case UC1 where the conversational agent's casual conversation response was interrupted, and then resumed after a calm period of time.

[0066] In Figure 5, the acquisition unit 131 acquires information as state information indicating that the conversational agent is responding. For example, the acquisition unit 131 acquires information from the speaker indicating that audio corresponding to the response sentence corresponding to the conversational agent's casual conversation response R21 is being output. The acquisition unit 131 also acquires image G12 from an external camera as environmental information, showing a person suddenly running into the road. The response control unit 132 inputs the environmental information and state information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate response control information C21. For example, the response control unit 132 generates response control information C21 indicating that the conversational agent's casual conversation response will be interrupted because a person has suddenly run into the road. The response control unit 132 may also generate response control information C21 based on information acquired from the vehicle's driver assistance system. For example, the response control unit 132 may acquire information from the vehicle's driver assistance system indicating that the emergency brake has been activated and generate response control information C21 based on the acquired information. Furthermore, the response generation unit 133 decides not to generate a response sentence corresponding to the casual conversation response if the response control unit 132 has generated the response control information C21. Also, the output control unit 134 stops outputting the audio corresponding to the response sentence corresponding to the casual conversation response from the speaker if the response control unit 132 has generated the response control information C21.

[0067] Furthermore, the acquisition unit 131 acquires the latest environmental information and status information each time. The response control unit 132 inputs the latest environmental information and status information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate response control information C22. For example, the response control unit 132 generates response control information C22 indicating that the emergency avoidance has ended and normal driving has returned, and therefore the dialogue agent will resume responding. The response generation unit 133 also generates a response statement indicating a response R22 corresponding to the response control information C22, based on the response control information C22. For example, the response generation unit 133 inputs the latest environmental information, status information, and response control information C22 acquired by the acquisition unit 131 into a multimodal large-scale language model to generate a response statement corresponding to a response R22 indicating that the emergency avoidance has ended and normal driving has returned, and therefore casual conversation will resume. The output control unit 134 also outputs the audio corresponding to the response statement generated by the response generation unit 133 from the speaker.

[0068] <1-1-4-3. Use Case UC3> Figure 6 shows an example of a dialogue in use case UC3. In use case UC3, information processing system 1 has a dialogue agent engage in a longer conversation when cruise control is being used on a straight road (also called a straight road). This allows information processing system 1 to have the dialogue agent speak more to the driver if the driver wants to be spoken to more.

[0069] In Figure 6, the acquisition unit 131 acquires dialogue history information showing the history of the conversation between the occupant and the dialogue agent. For example, the acquisition unit 131 acquires utterance information corresponding to the occupant's utterance U31 regarding a question about the name of a mountain located in front of the vehicle, as dialogue history information. The acquisition unit 131 also acquires response information corresponding to the dialogue agent's response R31 to the occupant's utterance U31, as dialogue history information. The acquisition unit 131 also acquires an image G13 as environmental information, showing the Hakone mountains in the distance along a straight road. The response control unit 132 inputs the environmental information and dialogue history information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate response control information C31. For example, the response control unit 132 generates response control information C31 indicating that the driver has time to begin a general knowledge response from the dialogue agent. The response generation unit 133 generates a response sentence corresponding to the response R32 in accordance with the response control information C31, based on the response control information C31. For example, the response generation unit 133 inputs the latest environmental information, dialogue history information, and response control information C31 acquired by the acquisition unit 131 into a multimodal large-scale language model to generate a response sentence corresponding to response R32, which shows trivia about the mountains of Hakone. The output control unit 134 also outputs the audio corresponding to the response sentence generated by the response generation unit 133 from the speaker.

[0070] <1-1-4-4. Use Case UC4> Figure 7 shows an example of a dialogue in use case UC4. In use case UC4, if the conversational agent is interacting with a passenger other than the driver, the information processing system 1 will respond to casual conversation from the conversational agent, even if the driver needs to concentrate on driving.

[0071] In Figure 7, the acquisition unit 131 acquires an image G14 from an external camera showing the Tokyo Monorail on the left side of the road as environmental information. The acquisition unit 131 also acquires information about the vehicle's travel path as state information. For example, the acquisition unit 131 acquires information indicating that a lane change is imminent as information about the vehicle's travel path. For example, the acquisition unit 131 acquires information indicating that a highway exit is approaching as information indicating that a lane change is imminent. The acquisition unit 131 also acquires information indicating that the dialogue agent is responding as state information. The acquisition unit 131 also acquires dialogue history information showing the history of conversations between the occupants and the dialogue agent. For example, the acquisition unit 131 acquires speech information showing the speech U41 of an occupant other than the driver as dialogue history information. The response control unit 132 inputs the environmental information, state information, and dialogue history information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate response control information C41. For example, the response control unit 132 generates response control information C41 indicating that a lane change is imminent, but the speaker is not the driver, and therefore initiates a casual conversation response. The response generation unit 133 then generates a response statement R41 based on the response control information C41. For example, the response generation unit 133 inputs the latest environmental information, state information, dialogue history information, and response control information C41 acquired by the acquisition unit 131 into a multimodal large-scale language model to generate a response statement R41 corresponding to the response to the utterance U41 of a passenger other than the driver. The output control unit 134 then outputs the audio corresponding to the response statement generated by the response generation unit 133 from the speaker.

[0072] <1-1-5. Information Processing Procedure> Figure 8 is a flowchart showing an example of information processing according to the first embodiment. In Figure 8, the acquisition unit 131 acquires environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object (step S11).

[0073] Furthermore, the response control unit 132 generates response control information, which is information for generating response information regarding the response of a dialogue agent that interacts with the occupants of a mobile vehicle in natural language, based on environmental information and state information (step S12). For example, the response control unit 132 inputs the environmental information and state information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate response control information corresponding to the environmental information and state information. For example, the response control unit 132 inputs a prompt into the multimodal large-scale language model instructing it to generate response control information, which is information indicating the behavior of the dialogue agent according to the environmental information and state information, and generates response control information, which is information indicating the behavior of the dialogue agent according to the environmental information and state information.

[0074] Furthermore, the response generation unit 133 generates response information corresponding to the response control information based on the response control information (step S13). For example, the response generation unit 133 inputs the environmental information and state information acquired by the acquisition unit 131 and the response control information generated by the response control unit 132 into a multimodal large-scale language model to generate response information corresponding to the environmental information, state information and response control information. For example, the response generation unit 133 inputs the environmental information, state information, and information indicating the behavior of the dialogue agent according to the environmental information and state information into a multimodal large-scale language model to generate the environmental information, state information, and the behavior of the dialogue agent according to the environmental information and state information. For example, the response generation unit 133 generates a response statement indicating the dialogue agent's response according to the dialogue agent's behavior as the behavior of the dialogue agent.

[0075] As described above, the information processing system 1 comprises an acquisition unit 131, a response control unit 132, and a response generation unit 133. The acquisition unit 131 acquires environmental information, which is information about the environment surrounding the mobile body, and state information, which is information about the state of the mobile body. The response control unit 132 generates response control information, which is information for generating response information related to the response of a dialogue agent that engages in natural language dialogue with the occupant of the mobile body, based on the environmental information and the state information. The response generation unit 133 generates response information based on the response control information. As a result, the information processing system 1 can, for example, refrain from having the dialogue agent speak when the occupant of the mobile body does not want to be spoken to, based on the environmental information and the state information. Also, the information processing system 1 can, for example, make the dialogue agent speak more when the occupant of the mobile body wants it to speak more, based on the environmental information and the state information. In this way, the information processing system 1 can provide responses that are appropriate to the environment surrounding the mobile body.

[0076] Furthermore, the acquisition unit 131 acquires response state information, which is information about the current response state of the dialogue agent, as state information. The response control unit 132 generates response control information based on the environment information and the response state information. As a result, the information processing system 1 can, for example, make a response according to the current response state of the dialogue agent based on the response state information.

[0077] <1-2. Second Embodiment> For example, a dialogue agent that uses visual information (images) to communicate (hereinafter sometimes referred to as a visual dialogue agent) may refer to an object it is seeing, saying "This building is..." However, due to its own movement or the movement of the object, the object may no longer be visible by the time of the next response. In such cases, the occupants may feel uncomfortable when the dialogue agent refers to an object that is no longer visible using pronouns such as "this" or "that." For example, if a white building is visible in an image equivalent to what the occupants see, the dialogue agent may respond, "I will look into that white building." However, while the dialogue agent is generating the response sentence for the next response, the vehicle moves, and the white building is no longer visible in the image equivalent to what the occupants see. If, despite this, the dialogue agent responds, "That building is XX," it may cause the occupants to doubt whether the system is actually recognizing the object. In contrast, in the second embodiment, the dialogue agent's response sentence is changed from "That building is XX" to "The building from earlier is XX." In other words, in the second embodiment, the information processing system outputs a voice response in which pronouns such as "this" and "that" that indicate an object that is no longer visible to the occupant are replaced with words such as "the one from earlier." This enables the information processing system to allow the occupant to have a natural conversation with the dialogue agent.

[0078] <1-2-1. Overview of Information Processing> Figure 9 is a diagram illustrating the overview of the processing according to the second embodiment. In Figure 9, the acquisition unit 131 acquires an image G21 showing a building O2 along the road as environmental information. The acquisition unit 131 also acquires dialogue history information showing the history of the dialogue between the occupant and the dialogue agent. The response generation unit 133 identifies the object mentioned in the dialogue from among the objects identified based on the environmental information, based on the environmental information and the dialogue history information. Specifically, the response generation unit 133 inputs the environmental information and the dialogue history information into a multimodal large-scale language model to cause the multimodal large-scale language model to identify the object mentioned in the dialogue from among the objects identified based on the environmental information. For example, the response generation unit 133 inputs the environmental information and the dialogue history information into a multimodal large-scale language model to generate information indicating the object mentioned in the dialogue from among the objects identified based on the environmental information. In Figure 9, the response generation unit 133 inputs image G21 and dialogue history information into a multimodal large-scale language model to generate information indicating that the object mentioned in the dialogue from among the objects contained in image G21 is building O2.

[0079] Furthermore, the response generation unit 133 inputs information indicating the object mentioned in the dialogue and the image (video) acquired by the acquisition unit 131 into the tracking model and starts tracking the object mentioned in the dialogue among the objects included in the image. For example, the tracking model is an AI model that detects objects included in the image and tracks the detected objects. In Figure 9, the acquisition unit 131 acquires image G22, which shows building O2, as environmental information. The response generation unit 133 inputs information indicating that the object mentioned in the dialogue is building O2 and image G22 into the tracking model and tracks building O2 included in image G22.

[0080] Furthermore, the response generation unit 133 determines, based on environmental information, whether or not it is possible to track the object mentioned in the dialogue. In Figure 9, the acquisition unit 131 acquires image G23 as environmental information, in which building O2 is no longer visible due to the movement of the vehicle. Based on image G23, the response generation unit 133 determines that it is impossible to track building O2. If the response generation unit 133 determines that it is impossible to track building O2, it generates a response sentence that changes the name of building O2, which is the object mentioned in the dialogue. For example, the response generation unit 133 generates a response sentence in which the name "that" referring to building O2 in the response sentence is replaced with the name "the one mentioned earlier."

[0081] <1-2-2. Information Processing Procedure> Figure 10 is a flowchart showing an example of information processing according to the second embodiment. In Figure 10, the acquisition unit 131 acquires environmental information and dialogue history information showing the history of the dialogue between the crew and the dialogue agent (step S21).

[0082] Furthermore, the response generation unit 133 identifies the object mentioned in the dialogue from among the objects identified based on the environmental information, based on the environmental information and the dialogue history information (step S22). For example, the response generation unit 133 inputs the environmental information and the dialogue history information into a multimodal large-scale language model to generate information indicating the object mentioned in the dialogue from among the objects identified based on the environmental information. For example, the response generation unit 133 inputs a prompt into the multimodal large-scale language model instructing it to identify the object mentioned in the dialogue from among the objects identified based on the environmental information, and generates information indicating the object mentioned in the dialogue from among the objects identified based on the environmental information.

[0083] Furthermore, the response generation unit 133 tracks the target based on the latest environmental information (step S23). For example, the response generation unit 133 inputs information indicating the target mentioned in the dialogue and the latest environmental information into the tracking model to track the target included in the latest environmental information.

[0084] Furthermore, the response generation unit 133 determines whether or not the output of response information has been completed (step S24). For example, the response generation unit 133 determines whether or not the output of response information has been completed based on the response status information.

[0085] If the response generation unit 133 determines that the output of response information is complete (step S24; Yes), it terminates the process. On the other hand, if the response generation unit 133 determines that the output of response information is not complete (step S24; No), it determines whether or not it is possible to track the target (step S25). For example, the response generation unit 133 determines whether or not it is possible to track the target based on the output result of the tracking model. For example, the response generation unit 133 determines whether or not it is possible to track the target based on whether or not the tracking model detects the target mentioned in the dialogue. For example, if the tracking model detects the target mentioned in the dialogue, the response generation unit 133 determines that it is possible to track the target. On the other hand, if the tracking model does not detect the target mentioned in the dialogue, the response generation unit 133 determines that it is not possible to track the target.

[0086] If the response generation unit 133 determines that it is possible to track the target (step S25; Yes), it tracks the target based on the latest environmental information (step S23). For example, if the response generation unit 133 determines that it is possible to track the target, it inputs the latest environmental information into the tracking model and tracks the target included in the environmental information.

[0087] On the other hand, if the response generation unit 133 determines that it is impossible to track the target (step S25; No), it generates response information with a changed designation for the target (step S26). For example, the response generation unit 133 generates a response sentence in which demonstrative pronouns such as "this," "that," and "that over there" that indicate the target are replaced with the phrase "the one from earlier."

[0088] The acquisition unit 131 may also acquire modal information other than images as environmental information. For example, the acquisition unit 131 may acquire modal information other than images from various sensors mounted on the vehicle. For example, the acquisition unit 131 acquires ambient sounds around the vehicle from sound sensors (e.g., microphones) mounted on the vehicle. The response generation unit 133 inputs the ambient sounds around the vehicle and the dialogue history information into a multimodal large-scale language model to allow the multimodal large-scale language model to identify the sounds mentioned in the dialogue from among the sounds included in the ambient sounds around the vehicle. The response generation unit 133 also generates response information with the names of the sounds mentioned in the dialogue changed depending on whether or not the sounds mentioned in the dialogue are heard.

[0089] Furthermore, the acquisition unit 131 acquires information indicating the smell around the vehicle from an odor sensor mounted on the vehicle. The response generation unit 133 inputs the information indicating the smell around the vehicle and the dialogue history information into a multimodal large-scale language model, causing the multimodal large-scale language model to identify the smell mentioned in the dialogue from among the smells around the vehicle. The response generation unit 133 also generates response information that modifies the name of the smell mentioned in the dialogue, depending on whether or not the smell mentioned in the dialogue is present.

[0090] Furthermore, the response generation unit 133 may, for example, calculate the field of view from a person seated in the rear seat by combining an external omnidirectional camera and an internal camera, and based on the determination result of whether or not the object is no longer visible within that range, generate response information that modifies the name of the object mentioned in the conversation.

[0091] As described above, the acquisition unit 131 acquires dialogue history information showing the history of conversations between the crew and the dialogue agent. The response generation unit 133 generates response information that modifies the designation of the object mentioned in the conversation, based on the environmental information and the dialogue history information. This allows the information processing system 1 to appropriately use different designations for the object mentioned in the conversation, for example, depending on whether or not the object mentioned in the conversation can be identified. Therefore, the information processing system 1 can enable the crew to have a natural conversation with the dialogue agent.

[0092] Furthermore, the response generation unit 133 identifies the object mentioned in the dialogue from among the objects identified based on the environmental information, tracks the object based on the environmental information, and generates response information with a changed designation for the object if tracking the object becomes impossible. This allows the information processing system 1 to appropriately use different designations for the object mentioned in the dialogue depending on whether or not it can identify the object mentioned in the dialogue based on the environmental information.

[0093] Furthermore, the acquisition unit 131 acquires images of the environment surrounding the moving object as environmental information. The response generation unit 133 generates response information that modifies the names of the objects included in the images based on the images and the dialogue history information. As a result, the information processing system 1 can appropriately use different names for the objects mentioned in the dialogue, depending on whether or not it is possible to identify the objects mentioned in the dialogue based on the images of the environment surrounding the moving object.

[0094] Furthermore, the response generation unit 133 identifies the object mentioned in the dialogue from among the objects included in the image based on the image and the dialogue history information, tracks the object based on the image, and if tracking the object becomes impossible, generates response information with a changed designation for the object. As a result, the information processing system 1 can appropriately use different designations for the object mentioned in the dialogue depending on whether or not it is possible to identify the object mentioned in the dialogue based on the image of the environment surrounding the moving object.

[0095] <1-3. Third Embodiment> For example, if a crew member speaks while interrupting a dialogue agent, there are cases where the response to the crew member's speech should be prioritized and cases where it should not. Note that a crew member's speech during a dialogue agent's speech is sometimes referred to as an "interrupting utterance." In contrast, in the third embodiment, the information processing system determines whether or not the response to the interrupting utterance should be prioritized and provides a response according to the result of that determination. This allows the information processing system to enable crew members to speak while interrupting dialogue agents, and to reject unintentional interrupting utterances.

[0096] <1-3-1. Overview of Information Processing> Figures 11-12 are diagrams illustrating the overview of the processing according to the third embodiment. In Figure 11, the acquisition unit 131 acquires voice information corresponding to the crew member's first utterance. For example, in Figure 11, the acquisition unit 131 acquires voice information corresponding to the crew member's first utterance U31, "Is there a restaurant at the next service area?", as voice information corresponding to the crew member's first utterance. The acquisition unit 131 may also convert the acquired voice information into text and acquire the converted text. The acquisition unit 131 also acquires response state information, which is information regarding the current response state of the dialogue agent. For example, the acquisition unit 131 acquires information indicating that it is in the process of making a response R31 to the first utterance U31, "I will look into restaurants." (responding). The acquisition unit 131 also acquires voice information corresponding to the crew member's second utterance. For example, the acquisition unit 131 acquires voice information corresponding to the crew member's second utterance U32, "Wait, where is the monorail on the left going?", as voice information corresponding to the crew member's second utterance. The acquisition unit 131 may also convert the acquired voice information into text and acquire the converted text. The response control unit 132 determines whether the crew member's second utterance is an interruption utterance based on the response state information and the voice information. For example, in Figure 11, the response control unit 132 determines that the crew member's second utterance is an interruption utterance.

[0097] Furthermore, if the response control unit 132 determines that the occupant's second utterance is an interruption utterance, it causes the generative model to determine which response should be prioritized: the response to the occupant's second utterance or a response other than the response to the occupant's second utterance (for example, the response to the occupant's first utterance). For example, the response control unit 132 inputs the response state information and speech information into the multimodal large-scale language model and causes the multimodal large-scale language model to determine which response should be prioritized: the response to the occupant's second utterance or a response other than the response to the occupant's second utterance (for example, the response to the occupant's first utterance). For example, in Figure 11, the multimodal large-scale language model determines that the response to the occupant's second utterance U32 should be prioritized. The response control unit 132 determines that the multimodal large-scale language model has determined that the response to the occupant's second utterance U32 should be prioritized.

[0098] On the other hand, Figure 12 differs from Figure 11 in that the acquisition unit 131 acquires voice information corresponding to the crew member's second utterance U33, "Yes, please," as voice information corresponding to the crew member's second utterance. In Figure 12, the multimodal large-scale language model determines that responses other than the response to the crew member's second utterance U33 (for example, the response to the crew member's first utterance U31) should be prioritized. The response control unit 132 determines that the multimodal large-scale language model has determined that responses other than the response to the crew member's second utterance U33 (for example, the response to the crew member's first utterance U31) should be prioritized.

[0099] <1-3-2. Information Processing Procedure> Figure 13 is a flowchart showing an example of information processing according to the third embodiment. In Figure 13, the acquisition unit 131 acquires voice information corresponding to the occupant's speech and response state information, which is information regarding the current response state of the dialogue agent (step S31).

[0100] Furthermore, the response control unit 132 determines whether the crew member's utterance is an interruption utterance based on the response state information and the voice information (step S32). Specifically, the response control unit 132 determines, based on the response state information and the voice information, whether the crew member's utterance was acquired while the acquisition unit 131 was generating a response sentence or outputting a response sentence audibly. For example, if the response control unit 132 determines that the crew member's utterance was acquired while the acquisition unit 131 was generating a response sentence or outputting a response sentence audibly, it determines that the crew member's utterance is an interruption utterance. On the other hand, if the response control unit 132 determines that the crew member's utterance was not acquired while the acquisition unit 131 was generating a response sentence or outputting a response sentence audibly, it determines that the crew member's utterance is not an interruption utterance.

[0101] If the response control unit 132 determines that the occupant's utterance is not an interruption utterance (step S32; No), it terminates processing. On the other hand, if the response control unit 132 determines that the occupant's utterance is an interruption utterance (step S32; Yes), it causes the generation model to determine which response should be prioritized: a response to the occupant's utterance or a response other than a response to the occupant's utterance (step S33). For example, the response control unit 132 inputs speech information and dialogue history information corresponding to the occupant's utterance into a multimodal large-scale language model and causes the multimodal large-scale language model to determine which response should be prioritized: a response to the occupant's utterance or a response other than a response to the occupant's utterance. For example, the response control unit 132 inputs speech information and dialogue history information corresponding to the occupant's utterance into a multimodal large-scale language model and generates information indicating which response should be prioritized: a response to the occupant's utterance or a response other than a response to the occupant's utterance. For example, the response control unit 132 inputs a prompt to the multimodal large-scale language model instructing it to determine which response should be prioritized: a response to a crew member's utterance or a response other than a response to a crew member's utterance. The response control unit 132 then generates information indicating which response should be prioritized. The response control unit 132 obtains information indicating which response should be prioritized: a response to a crew member's utterance or a response other than a response to a crew member's utterance.

[0102] Furthermore, the response control unit 132 determines whether or not it has been determined that a response to a crew member's utterance should be prioritized (step S34). For example, if the response control unit 132 obtains information indicating that a response to a crew member's utterance should be prioritized, it determines that a response to a crew member's utterance should be prioritized. On the other hand, if the response control unit 132 obtains information indicating that a response other than a response to a crew member's utterance should be prioritized, it determines that a response other than a response to a crew member's utterance should be prioritized.

[0103] If the response control unit 132 determines that it has not determined that a response to the crew member's utterance should be prioritized (step S34; No), it terminates processing. On the other hand, if the response control unit 132 determines that it has determined that a response to the crew member's utterance should be prioritized (step S34; Yes), it interrupts responses other than those to the crew member's utterance and generates instruction information instructing the system to generate response information corresponding to the response to the crew member's utterance (step S35).

[0104] The response generation unit 133 generates response information based on the instruction information (step S36). For example, the response generation unit 133 inputs the instruction information generated by the response control unit 132 into a multimodal large-scale language model to generate response information. For example, the response generation unit 133 inputs the instruction information and dialogue history information into a multimodal large-scale language model to generate response information.

[0105] Figure 14 is a diagram illustrating the method for determining an interrupt utterance according to the third embodiment. In Figure 14, the acquisition unit 131 acquires voice information corresponding to the crew member's first utterance between times t1 and t2. The response generation unit 133 generates a response statement indicating a response to the first utterance between times t2 and t3. The output control unit 134 outputs the response statement indicating a response to the first utterance by voice between times t3 and t4. The response control unit 132 determines that the crew member's second utterance is an interrupt utterance if it acquires voice information corresponding to the crew member's second utterance between times t2 and t4. In other words, if the response control unit 132 acquires voice information corresponding to the crew member's second utterance while generating a response statement indicating a response to the first utterance or while outputting a response statement indicating a response to the first utterance by voice, it determines that the crew member's second utterance is an interrupt utterance.

[0106] As described above, the acquisition unit 131 acquires voice information corresponding to the occupant's utterance as environmental information, and acquires response state information, which is information about the current response state of the dialogue agent, as state information. The response control unit 132 determines whether the occupant's utterance is an interruption utterance based on the response state information and voice information, and if it determines that the occupant's utterance is an interruption utterance, it causes the generation model to determine which response should be prioritized: a response to the occupant's utterance or a response other than a response to the occupant's utterance. As a result, the information processing system 1 can determine which response should be prioritized and then provide an appropriate response.

[0107] Furthermore, the response control unit 132 determines whether the acquisition unit 131 has acquired voice information while outputting audio corresponding to a predetermined utterance. If it determines that the acquisition unit 131 has acquired voice information while outputting audio corresponding to a predetermined utterance, it determines that the occupant's utterance is an interruption utterance. This allows the information processing system 1 to appropriately determine whether or not the occupant's utterance is an interruption utterance.

[0108] Furthermore, if the response control unit 132 determines that a response to a crew member's utterance should be prioritized, it generates instruction information that instructs the system to interrupt responses other than those to the crew member's utterance and to generate response information corresponding to the response to the crew member's utterance. The response generation unit 133 generates response information based on the instruction information. As a result, the information processing system 1 can provide an appropriate response when a crew member's utterance is an interruption utterance and the response to the interruption utterance should be prioritized.

[0109] <1-4. Fourth Embodiment> For example, in a visually-enabled dialogue agent, if the internal knowledge held by the information processing system is insufficient to provide an accurate response to the occupant's utterance, it may be necessary to determine which search method to use to acquire external knowledge. Furthermore, image retrieval is time-consuming and costly, so it is desirable to avoid using image retrieval unless absolutely necessary. In contrast, in the fourth embodiment, the information processing system inputs the occupant's utterance, the history held by the information processing system, and images corresponding to the currently visible scenery into a multimodal large-scale language model, and determines a search method to be implemented by setting up multiple conditional branches. This enables the information processing system to provide an accurate response to the occupant's utterance. In addition, the information processing system can avoid unnecessary searches, thereby reducing time and financial costs.

[0110] <1-4-1. Overview of Information Processing> Figures 15-16 illustrate the overview of the processing according to the fourth embodiment. In Figure 15, the acquisition unit 131 acquires an image G41 relating to the environment around the vehicle. The acquisition unit 131 also acquires voice information corresponding to the occupant's speech. For example, the acquisition unit 131 acquires question information indicating the occupant's question as voice information. For example, the acquisition unit 131 acquires question information indicating the occupant's question, "Are there any good restaurants around here?" The acquisition unit 131 also refers to the state information storage unit 122 to acquire dialogue history information indicating the history of the dialogue between the occupant and the dialogue agent, and search history information indicating the history of searches by the dialogue agent. The response generation unit 133 determines whether it is possible to generate response information to the question "Are there any good restaurants around here?" based on the image G41, the question information, the dialogue history information, and the search history information. In Figure 15, the response generation unit 133 determines that it is impossible to generate response information to the question based on image G41, question information, dialogue history information, and search history information. If the response generation unit 133 determines that it is impossible to generate response information to the question, it then determines whether or not the image G41 contains an object of interest to the crew member. In Figure 15, the response generation unit 133 determines that the image G41 does not contain an object of interest to the crew member. If the response generation unit 133 determines that the image G41 does not contain an object of interest to the crew member, it decides to perform a text-based search. For example, the response generation unit 133 decides to perform a web search.

[0111] In Figure 16, the acquisition unit 131 acquires an image G42 of the environment surrounding the vehicle. The acquisition unit 131 also acquires question information indicating the occupant's question, "What's that on the right?" The acquisition unit 131 also refers to the state information storage unit 122 to acquire dialogue history information and search history information. The response generation unit 133 then determines whether it is possible to generate response information to the question "What's that on the right?" based on the image G42, the question information, the dialogue history information, and the search history information. In Figure 16, the response generation unit 133 determines that it is impossible to generate response information to the question based on the image G42, the question information, the dialogue history information, and the search history information. If the response generation unit 133 determines that it is impossible to generate response information to the question, it then determines whether the image G42 contains an object of interest to the occupant. In Figure 16, the response generation unit 133 determines that the image G42 contains building O4, which is an object of interest to the occupant. Furthermore, if the response generation unit 133 determines that the image G42 contains building O4, which is of interest to the occupants, it determines whether it is necessary to identify the proper noun that represents building O4. In Figure 16, the response generation unit 133 determines that it is necessary to identify the proper noun that represents building O4, which is of interest to the occupants. If the response generation unit 133 determines that it is necessary to identify the proper noun that represents building O4, which is of interest to the occupants, it decides to perform at least an image search. For example, if the response generation unit 133 determines that it is necessary to identify the proper noun that represents building O4, which is of interest to the occupants, it decides to perform an image search, a map search, and a web search.

[0112] <1-4-2. Information Processing Procedure> Figure 17 is a flowchart showing an example of information processing according to the fourth embodiment. In Figure 17, the acquisition unit 131 acquires an image of the environment surrounding the moving object, question information indicating the occupant's questions, dialogue history information indicating the history of the dialogue between the occupant and the dialogue agent, and search history information indicating the history of searches by the dialogue agent (step S41).

[0113] Furthermore, the response generation unit 133 determines whether or not it is possible to generate response information to a question based on the image, question information, dialogue history information, and search history information (step S42). For example, the response generation unit 133 inputs the image, question information, dialogue history information, and search history information into a multimodal large-scale language model and causes the multimodal large-scale language model to determine whether or not it is possible to generate response information to a question. The response generation unit 133 inputs the image, question information, dialogue history information, and search history information into a multimodal large-scale language model and generates information indicating whether or not it is possible to generate response information to a question. For example, the response generation unit 133 inputs a prompt into the multimodal large-scale language model instructing it to determine whether or not it is possible to generate response information to a question based on the image, question information, dialogue history information, and search history information, and generates information indicating whether or not it is possible to generate response information to a question. The response generation unit 133 also obtains information indicating whether or not it is possible to generate response information to a question. For example, if the response generation unit 133 obtains information indicating that it is possible to generate response information to a question, it determines that it is possible to generate response information to a question. On the other hand, if the response generation unit 133 obtains information indicating that it is impossible to generate response information to a question, it determines that it is impossible to generate response information to a question.

[0114] If the response generation unit 133 determines that it is possible to generate response information to the question (step S42; Yes), it decides not to perform the search (step S43). On the other hand, if the response generation unit 133 determines that it is impossible to generate response information to the question (step S42; No), it determines whether or not the image contains an object of interest to the occupant (step S44). For example, the response generation unit 133 inputs the image and the question information into a multimodal large-scale language model to have the multimodal large-scale language model determine whether or not the image contains an object of interest to the occupant. The response generation unit 133 inputs the image and the question information into a multimodal large-scale language model to generate information indicating whether or not the image contains an object of interest to the occupant. For example, the response generation unit 133 inputs a prompt into the multimodal large-scale language model instructing it to determine whether or not the image contains an object of interest to the occupant, and generates information indicating whether or not the image contains an object of interest to the occupant. The response generation unit 133 also obtains information indicating whether or not the image contains an object of interest to the occupant. For example, if the response generation unit 133 obtains information indicating that the image contains an object of interest to the occupant, it determines that the image contains an object of interest to the occupant. On the other hand, if the response generation unit 133 obtains information indicating that the image does not contain an object of interest to the occupant, it determines that the image does not contain an object of interest to the occupant.

[0115] If the response generation unit 133 determines that the image does not contain the occupant's object of interest (step S44; No), it decides to perform a web search (step S45). If the response generation unit 133 decides to perform a web search, it performs the web search. On the other hand, if the response generation unit 133 determines that the image contains the occupant's object of interest (step S44; Yes), it determines whether or not it is necessary to identify the proper noun that indicates the occupant's object of interest (step S46). For example, the response generation unit 133 inputs the image and the question information into a multimodal large-scale language model and has the multimodal large-scale language model determine whether or not it is necessary to identify the proper noun that indicates the occupant's object of interest. The response generation unit 133 inputs the image and the question information into a multimodal large-scale language model and generates information indicating whether or not it is necessary to identify the proper noun that indicates the occupant's object of interest. For example, the response generation unit 133 inputs a prompt to the multimodal large-scale language model instructing it to determine whether or not it is necessary to identify a proper noun that represents the crew's object of interest, and generates information indicating whether or not it is necessary to identify a proper noun that represents the crew's object of interest. The response generation unit 133 also acquires information indicating whether or not it is necessary to identify a proper noun that represents the crew's object of interest. For example, if the response generation unit 133 acquires information indicating that it is necessary to identify a proper noun that represents the crew's object of interest, it determines that it is necessary to identify a proper noun that represents the crew's object of interest. On the other hand, if the response generation unit 133 acquires information indicating that it is not necessary to identify a proper noun that represents the crew's object of interest, it determines that it is not necessary to identify a proper noun that represents the crew's object of interest.

[0116] If the response generation unit 133 determines that it is not necessary to identify a proper noun indicating the crew's interest (step S46; No), it decides to perform a web search (step S45). If the response generation unit 133 decides to perform a web search, it performs a web search. On the other hand, if the response generation unit 133 determines that it is necessary to identify a proper noun indicating the crew's interest (step S46; Yes), it decides to perform an image search, a map search, and a web search (step S47). If the response generation unit 133 decides to perform an image search, a map search, and a web search, it performs an image search, a map search, and a web search.

[0117] As described above, the acquisition unit 131 acquires environmental information, including images of the environment surrounding the moving object and question information indicating the occupant's questions. It also acquires dialogue history information indicating the history of conversations between the occupant and the dialogue agent, and search history information indicating the history of searches by the dialogue agent. The response generation unit 133 determines whether it is possible to generate response information to the question based on the images, question information, dialogue history information, and search history information. This enables the information processing system 1 to provide accurate responses to the occupant's speech. Furthermore, the information processing system 1 can avoid unnecessary searches, thereby reducing time and financial costs.

[0118] Furthermore, if the response generation unit 133 determines that it is impossible to generate response information to a question, it determines whether or not the image contains an object of interest to the occupant, and based on the determination result of whether or not the image contains an object of interest to the occupant, it selects the optimal search method from among multiple search methods and decides to execute a search using the selected search method. In this way, the information processing system 1 can select the optimal search method from among multiple search methods based on the determination result of whether or not the image contains an object of interest to the occupant.

[0119] Furthermore, if the response generation unit 133 determines that the image does not contain the occupant's object of interest, it decides to perform a text-based search. This allows the information processing system 1 to decide to perform a text-based search (e.g., a web search) if it determines that the image does not contain the occupant's object of interest, thereby reducing time and financial costs.

[0120] Furthermore, if the response generation unit 133 determines that the image contains an object of interest to the occupant, it determines whether or not it is necessary to identify the proper noun that represents the object of interest to the occupant. Based on the determination result of whether or not it is necessary to identify the proper noun that represents the object of interest to the occupant, it selects the most suitable search method from among multiple search methods and decides to execute the search using the selected search method. In this way, the information processing system 1 can select the most suitable search method from among multiple search methods based on the determination result of whether or not it is necessary to identify the proper noun that represents the object of interest to the occupant.

[0121] Furthermore, if the response generation unit 133 determines that it is not necessary to identify proper nouns, it decides to perform a text-based search. This allows the information processing system 1 to decide to perform a text-based search when it determines that it is not necessary to identify proper nouns, thereby reducing time and monetary costs.

[0122] Furthermore, if the response generation unit 133 determines that it is necessary to identify a proper noun, it decides to perform at least an image search. This allows the information processing system 1 to perform the image search, which is the most time- and financially costly, only when it determines that it is necessary to identify a proper noun.

[0123] <1-5. Fifth Embodiment> For example, a visually-enabled dialogue agent may perform a map search to generate an accurate response to a passenger's utterance. In this case, it is desirable to perform preprocessing to enable an effective map search for the passenger's object of interest. In this fifth embodiment, the information processing system uses a multimodal large-scale language model to generate information indicating a general name representing the passenger's object of interest, as well as the distance and direction from the vehicle to the object of interest, as preprocessing for performing the map search. The information processing system then narrows the search range in the map search based on the information indicating the distance and direction from the vehicle to the object of interest, and then performs a map search using the general name representing the passenger's object of interest as a query. This allows the information processing system to perform the map search efficiently. Furthermore, the information processing system can provide an appropriate response to the passenger's utterance based on the search results of the map search.

[0124] <1-5-1. Overview of Information Processing> Figure 18 is a diagram illustrating the overview of the processing according to the fifth embodiment. In Figure 18, the response generation unit 133 generates a response sentence indicating a response to the occupant's question utterance, "What is that on the right?". For example, the response generation unit 133 generates a response sentence indicating a response to the occupant's question, "I'll look into the tower on the right." The response generation unit 133 also decides to perform a map search when it has generated a response sentence. The acquisition unit 131 acquires an image G51 showing a tower on the right side of the road. The acquisition unit 131 also acquires dialogue history information including the occupant's utterance, "What is that on the right?", and the dialogue agent's response, "I'll look into the tower on the right." The response generation unit 133 inputs the image G51 and dialogue history information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate general name information indicating the general name of the tower that is the object of interest to the occupant contained in image G51, distance information indicating the distance from the vehicle to the tower, and direction information indicating the direction of the tower relative to the direction of travel of the vehicle. For example, the response generation unit 133 generates the string "name:tower" as general name information, indicating the general name of the tower that is the object of interest to the occupant contained in image G51. The response generation unit 133 also generates the string "distance:800m" as distance information, indicating the distance from the vehicle to the tower. The response generation unit 133 also generates the string "direction:30.0°" as direction information, indicating the direction of the tower relative to the direction of travel of the vehicle.

[0125] Figure 19 is a diagram illustrating an example of information processing according to the fifth embodiment. In Figure 19, following Figure 18, the response generation unit 133 decides to perform a map search using the general name "tower" within a predetermined range A51 centered on point P51, which is 800m away from the vehicle and at an angle of 30° from the vehicle's current location relative to the direction of travel of the tower. The response generation unit 133 also performs a map search using the general name "tower" within a predetermined range centered on point 800m away from the vehicle and at an angle of 30° from the vehicle's current location relative to the direction of travel of the tower relative to the direction of travel of the vehicle.

[0126] Figure 20 is a diagram illustrating an example of information processing according to the fifth embodiment. In Figure 20, the response generation unit 133 adjusts the radius indicating the map search range to include the rear of the vehicle, for example, if the distance to the target is short, taking into consideration the possibility that the target has been passed. In Figure 20, the response generation unit 133 performs a map search using the general name of the target in the area A52 that includes the rear of the vehicle's current location on the map M52 which includes the vehicle's current location.

[0127] Figure 21 is a diagram illustrating an example of information processing according to the fifth embodiment. In Figure 21, the response generation unit 133 performs a map search by changing the distance multiple times relative to the estimation direction, for example, when the distance to the target is far, because it is difficult to estimate the distance to the target. In Figure 21, the response generation unit 133 performs a map search using the general name of the target in each of the predetermined ranges A53, A54, and A55, centered on points P53, P54, and P55, respectively, on a map M53 that includes the vehicle's current location, starting from the point closest to the vehicle.

[0128] <1-5-2. Information Processing Procedure> Figure 22 is a flowchart showing an example of information processing according to the fifth embodiment. In Figure 22, the acquisition unit 131 acquires an image of the environment surrounding the moving object and dialogue history information showing the history of the conversation between the occupant and the dialogue agent (step S51).

[0129] Furthermore, the response generation unit 133 inputs the image and dialogue history information into a multimodal large-scale language model to generate general name information indicating a general name of an object of interest of the occupant included in the image, distance information indicating the distance from the moving object to the object, and direction information indicating the direction of the object relative to the direction of movement of the moving object (step S52). For example, the response generation unit 133 inputs the image and dialogue history information into a multimodal large-scale language model to generate general name information indicating a general name of an object of interest of the occupant included in the image, distance information indicating the distance from the moving object to the object, and direction information indicating the direction of the object relative to the direction of movement of the moving object. For example, the response generation unit 133 determines the object of interest of the occupant included in the image, and inputs a prompt into the multimodal large-scale language model instructing it to generate general name information indicating a general name of the determined object, distance information indicating the distance from the moving object to the object, and direction information indicating the direction of the object relative to the direction of movement of the moving object, thereby generating general name information, distance information, and direction information.

[0130] Furthermore, the response generation unit 133 decides to perform a map search based on the general name information, distance information, and direction information (step S53). If the response generation unit 133 decides to perform a map search based on the general name information, distance information, and direction information, it performs a map search based on the general name information, distance information, and direction information. For example, the response generation unit 133 performs a map search using the general name information as a query within a search range based on the distance information and direction information on the map.

[0131] As described above, the acquisition unit 131 acquires images of the environment surrounding the moving object as environmental information, and acquires dialogue history information showing the history of the conversation between the occupant and the dialogue agent. The response generation unit 133 inputs the images and the dialogue history information into a multimodal large-scale language model to generate general name information indicating the general name of the object of interest of the occupant contained in the image, distance information indicating the distance from the moving object to the object, and direction information indicating the direction of the object relative to the direction of movement of the moving object. The unit then decides to perform a map search based on the general name information, distance information, and direction information. As a result, the information processing system 1 can efficiently perform a map search based on the general name information, distance information, and direction information. Furthermore, the information processing system can provide an appropriate response to the occupant's speech based on the search results of the map search.

[0132] Furthermore, the response generation unit 133 decides to perform a map search using general name information within a predetermined range centered on a point that is the same distance from the moving object to the target in the direction of the target relative to the moving object's direction of travel, from the moving object's current location. This allows the information processing system 1 to efficiently perform a map search by narrowing the search range to a predetermined range centered on a point that is the same distance from the moving object to the target.

[0133] Furthermore, the response generation unit 133 generates multiple distance information representing each of multiple distances, and decides to perform a map search using general name information within multiple predetermined ranges centered on each of multiple points that are each a distance away from the current location of the moving object in the direction of the object relative to the direction of the moving object's movement. As a result, the information processing system 1 can efficiently perform a map search by performing a map search within multiple search ranges, even when the distance to the object is far and it is difficult to estimate the distance.

[0134] <1-6. Sixth Embodiment> For example, a visually-enabled dialogue agent may perform image retrieval to generate accurate responses to a crew member's speech. In such cases, it is desirable to perform preprocessing to enable effective image retrieval of the crew member's interests. For instance, if an image retrieval is performed using the entire image as the query, it may be difficult to obtain information about the crew member's interests as a search result because the image contains many objects. In contrast, in the sixth embodiment, the information processing system uses a multimodal large-scale language model and an object detection model to perform image retrieval by preprocessing the image to identify the region showing the crew member's interests, and then performing an image retrieval using the region mainly containing the crew member's interests as the query. This allows the information processing system to perform image retrieval efficiently. Furthermore, the information processing system can provide appropriate responses to the crew member's speech based on the search results of the image retrieval.

[0135] <1-6-1. Overview of Information Processing> Figures 23-24 are diagrams illustrating the overview of the processing according to the sixth embodiment. In Figure 23, the response generation unit 133 generates a response sentence indicating a response to the occupant's question utterance, "What's that on the right?". For example, the response generation unit 133 generates a response sentence indicating a response to the occupant's question, "I'll look into the tower on the right." The response generation unit 133 also decides to perform an image search once it has generated a response sentence. The acquisition unit 131 acquires an image G61 showing a tower on the right side of the road. The acquisition unit 131 also acquires dialogue history information including the occupant's utterance, "What's that on the right?", and the dialogue agent's response, "I'll look into the tower on the right." The response generation unit 133 inputs the image G61 and dialogue history information acquired by the acquisition unit 131 into a multimodal large-scale language model to generate general name information indicating the general name of the tower that is the subject of interest to the crew and included in the image G61, and region information relating to region O6 in the image G61 that indicates the tower that is the subject of interest to the crew. For example, the response generation unit 133 generates the string "name:Tower" as general name information, indicating the general name of the tower that is the subject of interest to the crew and included in the image G61. The response generation unit 133 also generates the string "Box:[2800,700,2900,900]" as region information, indicating the location of region O6 in the image G61 that indicates the tower that is the subject of interest to the crew. For example, the response generation unit 133 generates the string indicating the coordinates that indicate the location of region O6, which is a bounding box, as region information.

[0136] Furthermore, in Figure 24, the response generation unit 133 inputs the string "name:Tower," which indicates the general name of the tower that is the object of interest to the occupants, contained in image G62 (the same image as image G61 shown in Figure 23), and image G62 into the object detection model to generate the precise bounding box of the tower that is the object of interest to the occupants in image G62. The response generation unit 133 generates new region information for a new region obtained by deleting the region of image G62 that represents the tower that is the object of interest to the occupants, excluding the tower itself. For example, the response generation unit 133 generates the string "Box:[2828,710,2897,866]" as new region information, which indicates the position of the precise bounding box of the tower that is the object of interest to the occupants in image G62.

[0137] <1-6-2. Information Processing Procedure> Figure 25 is a flowchart showing an example of information processing according to the sixth embodiment. In Figure 25, the acquisition unit 131 acquires an image of the environment surrounding the moving object and dialogue history information showing the history of the conversation between the occupant and the dialogue agent (step S61).

[0138] Furthermore, the response generation unit 133 inputs the image and dialogue history information into a multimodal large-scale language model to generate general name information indicating a general name that represents the occupant's object of interest included in the image, and region information relating to the region in the image that represents the object (step S62). For example, the response generation unit 133 inputs the image and dialogue history information into a multimodal large-scale language model to generate general name information indicating a general name that represents the occupant's object of interest included in the image, and region information relating to the region in the image that represents the object. For example, the response generation unit 133 determines the occupant's object of interest included in the image, and inputs a prompt into the multimodal large-scale language model instructing it to generate general name information indicating a general name that represents the determined object, and region information relating to the region in the image that represents the object, thereby generating general name information and region information.

[0139] Furthermore, the response generation unit 133 decides to perform an image search based on the general name information and the region information (step S63). If the response generation unit 133 decides to perform an image search based on the general name information and the region information, it performs an image search based on the general name information and the region information. For example, the response generation unit 133 performs an image search using the region information as a query.

[0140] As described above, the acquisition unit 131 acquires images of the environment surrounding the moving object as environmental information, and acquires dialogue history information showing the history of the conversation between the occupant and the dialogue agent. The response generation unit 133 inputs the images and the dialogue history information into a multimodal large-scale language model to generate general name information indicating the general name of the object of interest of the occupant contained in the image, and region information relating to the region of the image that indicates the object, and decides to perform an image search based on the general name information and region information. As a result, the information processing system 1 can efficiently perform an image search based on the general name information and region information. Furthermore, the information processing system can provide an appropriate response to the occupant's speech based on the search results of the image search.

[0141] Furthermore, the response generation unit 133 generates new region information for a new region obtained by deleting the region of the image that represents the target, and decides to perform an image search based on the new region information. As a result, the information processing system 1 can improve the search accuracy of the image search by performing an image search based on the new region information.

[0142] <1-7. Seventh Embodiment> For example, a visually-enabled dialogue agent may simultaneously perform searches using multiple search methods to generate accurate responses to the occupant's speech (hereinafter sometimes referred to as parallel search). In such cases, each search method may search for different targets. In contrast, in the seventh embodiment, the information processing system provides a specific initial response about what to investigate before starting the parallel search, adds this to the dialogue history, and then inputs the search queries for each search method when generating them, thereby aligning the search targets in the parallel search. This allows the information processing system to perform parallel searches efficiently.

[0143] <1-7-1. Overview of Information Processing> Figures 26-27 illustrate the overview of the processing according to the seventh embodiment. In Figure 26, the acquisition unit 131 acquires image G71, which shows a red vehicle on the road and a tower on the right side of the road. The acquisition unit 131 also acquires dialogue history information, including the occupant's utterance, "What's that on the right?" The response generation unit 133 then decides to input image G71 and the dialogue history information, including the occupant's utterance "What's that on the right?", acquired by the acquisition unit 131 into a multimodal large-scale language and perform image search, map search, and web search. For example, the response generation unit 133 performs an image search using a cropped area of ​​the red car contained in image G71 as the query. The response generation unit 133 also performs a map search using the keyword "tower" as the query. The response generation unit 133 also performs a web search using the keyword "Shibukawa City tall tower" as the query. As shown in Figure 26, image search, map search, and web search each examine different search targets, making it impossible to perform parallel searches efficiently.

[0144] In Figure 27, the acquisition unit 131 acquires image G72, which is the same image G71 described in Figure 26. The acquisition unit 131 also acquires dialogue history information, including the occupant's utterance, "What's that on the right?". The response generation unit 133 inputs image G71 and the dialogue history information, including the occupant's utterance "What's that on the right?", acquired by the acquisition unit 131, into a multimodal large-scale language to identify the object of the occupant's interest and perform a search related to that object. In Figure 27, the response generation unit 133 identifies the object of the occupant's interest contained in image G72 as tower O7, which is pictured on the right side of the road. The response generation unit 133 also generates a response sentence indicating that it will perform a search related to the object indicated by the general name, based on the general name information. For example, the response generation unit 133 generates the response sentence, "I'll look into that tall tower on the right." The output control unit 134 outputs the generated response sentence as speech. Furthermore, the acquisition unit 131 acquires dialogue history information including the response sentence, "I'll look into that tall tower on the right." The response generation unit 133 then inputs the image G72 acquired by the acquisition unit 131 and the dialogue history information including the response sentence, "I'll look into that tall tower on the right," into a multimodal large-scale language and decides to perform image search, map search, and web search. For example, the response generation unit 133 performs an image search using a cropped area of ​​image G72 showing tower O7 as the query. The response generation unit 133 also performs a map search using the keyword "tower" as the query. The response generation unit 133 also performs a web search using the keywords "Shibukawa City tall tower" as the query. In this way, as shown in Figure 27, since each of the image search, map search, and web search searches the same target, parallel searches can be performed efficiently.

[0145] <1-7-2. Information Processing Procedures> Figure 28 is a flowchart showing an example of information processing according to the seventh embodiment. In Figure 28, the acquisition unit 131 acquires an image of the environment surrounding the moving object and dialogue history information showing the history of the conversation between the occupant and the dialogue agent (step S71).

[0146] Furthermore, the response generation unit 133 inputs the image and dialogue history information into a multimodal large-scale language model to identify the occupant's object of interest and generates response information indicating that a search related to the occupant's object of interest will be performed (step S72). For example, the response generation unit 133 inputs a prompt into the multimodal large-scale language model instructing it to identify the occupant's object of interest, identify a general name representing the identified object, and generate a response statement indicating that a search related to the identified general name will be performed, thereby generating a response statement indicating that a search related to the occupant's object of interest will be performed.

[0147] Furthermore, the response generation unit 133 decides to execute a search using multiple search methods based on the dialogue history information including the response information (step S73). For example, the response generation unit 133 inputs the dialogue history information, which includes a response statement indicating that it will perform a search on the object indicated by the image and general name acquired by the acquisition unit 131, into the multimodal large-scale language model and generates a query for each of the multiple search methods. Furthermore, the response generation unit 133 decides to execute each of the multiple search methods based on the query for each of the multiple search methods.

[0148] As described above, the acquisition unit 131 acquires images of the environment surrounding the moving object as environmental information, and acquires dialogue history information showing the history of the conversation between the occupant and the dialogue agent. The response generation unit 133 inputs the images and the dialogue history information into a multimodal large-scale language model to generate general name information showing general names indicating the objects of interest of the occupant contained in the images, generates response information indicating that a search will be performed on the objects indicated by the general names based on the general name information, and decides to perform a search using multiple search methods based on the dialogue history information including the response information. As a result, the information processing system 1 can align the search targets in parallel searches, and thus can perform parallel searches efficiently.

[0149] <1-8. About Generative Models> Furthermore, for each of the generative models used in the processes described above, the internal structure can be any structure, not limited to the examples described in each section, as long as it can output the desired information in response to the input. Any combination of input, output, and internal structure of the generative model can be adopted, as long as it can output the desired information.

[0150] The input to the generative model may be text, images, audio, or a combination thereof. The output of the generative model may also be text, images, audio, or other similar formats. Note that the above-mentioned inputs and outputs are merely examples, and the generative model itself may use any input and output.

[0151] Furthermore, the internal structure of the generative model can be any structure depending on the combination of inputs and outputs. In other words, the internal structure of the generative model can be any structure as long as it allows for the desired output for a given input.

[0152] For example, a generative model may have a structure related to a Transformer. For example, a generative model may have a structure related to a Transformer and perform processing that takes into account the context, such as the relative positions within the data, such as text or time-series data. For example, a generative model may have a self-attention mechanism. For example, a generative model may have any attention mechanism such as Single-Head Attention or Multi-Head Attention. Note that a generative model does not have to have an attention mechanism.

[0153] The generative model may have a mechanism for extracting features from the input. For example, the generative model may have an encoder. The generative model may have a mechanism for generating information based on the extracted features. For example, the generative model may have a decoder.

[0154] The generative model may have a structure related to a Convolutional Neural Network (CNN). For example, when performing image processing, the generative model may have a structure related to a CNN. For example, the generative model may have at least one of the following: a convolutional layer, a pooling layer, a fully connected layer, etc.

[0155] The internal structure described above is merely an example, and the generative model described above may have any internal structure. For example, the generative model may have skip connections. Furthermore, the generative model may have a structure related to the diffusion model.

[0156] Furthermore, the generative models described above may be generated (trained) by any learning process. The generative models may also be machine learning models trained using any machine learning technique. For example, a generative model may be a model generated by fine-tuning a so-called Foundation Model to apply that Foundation Model to a specific task (e.g., prompt generation). For example, a generative model like the LLM described above may be a model generated by fine-tuning a Foundation Model to apply it to a specific task.

[0157] The foundational model referred to here is a model that has been trained to be applicable to various tasks, for example, to be able to perform a wide variety of tasks. For example, the foundational model is a neural network pre-trained on a large dataset of unlabeled data. The foundational model may have any structure, such as a Transformer-based architecture. For example, the foundational model is generated by self-supervised learning using data without correct labels. As described above, the foundational model is fine-tuned to be adaptable to a wide range of downstream tasks.

[0158] For example, when applied to a dialogue task, the underlying model is fine-tuned to adapt to the dialogue task, and a generative model adapted to the dialogue task is generated.

[0159] The generative model described above is trained by any learning process depending on the input, output, and internal structure of the generative model. For example, the generative model may be trained using unsupervised learning methods such as GAN (Generative Adversarial Network). Alternatively, the generative model may be trained in a distributed state without aggregating the data, such as in federated learning. In this case, each information processing service device (server, etc.) may generate a local model collected by that service, and a server (aggregation server) that aggregates the information (parameters, etc.) of the local models generated by each information processing service device (server, etc.) may generate a global model using the local model information. In this case, the information processing system 1 may receive the global model generated by the aggregation server from the aggregation server and use the received global model as the generative model for processing.

[0160] Thus, the generative models described above can be generated (learned) by any computer. In other words, the learning process for generating the generative models may be performed by any device (computer, etc.) within the information processing system 1, or by a device outside the information processing system 1. For example, if a device outside the information processing system 1 generates at least one of the generative models described above, the information processing system 1 retrieves that generative model from the device outside the information processing system 1 and processes the data using the retrieved generative model.

[0161] <2. Others> The processes described in each embodiment above may be carried out in various different forms (modifications) other than the embodiments and modifications described above. Furthermore, all or part of the processes described as being performed automatically in each embodiment may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0162] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0163] Furthermore, the embodiments and modifications described above can be combined as appropriate, provided that the processing content is not inconsistent.

[0164] Furthermore, the effects described herein are merely illustrative and not limiting; other effects may also occur.

[0165] <3. Hardware Configuration> The information processing device (terminal device 10 and information processing device 100, etc.) of the information processing system 1 according to the first embodiment described above is realized by a computer 1000 having a configuration such as that shown in Figure 29. The information processing device 100 will be used as an example for explanation. Figure 29 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. The computer 1000 has a processing circuitry 1100, RAM 1200, ROM 1300, secondary storage device 1400, communication interface 1500, input / output interface 1600, display unit 1700, camera unit 1800, microphone 1900, and speaker 2000. The various parts of the computer 1000 are connected by a bus 1050.

[0166] The processing circuit 1100 operates based on a program stored in the ROM 1300 or secondary storage device 1400, and controls each part. For example, the processing circuit 1100 loads the program stored in the ROM 1300 or secondary storage device 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0167] ROM1300 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the processing circuit 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0168] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily stores programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that stores programs for each process of the information processing device 100 according to the first embodiment, which is an example of program data 1450.

[0169] The communication interface 1500 is an interface for the computer 1000 to connect to the external network 1550. The communication interface 1500 corresponds to the communication unit 110 of the information processing device 100 (or the communication unit 11 of the terminal device 10). For example, the processing circuit 1100 receives data from other devices or transmits data generated by the processing circuit 1100 to other devices via the communication interface 1500.

[0170] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from input devices such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to output devices such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.

[0171] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescent display (Organic Electro Luminescence Display). Alternatively, the display unit 1700 may be a touch panel display device or an image projection device.

[0172] The camera unit 1800 is an interface for the computer 1000 to capture images. The microphone 1900 is an interface for the computer 1000 to capture audio. The speaker 2000 is an interface for the computer 1000 to output the processed audio. The various parts of the computer 1000 are connected by the bus 1050. Each interface does not necessarily have to be located inside the computer 1000; it may be located outside the computer 1000 via a network or the like. Furthermore, each part of the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100, but by a circuit dedicated to display processing that is provided within the display unit 1700.

[0173] For example, when computer 1000 functions as an information processing device 100 (or terminal device 10) according to the first embodiment, the processing circuit 1100 of computer 1000 functions as a control unit 130 (or control unit 13) by executing a program loaded onto RAM 1200. The secondary storage device 1400 stores the information processing program according to this disclosure and various data stored by the storage unit 120. The processing circuit 1100 reads and executes the program data 1450 from the secondary storage device 1400, but as another example, these programs may be obtained from other devices via an external network 1550. In other words, the secondary storage device 1400 is not limited to being inside computer 1000, but may be located outside computer 1000. The processing circuit 1100 is an example of an integrated circuit, and CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered integrated circuits.

[0174] Furthermore, this technology can also be configured as follows. (1) An acquisition unit that acquires environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object. A response control unit generates response control information, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupant of the mobile body, based on the environmental information and the state information. A response generation unit that generates the response information based on the response control information, An information processing system equipped with the following features. (2) The acquisition unit is, As the aforementioned state information, response state information, which is information regarding the current response state of the dialogue agent, is obtained. The response control unit, Based on the aforementioned environmental information and the aforementioned response state information, the response control information is generated. (1) The information processing system described above. (3) The acquisition unit is, The system acquires dialogue history information showing the history of the conversation between the crew member and the dialogue agent. The response generation unit, Based on the environmental information and the dialogue history information, the response information is generated with a modified designation indicating the object mentioned in the dialogue. The information processing system described in (1) or (2). (4) The response generation unit, Based on the environmental information and the dialogue history information, the system identifies the object mentioned in the dialogue from among the objects identified based on the environmental information, tracks the object based on the environmental information, and if tracking the object becomes impossible, generates the response information with a changed designation for the object. (3) The information processing system described above. (5) The acquisition unit is, As environmental information, images of the environment surrounding the moving object are acquired. The response generation unit, Based on the aforementioned image and the aforementioned dialogue history information, the response information is generated with a modified designation indicating the object included in the aforementioned image. (3) or (4) the information processing system described above. (6) The response generation unit, Based on the image and the dialogue history information, the system identifies the object mentioned in the dialogue from among the objects included in the image, tracks the object based on the image, and if tracking the object becomes impossible, generates the response information with a changed designation for the object. (5) The information processing system described above. (7) The acquisition unit is, As environmental information, voice information corresponding to the occupant's speech is acquired, and as state information, response state information, which is information regarding the current response state of the dialogue agent, is acquired. The response control unit, Based on the response state information and the voice information, it is determined whether the occupant's utterance is an interruption utterance. If it is determined that the occupant's utterance is an interruption utterance, the generative model is instructed to determine which response should take priority: the response to the occupant's utterance or a response other than the response to the occupant's utterance. An information processing system described in any one of (1) to (6). (8) The response control unit, While outputting audio corresponding to a predetermined utterance, the system determines whether the acquisition unit has acquired the audio information. If it determines that the acquisition unit has acquired the audio information while outputting audio corresponding to the predetermined utterance, it determines that the occupant's utterance is an interruption utterance. (7) The information processing system described above. (9) The response control unit, If it is determined that a response to the crew member's utterance should be prioritized, the system will interrupt all responses other than those to the crew member's utterance and generate instruction information instructing the system to generate response information corresponding to the response to the crew member's utterance. The response generation unit, Based on the instruction information, the response information is generated. The information processing system described in (7) or (8). (10) The acquisition unit is, As environmental information, images relating to the environment surrounding the mobile body and question information indicating the questions asked by the occupant are acquired, dialogue history information indicating the history of the dialogue between the occupant and the dialogue agent and search history information indicating the history of searches performed by the dialogue agent are acquired. The response generation unit, Based on the image, the question information, the dialogue history information, and the search history information, it is determined whether or not it is possible to generate the response information for the question. An information processing system described in any one of (1) to (9). (11) The response generation unit, If it is determined that it is impossible to generate the response information to the question, the system determines whether the image contains the occupant's object of interest, selects the most suitable search method from among multiple search methods based on the result of determining whether the image contains the occupant's object of interest, and decides to perform the search using the selected search method. (10) The information processing system described above. (12) The response generation unit, If it is determined that the image does not contain the object of interest of the crew member, it is decided to perform a text-based search. (11) The information processing system described above. (13) The response generation unit, If it is determined that the image contains an object of interest to the crew member, it is determined whether or not it is necessary to identify the proper noun representing the object of interest to the crew member. Based on the result of this determination, the system selects the most suitable search method from among several search methods and decides to perform the search using the selected search method. The information processing system described in (11) or (12). (14) The response generation unit, If it is determined that it is not necessary to identify the aforementioned proper noun, a text-based search will be performed. (13) The information processing system described above. (15) The response generation unit, If it is determined that it is necessary to identify the aforementioned proper noun, it is decided to at least perform an image search. The information processing system described in (13) or (14). (16) The acquisition unit is, As environmental information, images of the environment surrounding the mobile body are acquired, and dialogue history information showing the history of the dialogue between the occupant and the dialogue agent is acquired. The response generation unit, The image and the dialogue history information are input into a multimodal large-scale language model to generate general name information indicating a general name representing the object of interest of the occupant contained in the image, distance information indicating the distance from the moving object to the object, and direction information indicating the direction of the object relative to the direction of travel of the moving object. The system then decides to perform a map search based on the general name information, the distance information, and the direction information. An information processing system described in any one of (1) to (15). (17) The response generation unit, It is decided to perform the map search using the general name information within a predetermined range centered on a point that is a distance from the moving body to the target in the direction of the target relative to the direction of travel of the moving body, from the current location of the moving body. (16) The information processing system described above. (18) The response generation unit, The system generates multiple distance information representing each of multiple distances, and decides to perform the map search using the general name information within a plurality of predetermined ranges centered on each of the plurality of points that are each a certain distance away from the current location of the moving object in the direction of the object relative to the direction of movement of the moving object. The information processing system described in (16) or (17). (19) The acquisition unit is, As environmental information, images of the environment surrounding the mobile body are acquired, and dialogue history information showing the history of the dialogue between the occupant and the dialogue agent is acquired. The response generation unit, The image and the dialogue history information are input into a multimodal large-scale language model to generate general name information indicating a general name representing the object of interest of the occupant contained in the image, and region information relating to the region of the image that represents the object. The system then decides to perform an image search based on the general name information and the region information. An information processing system described in any one of (1) to (18). (20) The response generation unit, It is decided to generate new region information for a new region obtained by deleting the region of the image that represents the target, excluding the region representing the target, and to perform the image search based on the new region information. (19) The information processing system described above. (twenty one) The acquisition unit is, As environmental information, images of the environment surrounding the mobile body are acquired, and dialogue history information showing the history of the dialogue between the occupant and the dialogue agent is acquired. The response generation unit, The image and the dialogue history information are input into a multimodal large-scale language model to identify the occupant's object of interest, generate response information indicating that a search related to the occupant's object of interest will be performed, and decide to perform a search using multiple search methods based on the dialogue history information including the response information. An information processing system described in any one of (1) to (20). (twenty two) A method of information processing performed by a computer, To acquire environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object, Based on the aforementioned environmental information and the aforementioned state information, response control information is generated, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupant of the mobile body. Based on the response control information, the response information is generated, Information processing methods including (twenty three) To acquire environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object, Based on the aforementioned environmental information and the aforementioned state information, response control information is generated, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupant of the mobile body. Based on the response control information, the response information is generated, An information processing program that causes a computer to execute something. [Explanation of symbols]

[0175] 1. Information Processing System 10 Terminal devices 100 Information Processing Devices 110 Communications Department 120 Storage section 130 Control Unit 131 Acquisition Department 132 Response Control Unit 133 Response generation unit 134 Output Control Unit

Claims

1. An acquisition unit that acquires environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object. A response control unit generates response control information, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupant of the mobile body, based on the environmental information and the state information. A response generation unit that generates the response information based on the response control information, An information processing system equipped with the following features.

2. The acquisition unit is, As the aforementioned state information, response state information, which is information regarding the current response state of the dialogue agent, is obtained. The response control unit, Based on the aforementioned environmental information and the aforementioned response state information, the response control information is generated. The information processing system according to claim 1.

3. The acquisition unit is, The system acquires dialogue history information showing the history of the conversation between the crew member and the dialogue agent. The response generation unit, Based on the environmental information and the dialogue history information, the response information is generated with a modified designation indicating the object mentioned in the dialogue. The information processing system according to claim 1.

4. The response generation unit, Based on the environmental information and the dialogue history information, the system identifies the object mentioned in the dialogue from among the objects identified based on the environmental information, tracks the object based on the environmental information, and if tracking the object becomes impossible, generates the response information with a changed designation for the object. The information processing system according to claim 3.

5. The acquisition unit is, As environmental information, images of the environment surrounding the moving object are acquired. The response generation unit, Based on the aforementioned image and the aforementioned dialogue history information, the response information is generated with a modified designation indicating the object included in the aforementioned image. The information processing system according to claim 3.

6. The response generation unit, Based on the image and the dialogue history information, the system identifies the object mentioned in the dialogue from among the objects included in the image, tracks the object based on the image, and if tracking the object becomes impossible, generates the response information with a changed designation for the object. The information processing system according to claim 5.

7. The acquisition unit is, As environmental information, voice information corresponding to the occupant's speech is acquired, and as state information, response state information, which is information regarding the current response state of the dialogue agent, is acquired. The response control unit, Based on the response state information and the voice information, it is determined whether the occupant's utterance is an interruption utterance. If it is determined that the occupant's utterance is an interruption utterance, the generative model is instructed to determine which response should take priority: the response to the occupant's utterance or a response other than the response to the occupant's utterance. The information processing system according to claim 1.

8. The response control unit, While outputting audio corresponding to a predetermined utterance, the system determines whether the acquisition unit has acquired the audio information. If it determines that the acquisition unit has acquired the audio information while outputting audio corresponding to the predetermined utterance, it determines that the occupant's utterance is an interruption utterance. The information processing system according to claim 7.

9. The response control unit, If it is determined that a response to the crew member's utterance should be prioritized, the system will interrupt all responses other than those to the crew member's utterance and generate instruction information instructing the system to generate response information corresponding to the response to the crew member's utterance. The response generation unit, Based on the instruction information, the response information is generated. The information processing system according to claim 7.

10. The acquisition unit is, As environmental information, images relating to the environment surrounding the mobile body and question information indicating the questions asked by the occupant are acquired, dialogue history information indicating the history of the dialogue between the occupant and the dialogue agent and search history information indicating the history of searches performed by the dialogue agent are acquired. The response generation unit, Based on the image, the question information, the dialogue history information, and the search history information, it is determined whether or not it is possible to generate the response information for the question. The information processing system according to claim 1.

11. The response generation unit, If it is determined that it is impossible to generate the response information to the question, the system determines whether the image contains the occupant's object of interest, selects the most suitable search method from among multiple search methods based on the result of determining whether the image contains the occupant's object of interest, and decides to perform the search using the selected search method. The information processing system according to claim 10.

12. The response generation unit, If it is determined that the image does not contain the object of interest of the crew member, it is decided to perform a text-based search. The information processing system according to claim 11.

13. The acquisition unit is, As environmental information, images of the environment surrounding the mobile body are acquired, and dialogue history information showing the history of the dialogue between the occupant and the dialogue agent is acquired. The response generation unit, The image and the dialogue history information are input into a multimodal large-scale language model to generate general name information indicating a general name representing the object of interest of the occupant contained in the image, distance information indicating the distance from the moving object to the object, and direction information indicating the direction of the object relative to the direction of travel of the moving object. The system then decides to perform a map search based on the general name information, the distance information, and the direction information. The information processing system according to claim 1.

14. The response generation unit, It is decided to perform the map search using the general name information within a predetermined range centered on a point that is a distance from the moving body to the target in the direction of the target relative to the direction of travel of the moving body, from the current location of the moving body. The information processing system according to claim 13.

15. The response generation unit, The system generates multiple distance information representing each of multiple distances, and decides to perform the map search using the general name information within a plurality of predetermined ranges centered on each of the plurality of points that are each a certain distance away from the current location of the moving object in the direction of the object relative to the direction of movement of the moving object. The information processing system according to claim 13.

16. The acquisition unit is, As environmental information, images of the environment surrounding the mobile body are acquired, and dialogue history information showing the history of the dialogue between the occupant and the dialogue agent is acquired. The response generation unit, The image and the dialogue history information are input into a multimodal large-scale language model to generate general name information indicating a general name representing the object of interest of the occupant contained in the image, and region information relating to the region of the image that represents the object. The system then decides to perform an image search based on the general name information and the region information. The information processing system according to claim 1.

17. The response generation unit, It is decided to generate new region information for a new region obtained by deleting the region of the image that represents the target, excluding the region representing the target, and to perform the image search based on the new region information. The information processing system according to claim 16.

18. The acquisition unit is, As environmental information, images of the environment surrounding the mobile body are acquired, and dialogue history information showing the history of the dialogue between the occupant and the dialogue agent is acquired. The response generation unit, The image and the dialogue history information are input into a multimodal large-scale language model to identify the occupant's object of interest, generate response information indicating that a search related to the occupant's object of interest will be performed, and decide to perform a search using multiple search methods based on the dialogue history information including the response information. The information processing system according to claim 1.

19. A method of information processing performed by a computer, To acquire environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object, Based on the aforementioned environmental information and the aforementioned state information, response control information is generated, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupants of the mobile body. Based on the response control information, the response information is generated, Information processing methods including

20. To acquire environmental information, which is information about the environment surrounding the moving object, and state information, which is information about the state of the moving object, Based on the aforementioned environmental information and the aforementioned state information, response control information is generated, which is information for generating response information relating to the response of a dialogue agent that engages in natural language dialogue with the occupants of the mobile body. Based on the response control information, the response information is generated, An information processing program that causes a computer to execute something.

Citation Information

Patent Citations

  • Agent device

    JP2022165339A