Information processing apparatus, information processing method, and program

The information processing device addresses the limitations of existing vehicle information systems by using multiple agents to provide personalized and real-time information and continuous conversation based on the vehicle's situation, improving user interaction.

WO2026004565A1PCT designated stage Publication Date: 2026-01-02SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/020793
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-06-09
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing vehicle information systems, such as radio broadcasts and conventional AI technologies, fail to provide personalized and real-time information tailored to the specific situation of each vehicle or user, and require user-initiated questions for conversations.

Method used

An information processing device and method that utilizes multiple agents with different personalities to generate and control output information based on the external and internal situations of a moving object, enabling continuous conversation without user input.

Benefits of technology

Enables personalized, real-time information delivery and continuous conversation tailored to the vehicle's situation, enhancing user experience and convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025020793_02012026_PF_FP_ABST
    Figure JP2025020793_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing apparatus, an information processing method, and a program that make it possible to conduct a conversation which is appropriate for a situation. The information processing apparatus comprises: an information processing unit for controlling execution of a plurality of agents which are provided with mutually different individualities and which each generate output information that contains speech content related to a topic in accordance with at least one from among a situation outside a moving body and a situation inside the moving body; and an output control unit for controlling output of the output information generated by each of the agents. The present technology can be applied to, for example, a moving body which a user boards, such as a vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present technology relates to an information processing device, an information processing method, and a program, and in particular to an information processing device, an information processing method, and a program that are suitable for use in providing an agent that can have a conversation.

[0002] In recent years, a technology has been proposed for realizing a voice agent in a vehicle that recognizes the attributes of a speaker and the nature of an utterance and shows a reaction to the recognized utterance (see, for example, Patent Document 1).

[0003] International Publication No. 2023 / 090057

[0004] In response to this, there is a demand for an agent that can carry out appropriate conversations according to the situation.

[0005] The present technology has been developed in light of such circumstances, and makes it possible to carry out appropriate conversations according to the situation.

[0006] An information processing device according to one aspect of the present technology includes an information processing unit that controls the execution of a plurality of agents each having different personalities and each generating output information including speech content related to a topic corresponding to at least one of an external situation and an internal situation of a moving body, and an output control unit that controls the output of the output information generated by each of the agents.

[0007] An information processing method according to one aspect of the present technology includes a plurality of agents each having different personalities generating output information including speech content relating to a topic corresponding to at least one of an external situation and an internal situation of a moving object, and controlling output of the output information generated by each of the agents.

[0008] A program according to one aspect of the present technology causes a computer to execute a process including controlling the execution of a plurality of agents each having different personalities and each generating output information including speech content relating to a topic corresponding to at least one of an external situation and an internal situation of a moving object, and controlling the output of the output information generated by each of the agents.

[0009] In one aspect of the present technology, a plurality of agents each having different personalities generate output information including speech content related to a topic corresponding to at least one of an external situation and an internal situation of a moving object, and the output of the output information generated by each of the agents is controlled.

[0010] 1 is a block diagram showing an example of the configuration of a vehicle control system; FIG. 2 is a diagram showing an example of a sensing area of ​​an external recognition sensor of the vehicle control system of FIG. 1; FIG. 3 is a block diagram showing an example of the configuration of an information processing system to which the present technology is applied; FIG. 4 is a diagram showing an example of a conversation history recorded in a conversation history DB; FIG. 5 is a diagram showing an example of user preference information; FIG. 6 is a flowchart for explaining AI guide control processing; FIG. 7 is a schematic diagram of a vehicle windshield and display; FIG. 8 is a diagram showing an example of the display of an AI guide; FIG. 9 is a diagram showing an example of the display of an AI guide; FIG. 10 is a diagram showing an example of the configuration of a computer;

[0011] Hereinafter, embodiments of the present technology will be described. The description will be made in the following order: 0. Background of the present technology 1. Configuration example of a vehicle control system 2. Embodiment 3. Modification 4. Other

[0012] <<0. Background of the Present Technology>> First, the background of the present technology will be described.

[0013] Conventionally, radio has been used as a system for presenting audio information in a vehicle. However, radio broadcasts to an unspecified number of users and cannot present information tailored to the situation of each vehicle.

[0014] In contrast to this, the present technology makes it possible to individually present information according to the status of each moving body such as a vehicle.

[0015] Conventionally, applications that provide tour guides based on user location information have been provided, but such applications present information prepared in advance according to the user's location information, and are unable to present information in real time.

[0016] In contrast to this, the present technology makes it possible to present real-time information according to the location of a mobile object such as a user.

[0017] Conventionally, there are technologies that generate conversations in real time, such as artificial intelligence (AI). However, these conventional technologies are based on a question-and-answer format, where the user cannot receive an answer unless he or she asks a question. Therefore, the user cannot simply listen to the conversation and must constantly ask questions.

[0018] In contrast, the present technology allows one or more agents to continue the conversation without the user having to ask a question.

[0019] <<1. Configuration Example of Vehicle Control System>> FIG. 1 is a block diagram showing a configuration example of a vehicle control system 11, which is a non-limiting example of a mobility device control system to which the present technology is applied.

[0020] The vehicle control system 11 is provided in the vehicle 1 and performs processing related to automated driving of the vehicle 1. This automated driving includes levels 1 to 5 of automated driving, as well as remote driving and / or remote assistance of the vehicle 1 by a remote driver. The level of automated driving may refer to the Society of Automotive Engineers (SAE) J3016™ APL2021 Levels of Driving Automation, where SAE Level 0 denotes the lowest level of automated driving and SAE Level 5 denotes the highest level of automated driving. For example, SAE Level 1 automated driving may be composed of driver assistance functions that provide steering or braking / acceleration support to the driver, and SAE Level 5 automated driving may be composed of automated driving functions that can drive the vehicle under all conditions.

[0021] The vehicle control system 11 includes a vehicle control ECU (Electronic Control Unit) 21, a communication unit 22, a map information storage unit 23, a location information acquisition unit 24, an external recognition sensor 25, an in-vehicle sensor 26, a vehicle sensor 27, a memory unit 28, a driving automation control unit 29, a DMS (Driver Monitoring System) 30, an HMI (Human Machine Interface) 31, and a vehicle control unit 32.

[0022] Two or more (or in some cases, all) of the vehicle control ECU 21, communication unit 22, map information storage unit 23, position information acquisition unit 24, external recognition sensor 25, in-vehicle sensor 26, vehicle sensor 27, memory unit 28, driving automation control unit 29, DMS 30, HMI 31, and vehicle control unit 32 are communicatively connected to each other via a communication network 41. The communication network 41 is configured, for example, by an in-vehicle communication network or bus conforming to a digital bidirectional communication standard such as a Controller Area Network (CAN), a Local Interconnect Network (LIN), a Local Area Network (LAN), FlexRay (registered trademark), or Ethernet (registered trademark). In some embodiments, the communication network 41 may include two or more types of communication networks, and different types of communication networks may be used depending on the type of data being transmitted. For example, a CAN may be used for data related to vehicle control, and an Ethernet may be used for large-volume data. In some embodiments, two or more (or in some cases, all) units of the vehicle control system 11 may be directly connected using wireless communication (e.g., communication at a relatively short distance) without using the communication network 41. In some embodiments, the wireless communication may use a short-range wireless communication technology. Non-limiting examples of short-range wireless communication technologies include near field communication (NFC) and Bluetooth (registered trademark). In some embodiments, two or more (or in some cases, all) units of the vehicle control system 11 may be connected using the communication network 41 and a wireless communication technology (e.g., a short-range wireless communication technology).

[0023] Hereinafter, in an embodiment in which two or more units of the vehicle control system 11 communicate with each other via the communication network 41, the description of the communication network 41 will be omitted. For example, in an embodiment in which the vehicle control ECU 21 and the communication unit 22 communicate with each other via the communication network 41, it will simply be described that the vehicle control ECU 21 and the communication unit 22 communicate with each other.

[0024] The vehicle control ECU 21 is configured by various processors such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), etc. The vehicle control ECU 21 controls the entire or part of the functions of the vehicle control system 11.

[0025] The communication unit 22 communicates with various devices inside the vehicle 1 (hereinafter referred to as in-vehicle devices), various devices outside the vehicle 1 (hereinafter referred to as out-vehicle devices), other vehicles, base stations, etc., and transmits and receives various types of data. In some embodiments, the communication unit 22 may communicate using multiple communication technologies.

[0026] A non-limiting example of communication between the communication unit 22 and an external device will now be briefly described. In some embodiments, the communication unit 22 may communicate with a server (hereinafter referred to as an external server) or the like on an external network via a base station or an access point using wireless communication technology. Non-limiting examples of wireless communication technology include 5G (5th Generation Mobile Communication System), LTE (Long Term Evolution), DSRC (Dedicated Short Range Communications), etc. The external network with which the communication unit 22 can communicate is, for example, the Internet, a cloud network, or a network specific to an operator. The communication technology used by the communication unit 22 to communicate with the external network is not particularly limited as long as it is a wireless communication technology that enables digital two-way communication at a communication speed equal to or higher than a predetermined distance.

[0027] In some embodiments, the communication unit 22 may use P2P (Peer to Peer) technology to communicate with a terminal located near the vehicle. The terminal located near the vehicle may be, for example, a terminal attached to a mobile object moving at a relatively slow speed, such as a pedestrian or a bicycle, a terminal installed at a fixed location in a store, and / or an MTC (Machine Type Communication) terminal. In some embodiments, the communication unit 22 may perform V2X (Vehicle to Everything) communication. V2X communication generally refers to communication between the vehicle and another entity. Non-limiting examples of V2X communication include vehicle-to-vehicle communication with another vehicle, vehicle-to-infrastructure communication with a roadside unit, vehicle-to-home communication, and vehicle-to-pedestrian communication with a terminal carried or worn by a pedestrian.

[0028] In some embodiments, the communication unit 22 may receive a program for updating software that controls the operation of the vehicle control system 11 from outside the vehicle 1 (e.g., over the air). In some embodiments, the communication unit 22 may receive map information, traffic information, information about the surroundings of the vehicle 1, etc. from outside the vehicle 1. In some embodiments, the communication unit 22 may transmit information about the vehicle 1 or information about the surroundings of the vehicle 1, etc. to an external device or an external network. Non-limiting examples of information about the vehicle 1 that the communication unit 22 transmits to an external device or an external network include data indicating the status of the vehicle 1, recognition results by the recognition unit 73, etc. In some embodiments, the communication unit 22 may communicate with a vehicle emergency notification system. Non-limiting examples of a vehicle emergency notification system include eCall, etc.

[0029] In some embodiments, the communication unit 22 may receive electromagnetic waves transmitted by a road traffic information communication system. In some embodiments, the electromagnetic waves may be transmitted using a radio beacon, an optical beacon, FM multiplex broadcasting, or the like.

[0030] Non-limiting examples of communication with in-vehicle devices that can be performed by the communication unit 22 will now be briefly described. In some embodiments, the communication unit 22 may communicate with the in-vehicle devices using wireless communication. For example, in some embodiments, the communication unit 22 may communicate with the in-vehicle devices using wireless communication technology that enables bidirectional digital communication at a predetermined communication speed or higher. Non-limiting examples of wireless communication technology include wireless LAN, Bluetooth, NFC, and WUSB (Wireless USB). Alternatively, the communication unit 22 may communicate with the in-vehicle devices using wired communication (in addition to or as an alternative to wireless communication). For example, in some embodiments, the communication unit 22 may communicate with the in-vehicle devices using wired communication via a cable connected to a connection terminal (not shown). In some embodiments, the communication unit 22 may communicate with the in-vehicle devices using wired communication technology that enables bidirectional digital communication at a predetermined communication speed or higher. Non-limiting examples of wired communication technologies include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI)®, and Mobile High-definition Link (MHL).

[0031] Here, the in-vehicle devices refer to, for example, devices inside the vehicle 1 that are not connected to the communication network 41. The in-vehicle devices are divided into devices that constitute the vehicle control system 11 and devices that do not constitute the vehicle control system 11. Non-limiting examples of in-vehicle devices that do not constitute the vehicle control system 11 include mobile devices and wearable devices carried by users of the vehicle 1 (for example, the driver or passengers), and information devices that are temporarily installed inside the vehicle 1. These devices can, for example, be moved outside the vehicle 1 and become external devices.

[0032] The map information storage unit 23 stores maps acquired from an external device or an external network and / or maps created by the vehicle 1. For example, the map information storage unit 23 may store a three-dimensional high-precision map, a global map that has lower precision than a high-precision map and covers a wide area, or the like.

[0033] The high-precision map may be, for example, a dynamic map, a point cloud map, a vector map, etc. The dynamic map may be, for example, a map consisting of four layers of dynamic information, quasi-dynamic information, quasi-static information, and static information, and may be provided to the vehicle 1 from an external server or the like. The point cloud map may be a map composed of a point cloud (point cloud data). The vector map may be, for example, a map adapted for automated driving by associating traffic information such as the positions of lanes and traffic lights with the point cloud map.

[0034] The point cloud map and the vector map may be provided, for example, from an external server or the like, or may be created in the vehicle 1 based on sensing results from the camera 51, radar 52, LiDAR 53, etc. as a map for matching with a local map described later, and stored in the map information storage unit 23. Furthermore, when a high-precision map is provided from an external server or the like, map data of, for example, an area of ​​several hundred square meters regarding the planned route along which the vehicle 1 will travel may be acquired from the external server or the like in order to reduce communication capacity.

[0035] The position information acquisition unit 24 acquires position information of the vehicle 1. The acquired position information may be supplied to the driving automation control unit 29. In some embodiments, the position information acquisition unit 24 may receive GNSS (Global Navigation Satellite System) signals from GNSS satellites. In some embodiments, the position information acquisition unit 24 may receive signals from beacons or the like.

[0036] The external recognition sensor 25 includes various sensors used to recognize the situation outside the vehicle 1, and supplies sensor data from one or more (or in some cases, all) sensors to one or more (or in some cases, all) units of the vehicle control system 11. The type and number of sensors included in the external recognition sensor 25 are arbitrary.

[0037] In some embodiments, the external recognition sensor 25 may include a camera 51, a radar 52, a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) 53, and an ultrasonic sensor 54. Without being limited to this, the external recognition sensor 25 may be configured to include one or more types of sensors selected from the camera 51, the radar 52, the LiDAR 53, and the ultrasonic sensor 54. The number of cameras 51, radars 52, LiDARs 53, and ultrasonic sensors 54 is not particularly limited as long as they are numbers that can be realistically installed on the vehicle 1. Furthermore, the types of sensors included in the external recognition sensor 25 are not limited to this example, and the external recognition sensor 25 may include other types of sensors. Examples of sensing areas of the sensors included in the external recognition sensor 25 will be described later.

[0038] The camera 51 may use any suitable imaging method. In some embodiments, the camera 51 may use an imaging method capable of distance measurement. Non-limiting examples of cameras using imaging methods capable of distance measurement include a time-of-flight (ToF) camera, a stereo camera, a monocular camera, and an infrared camera. However, the camera 51 may simply acquire an image without distance measurement.

[0039] In some embodiments, the external recognition sensor 25 may include an environmental sensor for detecting characteristics of the environment around the vehicle 1. Non-limiting examples of environmental characteristics that may be detected include weather, climate, brightness, etc. In some embodiments, the environmental sensor may include various sensors such as a rain sensor, a fog sensor, a sunlight sensor, a snow sensor, and an illuminance sensor.

[0040] In some embodiments, the external recognition sensor 25 may include a microphone used to detect sounds around the vehicle 1 and the location of sound sources.

[0041] The interior sensor 26 includes various sensors for detecting information about the interior of the vehicle 1, and supplies sensor data from one or more (or in some cases, all) sensors to one or more (or in some cases, all) units of the vehicle control system 11. The types and number of the various sensors included in the interior sensor 26 are not particularly limited as long as they are of the types and number that can be realistically installed in the vehicle 1.

[0042] In some embodiments, the interior sensor 26 may include one or more sensors selected from the group consisting of a camera, radar, a seating sensor, a microphone, and a biometric sensor. In some embodiments, the camera included in the interior sensor 26 may use an imaging method capable of measuring distances. Non-limiting examples of cameras using imaging methods capable of measuring distances include a Time of Flight (ToF) camera, a stereo camera, a monocular camera, and an infrared camera. The camera included in the interior sensor 26 may also be a camera simply used to acquire captured images, regardless of distance measurement. The biometric sensor included in the interior sensor 26 may be provided, for example, on a seat or a steering wheel, and may detect various types of biometric information of the user.

[0043] The vehicle sensor 27 includes various sensors for detecting the state of the vehicle 1, and supplies sensor data from one or more (or in some cases, all) sensors to one or more (or in some cases, all) units of the vehicle control system 11. The types and number of the various sensors included in the vehicle sensor 27 are not particularly limited as long as they are of the types and number that can be realistically installed on the vehicle 1.

[0044] In some embodiments, the vehicle sensor 27 may include a speed sensor, an acceleration sensor, an angular velocity sensor (gyro sensor), and / or an inertial measurement unit (IMU) that integrates these. In some embodiments, the vehicle sensor 27 may include a steering angle sensor that detects the steering angle of the steering wheel, a yaw rate sensor, an accelerator sensor that detects the amount of accelerator pedal operation (e.g., pedal force, pedal stroke), and / or a brake sensor that detects the amount of brake pedal operation (e.g., pedal force, pedal stroke). In some embodiments, the vehicle sensor 27 may include a rotation sensor that detects the number of rotations of the engine or motor, an air pressure sensor that detects tire air pressure, a slip ratio sensor that detects tire slip ratio, and / or a wheel speed sensor that detects the rotation speed of the wheels. In some embodiments, the vehicle sensor 27 may include a battery sensor that detects the remaining battery level and temperature, and / or an impact sensor that can detect external impacts.

[0045] The storage unit 28 includes at least one of a non-volatile storage medium and a volatile storage medium, and stores data and programs. Non-limiting examples of storage media include magnetic storage devices such as electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), and / or hard disk drives (HDDs), semiconductor storage devices, optical storage devices, and magneto-optical storage devices. The storage unit 28 stores various programs and data used by one or more (or in some cases, all) units of the vehicle control system 11. In some embodiments, the storage unit 28 may include an event data recorder (EDR) or a data storage system for automated driving (DSSAD), and may store information about the vehicle 1 before and after an event such as an accident, as well as information acquired by the in-vehicle sensors 26.

[0046] The driving automation control unit 29 controls the driving automation function of the vehicle 1. In some embodiments, the driving automation control unit 29 may include an analysis unit 61, an action planning unit 62, and an operation control unit 63.

[0047] The analysis unit 61 performs an analysis process of the vehicle 1 and / or the surrounding situation. The analysis unit 61 includes a self-position estimation unit 71, a sensor fusion unit 72, and a recognition unit 73.

[0048] In some embodiments, the self-position estimation unit 71 may estimate the self-position of the vehicle 1 based on sensor data from the external recognition sensor 25 and a high-precision map stored in the map information storage unit 23. For example, the self-position estimation unit 71 may generate a local map based on the sensor data from the external recognition sensor 25 and estimate the self-position of the vehicle 1 by matching the local map with the high-precision map. The position of the vehicle 1 may be based on, for example, the center of the rear wheel pair axle.

[0049] In some embodiments, the local map may be a three-dimensional high-precision map, an occupancy grid map, or the like created using a technique such as SLAM (Simultaneous Localization and Mapping). The three-dimensional high-precision map may be, for example, the point cloud map described above. The occupancy grid map may be a map obtained by dividing a three-dimensional or two-dimensional space around the vehicle 1 into grids of a predetermined size and indicating the occupancy status of objects on a grid-by-grid basis. The occupancy status of an object may be indicated, for example, by the presence or absence of an object or a probability of its presence. In some embodiments, the local map may also be used, for example, in the detection process and / or recognition process of the situation outside the vehicle 1 by the recognition unit 73.

[0050] In some embodiments, the self-position estimation unit 71 may estimate the self-position of the vehicle 1 based on the position information acquired by the position information acquisition unit 24 and / or sensor data from the vehicle sensor 27 .

[0051] The sensor fusion unit 72 performs sensor fusion processing to obtain information by combining multiple different types of sensor data (for example, image data supplied from the camera 51 and sensor data supplied from the radar 52). Methods for combining different types of sensor data include, but are not limited to, compounding, integration, fusion, and association.

[0052] The recognition unit 73 executes a detection process for detecting the situation outside the vehicle 1 and / or a recognition process for recognizing the situation outside the vehicle 1 .

[0053] For example, the recognition unit 73 may perform detection processing and / or recognition processing of the situation outside the vehicle 1 based on information from the external recognition sensor 25, information from the self-position estimation unit 71, information from the sensor fusion unit 72, etc.

[0054] Specifically, for example, the recognition unit 73 may perform a detection process and / or a recognition process of objects around the vehicle 1. The object detection process may be, for example, a process of detecting the presence or absence, size, shape, position, movement, etc. of an object. The object recognition process may be, for example, a process of recognizing attributes such as the type of object, or a process of identifying a specific object. The detection process and the recognition process are not necessarily clearly separated, and there may be at least a partial overlap.

[0055] In some embodiments, the recognition unit 73 may detect objects around the vehicle 1 by performing clustering to classify a point cloud based on sensor data from the radar 52 and / or the LiDAR 53, etc. into clusters of points. This makes it possible to detect the presence, size, shape, and position of objects around the vehicle 1.

[0056] In some embodiments, the recognition unit 73 may detect the movement of objects around the vehicle 1 by tracking the movement of clusters of point clouds classified by clustering. This makes it possible to detect the speed and / or traveling direction (movement vector) of objects around the vehicle 1.

[0057] In some embodiments, the recognition unit 73 may detect and / or recognize vehicles (including bicycles), people, obstacles, structures, roads, traffic lights, traffic signs, road markings, etc. based on image data supplied from the camera 51. In some embodiments, the recognition unit 73 may recognize the type of object around the vehicle 1 by performing recognition processing such as semantic segmentation.

[0058] In some embodiments, the recognition unit 73 may perform a recognition process of traffic rules around the vehicle 1 based on the map stored in the map information storage unit 23, the result of estimation of the self-position by the self-position estimation unit 71, and / or the result of recognition of objects around the vehicle 1 by the recognition unit 73. Through this process, the recognition unit 73 may recognize the position and / or state of traffic lights, the contents of traffic signs and / or road markings, the contents of traffic regulations, and / or lanes that can be traveled, etc.

[0059] In some embodiments, the recognition unit 73 may perform recognition processing of the environment around the vehicle 1. In some embodiments, the recognition unit 73 may recognize weather characteristics (temperature, humidity, brightness) and / or road surface conditions, etc.

[0060] The behavior planning unit 62 creates a behavior plan for the vehicle 1. For example, the behavior planning unit 62 may create a behavior plan by performing route planning and route tracking.

[0061] In some embodiments, path planning may include global path planning and local path planning. Global path planning may include a process of planning a rough route from a start to a goal. Local path planning, also referred to as trajectory planning, may include generating a trajectory that allows the vehicle 1 to proceed safely and smoothly along a planned route in the vicinity of the vehicle 1, taking into account the motion characteristics of the vehicle 1, the presence of any obstacles, and the like.

[0062] In some embodiments, the path following may be a planning of an operation for safely and accurately traveling along a route planned by the route planner within a planned time. The behavior planning unit 62 may, for example, calculate a target speed and / or a target angular velocity of the vehicle 1 based on the result of the path following process.

[0063] The operation control unit 63 controls the operation of the vehicle 1 in order to realize the action plan created by the action planning unit 62 .

[0064] For example, in some embodiments, the operation control unit 63 may control the steering control unit 81, the brake control unit 82, and / or the drive control unit 83 included in the vehicle control unit 32 (described later) to perform lateral vehicle motion control and / or longitudinal vehicle motion control so that the vehicle 1 travels along the trajectory calculated by the trajectory plan. For example, the operation control unit 63 may perform control (e.g., lateral vehicle motion control, longitudinal vehicle motion control) for one or more driver assistance functions and / or driving automation. Non-limiting examples of driver assistance functions include collision avoidance or impact mitigation, following distance control (e.g., control to maintain a specific distance from a vehicle traveling in front of the vehicle 1), vehicle speed control (e.g., control to maintain a specific speed), vehicle collision warning, and lane departure warning. Non-limiting examples of driving automation include driving without operation by a driver or a remote driver.

[0065] In some embodiments, the DMS 30 may perform a driver authentication process and / or a driver state recognition process based on sensor data from the in-vehicle sensors 26 and / or input data input to the HMI 31 (described later), etc. Non-limiting examples of the driver state that may be recognized include physical condition, alertness, concentration, fatigue, gaze direction, level of intoxication, driving operation, posture, etc.

[0066] In some embodiments, the DMS 30 may perform authentication processing of a user other than the driver (e.g., a passenger) and / or recognition processing of the state of the user. In some embodiments, the DMS 30 may perform recognition processing of the interior situation of the vehicle 1 based on sensor data from the interior sensors 26. Non-limiting examples of characteristics of the interior situation of the vehicle 1 that can be recognized include temperature, humidity, brightness, odor, etc.

[0067] The HMI 31 receives various data, instructions, etc. as input, and presents the various data to the user.

[0068] The input of data to the HMI 31 will be briefly described. The HMI 31 includes an input device through which a person inputs data, instructions, etc. The HMI 31 generates an input signal based on the data, instructions, etc. input via the input device and supplies the signal to one or more (or in some cases, all) units of the vehicle control system 11. In some embodiments, the HMI 31 may include a touch panel, buttons, switches, and / or levers as input devices. Without being limited thereto, the HMI 31 may also include an input device that allows information to be input by a method other than manual operation, such as voice or gestures. In some embodiments, the HMI 31 may include an input device such as a remote control device using infrared and / or radio waves, or an externally connected device that can operate the vehicle control system 11. Non-limiting examples of externally connected devices include a mobile device (e.g., a smartphone) and a wearable device (e.g., a smart watch).

[0069] The presentation of data by the HMI 31 will be briefly described. The HMI 31 generates visual information, auditory information, and / or tactile information for the user and / or a person outside the vehicle 1. The HMI 31 may also perform output control, controlling the output, output content, output timing, and / or output method of each piece of generated information. Non-limiting examples of visual information that can be generated and output by the HMI 31 include information displayed by images or lights, such as an operation screen, a status display of the vehicle 1, a warning display, and a monitor image showing the situation around the vehicle 1. Non-limiting examples of auditory information that can be generated and output by the HMI 31 include voice guidance, warning sounds, warning messages, etc. Non-limiting examples of tactile information that can be generated and output by the HMI 31 include information imparted to the user's sense of touch by force, vibration, movement, etc.

[0070] In some embodiments, the HMI 31 may include, as an output device capable of outputting visual information, a display device that presents visual information by displaying an image itself or a projector device that presents visual information by projecting an image. In some embodiments, the display device may be, in addition to or instead of a typical display device, a device that displays visual information within the user's field of view, such as a head-up display, a see-through display, or a wearable device with an augmented reality (AR) function. In some embodiments, the HMI 31 may include, as an output device capable of outputting visual information, a display device included in a navigation device, an instrument panel, a camera monitoring system (CMS), an electronic mirror, a lamp, or the like provided in the vehicle 1.

[0071] In some embodiments, the HMI 31 may include an audio speaker, headphones, or earphones as output devices capable of outputting auditory information.

[0072] In some embodiments, the HMI 31 may include a haptic element using haptic technology as an output device capable of outputting tactile information. The haptic element may be provided on a part of the vehicle 1 that the user comes into contact with, such as the steering wheel or the seat.

[0073] The vehicle control unit 32 controls one or more (or in some cases, all) units of the vehicle 1. The vehicle control unit 32 includes a steering control unit 81, a brake control unit 82, a drive control unit 83, a body system control unit 84, a light control unit 85, and a horn control unit 86.

[0074] The steering control unit 81 detects and / or controls the state of the steering system of the vehicle 1. The steering system includes, for example, a steering mechanism including a steering wheel, an electric power steering, etc. The steering control unit 81 includes, for example, a steering ECU that controls the steering system, an actuator that drives the steering system, etc.

[0075] The brake control unit 82 detects and / or controls the state of the brake system of the vehicle 1. The brake system includes, for example, a brake mechanism including a brake pedal, an antilock brake system (ABS), a regenerative brake mechanism, etc. The brake control unit 82 includes, for example, a brake ECU that controls the brake system, an actuator that drives the brake system, etc.

[0076] The drive control unit 83 detects and / or controls the state of the drive system of the vehicle 1. The drive system includes, for example, an accelerator pedal, a drive force generating device for generating drive force such as an internal combustion engine or a drive motor, and a drive force transmission mechanism for transmitting the drive force to the wheels. The drive control unit 83 includes, for example, a drive ECU for controlling the drive system, and an actuator for driving the drive system.

[0077] The body system control unit 84 detects and / or controls the states of the body system systems of the vehicle 1. The body system systems include, for example, a keyless entry system, a smart key system, a power window device, a power seat, an air conditioning system, an airbag, a seat belt, a shift lever, etc. The body system control unit 84 includes, for example, a body system ECU that controls the body system systems, an actuator that drives the body system systems, etc.

[0078] The light control unit 85 detects and / or controls the states of various lights of the vehicle 1. Non-limiting examples of lights that can be controlled by the light control unit 85 include headlights, backlights, fog lights, turn signals, brake lights, projector lights, and bumper indicators. The light control unit 85 includes a light ECU that controls the lights, an actuator that drives the lights, and the like.

[0079] The horn control unit 86 detects and / or controls the state of the car horn of the vehicle 1. The horn control unit 86 includes, for example, a horn ECU that controls the car horn, an actuator that drives the car horn, and the like.

[0080] Fig. 2 is a diagram showing an example of a sensing area of ​​the camera 51, the radar 52, the LiDAR 53, the ultrasonic sensor 54, etc. of the external recognition sensor 25 in Fig. 1. Fig. 2 schematically shows the vehicle 1 as viewed from above.

[0081] Sensing area 101F and sensing area 101B show examples of sensing areas of the ultrasonic sensors 54. Sensing area 101F (e.g., sensing area of ​​the multiple ultrasonic sensors 54) covers the periphery of the front end of the vehicle 1. Sensing area 101B (e.g., sensing area of ​​the multiple ultrasonic sensors 54) covers the periphery of the rear end of the vehicle 1.

[0082] The sensing results in sensing area 101F and / or sensing area 101B may be used, for example, for parking assistance for vehicle 1.

[0083] Sensing area 102F, sensing area 102B, sensing area 102L, and sensing area 102R show examples of sensing areas of a short-range or medium-range radar 52. Sensing area 102F covers a position farther in front of the vehicle 1 than sensing area 101F. Sensing area 102B covers a position farther behind the vehicle 1 than sensing area 101B. Sensing area 102L covers the surrounding area behind the left side of the vehicle 1. Sensing area 102R covers the surrounding area behind the right side of the vehicle 1.

[0084] The sensing results in the sensing area 102F may be used, for example, to detect vehicles, pedestrians, etc. present in front of the vehicle 1. The sensing results in the sensing area 102B may be used, for example, for a collision prevention function behind the vehicle 1. The sensing results in the sensing area 102L and / or the sensing area 102R may be used, for example, to detect one or more objects in blind spots on the left and / or right sides of the vehicle 1.

[0085] Sensing area 103F, sensing area 103B, sensing area 103L, and sensing area 103R show examples of sensing areas sensed by camera 51. Sensing area 103F covers a position farther in front of vehicle 1 than sensing area 102F. Sensing area 103B covers a position farther behind vehicle 1 than sensing area 102B. Sensing area 103L covers the periphery on the left side of vehicle 1. Sensing area 103R covers the periphery on the right side of vehicle 1.

[0086] The sensing results in sensing area 103F may be used, for example, for recognizing traffic lights and traffic signs, a lane departure prevention assistance system, or an automatic headlight control system. The sensing results in sensing area 103B may be used, for example, for parking assistance and / or a surround view system. The sensing results in sensing area 103L and / or sensing area 103R may be used, for example, for a surround view system.

[0087] Sensing area 104 shows an example of the sensing area of ​​LiDAR 53. Sensing area 104 covers a position farther ahead of vehicle 1 than sensing area 103F. On the other hand, sensing area 104 has a narrower range in the left-right direction of vehicle 1 than sensing area 103F.

[0088] The sensing results in the sensing area 104 may be used to detect objects such as surrounding vehicles, for example.

[0089] Sensing area 105 shows an example of the sensing area of ​​the long-range radar 52. Sensing area 105 covers a position further ahead of the vehicle 1 than sensing area 104. On the other hand, sensing area 105 has a narrower range in the left-right direction of the vehicle 1 than sensing area 104.

[0090] The sensing results in the sensing area 105 may be used for, for example, adaptive cruise control (ACC), emergency braking, collision avoidance, and the like.

[0091] In some embodiments, the sensing area of ​​each of the external recognition sensors 25 (e.g., the camera 51, the radar 52, the LiDAR 53, and the ultrasonic sensor 54) may have various configurations other than the configuration shown in Fig. 2. Specifically, in some embodiments, the ultrasonic sensor 54 may also sense the sides of the vehicle 1, and the LiDAR 53 may sense the rear of the vehicle 1. Furthermore, the installation position of each sensor is not limited to the above-mentioned examples. Furthermore, the number of each sensor may be one or more.

[0092] The present technology makes it possible to realize an agent capable of carrying out appropriate conversation depending on the situation in a moving object such as a vehicle 1, for example.

[0093] <<2. Embodiment>> Next, an embodiment of the present technology will be described with reference to FIGS.

[0094] <Configuration Example of Information Processing System 201> FIG. 3 shows a configuration example of an information processing system 201 to which the present technology is applied.

[0095] The information processing system 201 is, for example, a system realized in a mobile body such as the above-described vehicle 1. The information processing system 201 is, for example, a system that executes a predetermined application program (hereinafter referred to as an AI guide APP) in the mobile body, and provides various types of information to a passenger of the mobile body (hereinafter also referred to as a user) through conversation using an AI guide, which is a character that uses AI.

[0096] The information processing system 201 includes an input unit 211, an external data acquisition unit 212, an internal data acquisition unit 213, a communication unit 214, a conversation history DB (database) 215, a user preference DB (database) 216, a control unit 217, an audio output unit 218, and a display unit 219.

[0097] The input unit 211 includes, for example, various input devices, and is used by the user to perform operations and input data to the information processing system 201. The input unit 211 supplies input data indicating the input content, etc. to an information processing unit 231 of the control unit 217.

[0098] For example, in the vehicle 1 of FIG. 1 , a part of the HMI 31 constitutes at least a part of the input unit 211 .

[0099] The external data acquisition unit 212 includes sensors that detect the external (surrounding) conditions of the mobile object, such as a camera, LiDAR, a GNSS receiver, a microphone, etc. The external data acquisition unit 212 acquires sensor data (hereinafter referred to as external data) that indicates the external conditions of the mobile object detected by each sensor, and supplies the data to the information processing unit 231.

[0100] For example, in the vehicle 1 of FIG. 1 , the position information acquisition unit 24 and at least a part of the external recognition sensor 25 constitute at least a part of the external data acquisition unit 212 .

[0101] The internal data acquisition unit 213 includes sensors that detect the internal conditions of the mobile object, such as a camera, LiDAR, a touch panel, a microphone, etc. The internal data acquisition unit 213 acquires sensor data (hereinafter referred to as internal data) that indicates the internal conditions of the mobile object detected by each sensor, and supplies the data to the information processing unit 231.

[0102] For example, in the vehicle 1 of FIG. 1 , at least a part of the interior sensor 26 , the vehicle sensor 27 , and the DMS 30 constitute at least a part of the internal data acquisition unit 213 .

[0103] The communication unit 214 communicates with, for example, an external server.

[0104] For example, in the vehicle 1 of FIG. 1 , at least a part of the communication unit 22 constitutes at least a part of the communication unit 214 .

[0105] The conversation history DB215 is a database that records the conversation history of each user and each AI guide within the vehicle.

[0106] FIG. 4 shows an example of a conversation history recorded in the conversation history DB 215. As shown in FIG.

[0107] The conversation history includes the date and time, the location, the content of the conversation, and the user's reaction.

[0108] The date and time indicates the date and time when the conversation took place.

[0109] The location indicates the location of the mobile when the conversation took place.

[0110] The conversation content includes specific conversation content between the user and the AI ​​guide. For example, the conversation content includes a user ID for identifying the user, an AI guide ID for identifying the AI ​​guide, and the speech content of each user and AI guide.

[0111] The user's reaction includes the recognition result of the user's reaction to the conversation, which is expressed by, for example, the user's facial expression, gaze, gestures, etc.

[0112] In addition, the conversation history DB215 may record the history of conversations between users, conversations between AI guides, conversations between users alone (user monologue), and conversations between AI guides alone (AI guide monologue).

[0113] The user preference DB 216 is a database that records preference information indicating the preferences of each user who uses the AI ​​guide APP in a mobile object.

[0114] FIG. 5 shows an example of preference information of a certain user stored in the user preference DB 216. As shown in FIG.

[0115] The preference information includes a category and an interest level.

[0116] The category indicates the type of conversation category.

[0117] The interest level is a score that indicates the level of interest the user has in the target category.

[0118] The control unit 217 includes a processor such as a CPU (Central Processing Unit), and controls the execution of various processes in the information processing system 201. The control unit 217 includes an information processing unit 231 and an output control unit 232.

[0119] The information processing unit 231 is realized, for example, by the control unit 217 executing the AI ​​guide APP. The information processing unit 231 includes a recognition unit 241, an information generation unit 242, an information selection unit 243, and a learning unit 244.

[0120] The recognition unit 241 recognizes the external situation of the moving object based on external data. The recognition unit 241 recognizes the internal situation of the moving object based on internal data. The recognition unit 241 supplies information indicating the recognition results of the external and internal situations of the moving object to the information generation unit 242, the information selection unit 243, and the learning unit 244.

[0121] The information generating unit 242 includes agents 251 - 1 to 251 - n that can be executed by the information processing unit 231 .

[0122] Hereinafter, when there is no need to distinguish between the agents 251-1 to 251-n, they will be simply referred to as agents 251.

[0123] Each agent 251 is a software agent that uses a pre-trained AI model (for example, a Large Language Model (LLM)). Each agent 251 operates independently and can be started or stopped at any time.

[0124] Each agent 251 realizes an AI guide, which is a character with a different personality, and presents it to the user. Each agent 251 uses the AI ​​guide to present various information to the user while conversing with the user or other AI guides as necessary.

[0125] Each agent 251 receives external information from an external server or the like via the communication unit 214. Each agent 251 generates output information based on at least one of the situation outside the mobile object, the situation inside the mobile object, the external information, and text information indicating the content of its own utterances. The output information includes, for example, text information and image information. The text information is information including, for example, text indicating the content of the utterances (lines) of the AI ​​guide. The image information is information including, for example, an image (video or still image) of the AI ​​guide. Each agent 251 supplies the output information to the information selection unit 243.

[0126] Each agent 251 can activate other agents 251. For example, the agent 251 activates other agents 251 as needed based on at least one of a situation outside the mobile body and a situation inside the mobile body.

[0127] In the following, the agent 251 corresponding to each AI guide, i.e., the agent 251 for realizing each AI guide, will be simply referred to as the agent 251 of each AI guide.

[0128] The information selection unit 243 selects output information to be presented to the user from the output information generated by each agent 251. For example, the information selection unit 243 selects output information to be presented to the user based on at least one of the external situation of the mobile object, the internal situation of the mobile object, the conversation history of each user recorded in the conversation history DB 215, and the preferences of each user recorded in the user preference DB 216. The information selection unit 243 supplies text information included in the selected output information to the learning unit 244 and the speech synthesis unit 261 of the output control unit 232. The information selection unit 243 supplies image information included in the selected output information to the display control unit 263 of the output control unit 232.

[0129] The learning unit 244 updates the conversation history DB 215 based on the external situation of the mobile object, the internal situation of the mobile object, and the text information of the output information selected by the information selection unit 243. The learning unit 244 learns the preferences of each user based on the conversation history of each user stored in the conversation history DB 215, and updates the user preference DB 216.

[0130] The output control unit 232 controls the output of the output information and includes a voice synthesis unit 261, a voice output control unit 262, and a display control unit 263.

[0131] The voice synthesis unit 261 converts text information into voice data using a voice synthesis technique (TTS (Text-to-Speech)), and supplies the voice data to the voice output control unit 262 .

[0132] Furthermore, for example, the voice synthesis unit 261 includes a plurality of TTSs that respectively realize the tone of voice, speaking style, etc. of each AI guide. For example, the voice synthesis unit 261 switches the TTS to be used depending on the AI ​​guide that outputs (speaks) the voice.

[0133] The audio output control unit 262 controls the output of audio based on audio data by the audio output unit 218. For example, the audio output control unit 262 controls the output timing, output position, volume, etc. of the audio.

[0134] The display control unit 263 converts the image information into image data that is laid out in a form that is easy for the user to understand using, for example, a GUI (Graphic User Interface) or the like. The display control unit 263 controls the display of an image based on the image data by the display unit 219. For example, the display control unit 263 controls the display timing, display position, display size, etc. of the image.

[0135] The audio output unit 218 includes an audio output device such as a speaker, and outputs audio based on the audio data. For example, the audio output unit 218 may include stereo speakers or surround speakers whose audio output position can be controlled.

[0136] For example, in the vehicle 1 , at least a part of the HMI 31 constitutes at least a part of the audio output unit 218 .

[0137] The display unit 219 includes a display device such as a display or a head-up display (HUD), and displays an image based on the image data.

[0138] For example, in the vehicle 1 , at least a part of the HMI 31 constitutes at least a part of the display unit 219 .

[0139] <AI Guide Control Processing> Next, the AI ​​guide control processing executed by the information processing system 201 will be described with reference to the flowchart of FIG.

[0140] This process is started, for example, when an operation to start the AI ​​guide APP is performed via the input unit 211, and is ended when an operation to stop the AI ​​guide APP is performed.

[0141] In step S1, the information processing system 201 acquires external data and internal data. For example, the external data acquisition unit 212 acquires external data indicating the external conditions of the mobile object detected by each sensor and supplies the data to the information processing unit 231. The internal data acquisition unit 213 acquires internal data indicating the internal conditions of the mobile object detected by each sensor and supplies the data to the information processing unit 231.

[0142] In step S2, the recognition unit 241 recognizes the external and internal conditions of the moving object.

[0143] For example, the recognition unit 241 recognizes the external situation of the moving object based on external data. For example, the recognition unit 241 recognizes the external situation of the moving object that is likely to be a main topic of conversation. Specifically, for example, the recognition unit 241 recognizes the current location of the moving object. For example, the recognition unit 241 recognizes other moving objects that have characteristics different from ordinary moving objects (for example, fast vehicles, ultra-luxury cars, etc.). For example, the recognition unit 241 recognizes people that have characteristics different from ordinary people (for example, queues, performers, etc.). For example, the recognition unit 241 recognizes buildings that have characteristics different from ordinary buildings (for example, famous buildings, buildings under construction, etc.).

[0144] For example, the recognition unit 241 recognizes the internal situation of the mobile object based on internal data. For example, the recognition unit 241 recognizes the user's situation inside the mobile object as the internal situation of the mobile object. Specifically, for example, the recognition unit 241 performs personal identification of the user. For example, the recognition unit 241 recognizes the content of the user's utterance through voice recognition. For example, the recognition unit 241 recognizes the user's facial expression, posture, gestures, gaze, etc. in order to recognize the user's emotions, reactions, interests, physical condition, etc.

[0145] For example, the recognition unit 241 recognizes the object of interest of the user by recognizing the situation in the direction in which the passenger is pointing based on the recognition result of the user's gesture and external data.

[0146] For example, the recognition unit 241 recognizes the user's subject of interest by recognizing the state of the user's line of sight direction based on the recognition result of the user's line of sight and external data.

[0147] For example, the recognition unit 241 recognizes the content of each AI guide's speech by voice recognition as the internal situation of the moving body.

[0148] The recognition unit 241 supplies information indicating the recognition results of the external and internal situations of the mobile object to the active agent 251 , the information selection unit 243 , and the learning unit 244 .

[0149] In step S3, the active agent 251 acquires external information from an external server or the like via the communication unit 214.

[0150] The content and source of the external information are not particularly limited as long as it includes information that is likely to become a topic of conversation with the user. For example, the external information includes current events such as news and weather, information about the user's friends on social media, etc.

[0151] In step S4, the information processing unit 231 determines whether or not to change the AI ​​guide.

[0152] Changes to an AI guide include, for example, adding, deleting, and replacing an AI guide. Adding an AI guide is when a new AI guide is started. Deleting an AI guide is when an AI guide that is currently running is stopped. Replacing an AI guide is, for example, when an AI guide that is currently running is stopped and a new AI guide is started instead.

[0153] The AI ​​guide may be changed by a user operation or by the AI ​​guide in operation.

[0154] For example, the user inputs an instruction to replace, add, or delete an AI guide via the input unit 211. At this time, the user can specifically specify the AI ​​guide to be newly activated and the AI ​​guide to be stopped. The input unit 211 supplies input data indicating an instruction to change the AI ​​guide to the information processing unit 231.

[0155] Then, when the information processing unit 231 recognizes an instruction to change the AI ​​guide based on the input data, it determines that the AI ​​guide should be changed, and the processing proceeds to step S5.

[0156] For example, an active AI guide can call (summon) other AI guides as needed. For example, an active AI guide can call other AI guides who are good at a topic that comes up during a conversation with a user. In other words, the agent 251 of the active AI guide can instruct the activation of the agent 251 of another AI guide who is good at that topic.

[0157] When the information processing unit 231 receives an instruction from the active agent 251 to activate another agent 251, it determines that the AI ​​guide should be changed, and the process proceeds to step S5.

[0158] For example, an operating AI guide can stop its own operation when necessary. For example, when a topic that the operating AI guide is not good at comes up during a conversation with a user, the operating AI guide can stop its own operation. In other words, the agent 251 of the operating AI guide can instruct the AI ​​guide to stop its own operation.

[0159] When the information processing unit 231 receives an instruction from the active agent 251 to stop its own operation, it determines that the AI ​​guide should be changed, and the process proceeds to step S5.

[0160] In step S5, the information processing unit 231 changes the AI ​​guide. For example, when starting a new AI guide, the information processing unit 231 starts the agent 251 of the AI ​​guide to be started. For example, when stopping an AI guide that is currently running, the information processing unit 231 stops the agent 251 of the AI ​​guide to be stopped.

[0161] Then, by combining the activation of a new AI guide and the stopping of an AI guide that is currently running, an AI guide can be added, deleted, or replaced.

[0162] Then, the process proceeds to step S6.

[0163] On the other hand, if it is determined in step S4 that the AI ​​guide is not to be changed, the process of step S5 is skipped and the process proceeds to step S6.

[0164] In step S6, the information generating unit 242 generates output information.

[0165] For example, the information generation unit 242 inputs to the operating agent 251 at least one of the recognition results of the situation outside the moving body, the recognition results of the situation inside the moving body, external information, and text information included in output information that the agent 251 has previously output (for example, output immediately before).

[0166] As a result, for example, the agent 251 generates output information including at least one of utterance content related to a topic corresponding to the external situation of the mobile body, utterance content related to a topic corresponding to the internal situation of the mobile body, utterance content related to a topic corresponding to external information, utterance content corresponding to the user's utterance content, and utterance content corresponding to the utterance content of another agent 251.

[0167] When multiple agents 251 are active, for example, each agent 251 may generate output information individually, or only some of the agents 251 may generate output information. Also, for example, each agent 251 may generate multiple pieces of output information each containing different utterance content (for example, utterance content related to different topics).

[0168] Each agent 251 supplies the generated output information to the information selection unit 243 .

[0169] In step S7, the information selection unit 243 selects output information. For example, when a plurality of pieces of output information have been generated, the information selection unit 243 selects the output information to be presented to the user based on at least one of the external situation of the mobile object, the internal situation of the mobile object, the conversation history of each user recorded in the conversation history DB 215, and the preferences of each user recorded in the user preference DB 216.

[0170] Specifically, for example, the information selection unit 243 calculates the degree of interest and the degree of novelty for each piece of output information.

[0171] The level of interest is an index indicating the level of interest of the user in a topic included in the speech content included in the output information (hereinafter referred to as a topic corresponding to the output information).

[0172] For example, in the user preference information stored in the user preference DB 216, the higher the level of interest in the category to which the topic corresponding to the output information belongs, the higher the level of interest in the output information. Conversely, for example, in the user preference information, the lower the level of interest in the category to which the topic corresponding to the output information belongs, the lower the level of interest in the output information. This allows the user's long-term interest and level of interest in the topic corresponding to the output information to be evaluated.

[0173] For example, the recognition unit 241 recognizes an object that the user is paying attention to (hereinafter referred to as an "attention object") based on the user's gaze or gestures. For example, an object that the user is looking at or pointing at is recognized as the attention object. For example, if the topic corresponding to the output information is related to the attention object, the level of interest in the output information increases. This allows the user's short-term interest or level of interest in the topic corresponding to the output information to be evaluated.

[0174] For example, the recognition unit 241 recognizes the user's reaction to the topic spoken by the AI ​​guide based on at least one of the user's facial expression, speech content, and gestures. The topic spoken by the AI ​​guide is a topic spoken by the AI ​​guide based on output information prior to the output information to be evaluated for interest level (a topic included in the speech content output by the AI ​​guide).

[0175] For example, the more the topic corresponding to the output information is a topic that receives a positive user response, the higher the level of interest in the output information. On the other hand, for example, the more the topic corresponding to the output information is a topic that receives a negative user response, the lower the level of interest in the output information. This allows the user's short-term interest in the topic corresponding to the output information to be evaluated.

[0176] For example, if the topic corresponding to the output information is related to a topic that appeared in the user's most recent conversation history in the conversation history DB 215, the level of interest in the output information increases. This increases the level of interest in output information that includes utterance content related to the topic the user was most recently talking about. This allows the user's short-term interest in the topic corresponding to the output information to be evaluated.

[0177] The novelty is an index showing the recency of a topic corresponding to output information for a user. For example, in the conversation history DB 215, if a topic corresponding to output information does not exist in the user's past conversation history, i.e., if the topic has not appeared in a conversation with the user in the past, the novelty of the output information is high. On the other hand, if a topic corresponding to output information exists in the user's past conversation history, i.e., if the topic has appeared in a conversation with the user in the past, the novelty of the output information is low.

[0178] Furthermore, in the user's past conversation history, the more recently a topic corresponding to the output information appears, the lower the novelty of the output information, and the older the appearance, the higher the novelty of the output information. In other words, the more recently the output information related to a topic that appeared in a conversation with the user, the lower the novelty, and the more recently the output information related to a topic that appeared in a conversation with the user, the higher the novelty.

[0179] The user's most recent conversation history may be excluded from the calculation of the novelty, which prevents a decrease in the novelty of output information related to, for example, the topic the user was talking about most recently (i.e., the topic in which the user is currently interested).

[0180] The information selection unit 243 calculates the degree of interest and the degree of novelty for each piece of output information, and calculates a selection score based on the calculated degrees of interest and novelty. The selection score increases as the degree of interest increases and decreases as the degree of interest decreases. The selection score increases as the degree of novelty increases and decreases as the degree of novelty decreases. For example, the information selection unit 243 selects, from among a plurality of pieces of output information, the output information with the highest selection score.

[0181] As a result, output information including speech content relating to topics that are of greater interest to the user is preferentially selected.

[0182] In addition, output information including utterance content related to a topic that has not been discussed with the user or a topic that has not been discussed with the user for a long period of time is preferentially selected, thereby prioritizing fresh topics and preventing the same topic from appearing repeatedly in a short span of time.

[0183] It should be noted that the selection score may be calculated based on only one of the interest level and the novelty level, for example.

[0184] The information selection unit 243 supplies the text information included in the selected output information to the learning unit 244 and the voice synthesis unit 261 , and supplies the image information to the display control unit 263 .

[0185] When only one piece of output information is generated, the information selection unit 243 supplies the text information included in the output information to the learning unit 244 and the voice synthesis unit 261 , and supplies the image information to the display control unit 263 .

[0186] For example, if the user is not interested in the current topic of the AI ​​guide, the user can input an instruction via the input unit 211 to skip the conversation on the current topic or to change the topic.

[0187] In response to this, the information selection unit 243 may select output information relating to a topic different from that of the output information currently being output, in accordance with a user instruction.

[0188] In step S8, the information processing system 201 outputs the output information.

[0189] Specifically, the voice synthesis unit 261 converts the text information into voice data using TTS for the AI ​​guide of the agent 251 that generated the selected output information. The voice synthesis unit 261 supplies the voice data to the voice output control unit 262.

[0190] The audio output unit 218 outputs audio based on the audio data under the control of the audio output control unit 262. As a result, the audio (speech sound) of the speech content indicated in the text information included in the output information is output in a voice corresponding to the AI ​​guide of the agent 251 that generated the output information.

[0191] The display control unit 263 converts the image information into image data in a predetermined format.

[0192] The display unit 219 displays an image based on the image data under the control of the display control unit 263. As a result, for example, an image including a state in which the AI ​​guide of the agent 251 that generated the selected output information speaks is displayed.

[0193] In step S9, the learning unit 244 updates the conversation history.

[0194] For example, when the user's utterance is recognized in the most recent processing of step S2, the learning unit 244 adds information about the content of the recognized user's utterance to the conversation history recorded in the conversation history DB 215. For example, the learning unit 244 adds information about the content of the AI ​​guide's utterance output in the most recent processing of step S8 to the conversation history recorded in the conversation history DB 215.

[0195] For example, if the learning unit 244 recognizes the user's reaction to the AI ​​guide's speech in the most recent processing of step S2, it adds information about the recognized user's reaction to the conversation history recorded in the conversation history DB 215.

[0196] In step S10, the learning unit 244 appropriately learns the preferences of the user. Specifically, the learning unit 244 learns the preferences of the user based on the conversation history of each user recorded in the conversation history DB 215, as necessary.

[0197] For example, the learning unit 244 increases the level of interest in the category of the user's preference information to be learned in the user preference DB 216, to which the user responded well during a conversation with the AI ​​guide.

[0198] Below is an example of a good response.

[0199] - A user nods during a conversation. - A user has a positive facial expression (e.g., smiles) during a conversation. - A user asks a related question. - A user starts a conversation with another user on a related topic.

[0200] On the other hand, for example, the learning unit 244 reduces the level of interest in categories containing topics to which the user responded negatively during a conversation with the AI ​​guide in the preference information of the user being learned in the user preference DB 216.

[0201] Below is an example of a bad reaction:

[0202] - The user is unresponsive. - The user has a negative facial expression (for example, a frown) during a conversation. - The user quickly tries to change the topic. - The user starts an unrelated conversation with another user.

[0203] In addition, the learning unit 244 may further use information regarding the user's preferences obtained from an external server, etc. (e.g., internet search / browsing history, shopping purchase history, restaurant reservation history, etc.) to update the user's preference information.

[0204] Furthermore, learning of user preferences may be performed, for example, periodically, at a predetermined timing, or when the accumulated amount of conversation history reaches or exceeds a predetermined threshold.

[0205] Thereafter, the process returns to step S1, and the processes of steps S1 to S10 are repeatedly executed.

[0206] <Specific Example of Processing by AI Guide APP> Next, with reference to FIGS. 7 to 10, a specific example of processing by the AI ​​guide APP when the information processing system 201 is provided in the vehicle 1 will be described.

[0207] Hereinafter, displaying an image of the AI ​​guide may be simply referred to as "displaying the AI ​​guide." Hereinafter, the AI ​​guide's agent 251 performing various processes may be simply referred to as "the AI ​​guide performing various processes." Specifically, for example, the AI ​​guide's agent 251 generating speech content may be simply referred to as "the AI ​​guide generating speech content."

[0208] 7 is a schematic diagram of a windshield 301 and a display 302 at the front of the vehicle 1. The display 302 is, for example, a touch panel display, and is disposed in front of the driver's seat and passenger seat on the dashboard below the windshield 301, extending left and right. In front of the driver's seat, a digital instrument panel 311 is displayed on the right side of the display 302.

[0209] 7 to 10 show the same image inside the windshield 301, but in reality the scenery inside the windshield 301 changes as the vehicle 1 moves.

[0210] For example, when the AI ​​guide APP is launched, a menu 312 for selecting the AI ​​guide is displayed in the center of the display 302 .

[0211] Information about multiple selectable AI guides is displayed in the menu 312. For example, the AI ​​guide's facial image, characteristics (e.g., name, personality, occupation, specialty, etc.), etc. are displayed. For example, the displayed AI guides (i.e., selectable AI guides) can be switched by scrolling the menu 312 left or right.

[0212] For example, the user can select the AI ​​guide to be activated by touching the desired AI guide from the menu 312. When the AI ​​guide is selected, an image of the AI ​​guide is displayed within the user's field of vision. At this time, the user can select multiple AI guides and activate multiple AI guides simultaneously.

[0213] For example, when historian James, who is knowledgeable about history, is selected from the AI ​​guides in menu 321, an image of historian James is displayed within the user's field of view. For example, as shown in Fig. 8 , an icon 331 of historian James is projected onto the real space through windshield 301 using AR (Augmented Reality) by the HUD included in display unit 219.

[0214] The icon 331 is spherical and displays the facial image of historian James inside. The facial image within the icon 331 may, for example, change its expression in accordance with the conversation. In reality, the icon 331 is semi-transparent, allowing the background of the icon 331 (the scenery ahead of the vehicle 1) to be seen through.

[0215] The method for displaying the AI ​​guide is not particularly limited. For example, the AI ​​guide may be displayed as a 3D avatar. For example, the AI ​​guide, which is a 3D avatar, may add gestures in accordance with the conversation.

[0216] The AI ​​guide may be displayed at a position that matches the position of the recognized user. For example, in the example of Fig. 8, the user sitting in the driver's seat is recognized, and an icon 331 is projected in front of the driver's seat through the windshield 301.

[0217] If the vehicle 1 does not have a HUD, the icon 331 may be displayed on the display 302.

[0218] Also, if the user is recognized as someone who has ridden vehicle 1 in the past, for example, historian James may greet the user at startup with a greeting such as "Hello, Mr. A."

[0219] For example, the AI ​​guide can hold a conversation regarding the current location of the vehicle 1.

[0220] Specifically, for example, when historian James recognizes that vehicle 1 is traveling in Ginza, he can engage in a conversation about the history of Ginza. For example, historian James might say, "Welcome to Tokyo sightseeing! This vehicle is currently traveling in Ginza. Speaking of Ginza, do you know the origin of the place name? It began in the Edo period when the then shogun, Tokugawa Ieyasu, moved the silver coin mint here to Ginza. The official name of the town was "Shinryogaecho," but it became commonly known as "Ginza."

[0221] In contrast, if the user does not terminate the AI ​​guide APP, Historian James continues the conversation about the history of Ginza. For example, by inputting text information corresponding to the previous utterance of the historian agent into Historian James's agent 251, Historian James can continue the conversation about the history of Ginza. For example, Historian James might utter, "By the way, at that time there was also a place called 'Ginza' in addition to 'Ginza'. That is where the Bank of Japan headquarters is now located."

[0222] For example, historian James can conduct a conversation about the situation outside vehicle 1.

[0223] Specifically, for example, if a car is recognized as having overtaken vehicle 1 at high speed, Historian James will say, "Wow, that's a really fast car. The car that just overtook you is a Ferrari 296 GTB. Incidentally, Ferrari was built in 1929 by the Italian Enzo Ferrari..."

[0224] In this case, historian James may provide more detailed information if the user is a car enthusiast, for example, or may quickly change the topic if the user is not a car enthusiast.

[0225] For example, if a famous building is recognized around vehicle 1, historian James may engage in a conversation related to the recognized building. For example, if "Ginza Matsuya" is recognized, historian James may say, "What you can see ahead to the right is a department store called Ginza Matsuya. Its predecessor was a kimono fabric store established in Yokohama in 1869, and the department store opened here in Ginza in 1925. At the time..."

[0226] During this speech, for example, historian James's agent 251 may obtain a photo of Ginza Matsuya when it first opened in 1925 from an external server via the communication unit 214, and generate image information including the photo. Then, based on the generated image information, a photo of Ginza Matsuya when it was open may be projected in front of the windshield 301 or displayed on the display 302.

[0227] Furthermore, for example, the display control unit 263 may perform a three-dimensional reconstruction process and use technology such as AR to replace the actual Ginza Matsuya building with an image of the building as it was when the store was first established, thereby providing the user with an enjoyable experience of time travel.

[0228] For example, the AI ​​guide can hold a conversation about the subject that the user is focusing on.

[0229] For example, if the user is recognized as looking at the queue in front of a store ahead of vehicle 1, James the Historian may engage in a conversation about that store.

[0230] In this case, Historian James may develop a conversation about the history of ramen based on his own personality, for example, Historian James might say, "There's a huge line! This ramen shop is called XX and it just opened a year ago. Its rich pork bone soup is popular, but the owner of this shop actually trained at the famous YY for five years before being allowed to open a branch..."

[0231] This allows users to enjoy different conversations depending on the AI ​​guide's personality, even if the topic is the same.

[0232] For example, each AI guide can carry on a conversation by itself, but if there are other AI guides with different personalities, the AI ​​guides can converse with each other, allowing the conversation to continue more naturally. Also, a conversation about a topic that an AI guide is not good at can be left to another AI guide who is good at that topic.

[0233] For example, if a user says to Historian James, "I want to eat lunch," the conversation may unfold as follows:

[0234] For example, historian James might say, "I'm a historian, so I know a lot about the history of restaurants, but I don't know much about the food itself. Let's take this opportunity to call Chef Kate. Kate!"

[0235] In this way, an active AI guide can call another AI guide depending on the situation. This allows, for example, Chef Kate, an AI guide who is good at cooking, to appear. For example, as shown in FIG. 9 , Chef Kate's icon 332 is projected onto the real space through the windshield 301. The display mode of the icon 331 is the same as that of the historian James's icon 331.

[0236] In addition, the user can also select a new AI guide to appear from among the AI ​​guides in menu 312.

[0237] The conversation then continues through dialogue between multiple AI guides.

[0238] For example, historian James might say, "Thanks for coming, Kate! ZZ is looking for a nice restaurant, but she's not very knowledgeable about food...do you have any recommendations?"

[0239] In response, chef Kate says, "Hello, I'm Kate! From here on, I'll join James as your guide. Ginza is a treasure trove of restaurants. Where would you recommend? Hmm... For example, how about Trattoria Milano, a hidden restaurant in Ginza? Their Sicilian-style carbonara pasta is excellent, and the interior is relaxing, so it's recommended for families too."

[0240] In response, historian James says, "Wow! That looks delicious. I'd love to try it."

[0241] In this way, by having multiple AI guides converse with each other, appropriate information can be presented to the user in a natural way through conversational exchanges.

[0242] At this time, since Chef Kate has learned about cooking, she may recommend the most suitable restaurant based on a database accessible to Chef Kate's agent 251. For example, in addition to recommending restaurants based on the user's preferences, Chef Kate may also recognize children in the back seat and, if it is determined that the car is a family, recommend a restaurant that is suitable for families.

[0243] Next, for example, if Historian James recognizes the user's facial expressions and conversations with other users in the car and realizes that the user seems to want to know more about other restaurant options, he or she will say, "But it seems like you all want to know more about other options. Is there anything else?"

[0244] In response, chef Kate says, "Yes, the All Day Dining here is also absolutely delicious. It might be the best steakhouse in Ginza. They have a variety of menu items, but I recommend the T-bone steak. It's the restaurant's most popular item. It has an American feel, so if you like wagyu beef..."

[0245] In this case, for example, if the restaurant recommended by Chef Kate is visible from vehicle 1, Chef Kate's icon 332 may move to a location near the restaurant as seen from the user's perspective, allowing the conversation to continue, as shown in Fig. 10. In other words, the display position of icon 332 may be controlled based on the location of the restaurant that is the subject of conversation in the speech of Chef Kate, who is the AI ​​guide. For example, the example in Fig. 10 shows that the restaurant recommended by Chef Kate is located in a building near icon 332.

[0246] This allows conversations to be developed that include demonstratives such as "It's here..." It also allows users to intuitively recognize the location of the restaurant that is the subject of the conversation.

[0247] For example, the output position of Chef Kate's speech may be moved in the same manner as the position of icon 332. In other words, the output position of Chef Kate's speech may be controlled based on the location of the restaurant that is the subject of conversation in the speech of Chef Kate, who is the AI ​​guide.

[0248] Furthermore, information that is difficult to convey through conversation alone (such as photos of restaurants and dishes) may be displayed by the GUI. For example, in the example of Fig. 10, information about recommended restaurants is displayed in window 333.

[0249] The AI ​​guide can provide conversations that users can enjoy just by listening, but it can also interact with the user.

[0250] For example, if two users are having a conversation like, "Japanese food would be nice if possible," chef Kate can expand the conversation by saying, "For Japanese food, how about Sachi? They recommend their seafood..."

[0251] Furthermore, if the user directly tells Chef Kate, "It looks delicious, please take me there," Chef Kate's agent 251 may provide restaurant guidance by working in cooperation with the restaurant's reservation system and the vehicle 1's driving automation control unit 29, for example.

[0252] For example, chef Kate's agent 251 makes a reservation at a target restaurant using the restaurant's reservation system via the communication unit 214. Also, for example, chef Kate's agent 251 cooperates with the driving automation control unit 29 of the vehicle 1 to move the vehicle 1 to the restaurant by autonomous driving. In response to this, chef Kate can develop a conversation, for example, by saying, "This place looks delicious too! I'll contact the restaurant myself. There should be a seat available by the time we arrive. I'll take you there by autonomous driving now."

[0253] Next, a specific example of a method for selecting a conversation (output information) will be described.

[0254] For example, a case will be described in which user A and user B are on board vehicle 1 and two AI guides, historian James and chef Kate, are operating.

[0255] For example, when a historical building is recognized ahead of the vehicle 1, the historian James' agent 251 generates output information including speech content related to the building, while the chef Kate's agent 251 generates output information including speech content related to a restaurant near the building.

[0256] In response to this, the information selection unit 243 calculates, for example, the selection score of the output information of the historian James and the selection score of the output information of the chef Kate.

[0257] For example, the information selection unit 243 adds the degree of interest in history in the preference information of user A and the degree of interest in history in the preference information of user B to the selection score of the output information of historian James. The information selection unit 243 adds the degree of interest in cooking in the preference information of user A and the degree of interest in cooking in the preference information of user B to the selection score of the output information of chef Kate.

[0258] Next, for example, the information selection unit 243 refers to the conversation history DB 215, and if at least one of user A and user B has had a conversation about the building in the past, it subtracts the selection score of the output information of historian James. Also, for example, the information selection unit 243 refers to the conversation history DB 215, and if at least one of user A and user B has had a conversation about the restaurant in the past, it subtracts the selection score of the output information of chef Kate. Note that the amount of subtraction decreases as the time when the topic last came up increases.

[0259] For example, if user A and user B had just had a conversation such as "I wonder what that building is?" while pointing at a building, this conversation is also recorded in the conversation history DB 215. In this case, the selection score of the output information of historian James is added. Note that since this is the most recent conversation, the amount of addition may be larger.

[0260] The information selection unit 243 then compares the final selection scores and selects the output information with the higher selection score. As a result, conversations based on the selected output information are presented to user A and user B. That is, conversations related to topics that are of greater interest to user A and user B and that have not appeared in conversations between user A or user B in the recent past are presented preferentially.

[0261] For example, in addition to the external and internal conditions of vehicle 1, conversations may be held based on external information obtained via communication unit 214 about current events such as news and weather, or updates on friends' status via social media.

[0262] For example, if vehicle 1 is traveling in Ginza, the AI ​​guide may present conversations based on external information such as, "It seems that a robbery occurred at the jewelry store on the right yesterday," "Your friend A will be holding a presentation event tomorrow on the 7th floor of the building in front of you," and "Unfortunately, the weather is bad today, and it seems that it will be cold tomorrow as well."

[0263] For example, if there is no topic that could become a conversation topic outside or inside the vehicle 1, output information including speech content related to a topic corresponding to external information may be generated and output, regardless of the situation outside or inside the vehicle 1. This allows a conversation related to a topic based on external information to be presented. For example, a conversation related to a topic unrelated to the situation outside or inside the vehicle 1 may be presented, such as, "There's not much interesting around here, so let's talk about XX that we were talking about yesterday."

[0264] In this way, an appropriate conversation can be carried out depending on the situation.

[0265] For example, based on the results of situational recognition both inside and outside the vehicle, it becomes possible to present conversations about events occurring in front of the user in real time, making it possible to present information according to the situation of each vehicle individually and in real time.

[0266] For example, one or more AI guides can continue a conversation without the user having to ask a question. This allows the user to enjoy listening to conversations between AI guides without having to speak to them.

[0267] For example, by using AI guides with different personalities, conversations can be held according to the user's preferences. For example, if a user wants to learn about history, they can enjoy conversations about history by selecting an AI guide who is good at history.

[0268] For example, by selecting the content of the AI ​​guide's speech based on the level of interest and novelty, information that will provide greater user satisfaction can be presented.

[0269] For example, it is possible to have a conversation about a subject in which the user has shown interest through their gaze or gestures, or to present information about a topic that the users are discussing.

[0270] <<3. Modifications>> Modifications of the above-described embodiments of the present technology will now be described.

[0271] <Modifications Regarding Allocation of Processing> For example, some or all of the processing of the control unit 217 of the information processing system 201 may be executed by an external server, etc. For example, some or all of the processing of the information processing unit 231 may be executed by an external server, etc. Specifically, for example, each agent 251 may be executed by an external server, etc.

[0272] For example, at least one of the conversation history DB 215 and the user preference DB 216 may be provided in an external server or the like.

[0273] For example, each agent 251 may be configured to perform the same processing as the recognition unit 241. That is, each agent 251 may be configured to recognize the external and internal conditions of the mobile object based on external data and internal data. Each agent 251 may then generate and output output information by receiving the external data and internal data. In this case, it is possible to delete the recognition unit 241. Also, for example, each agent 251 may supply information indicating the recognition results of the external and internal conditions of the mobile object to the information selection unit 243 and the learning unit 244.

[0274] <Modifications Regarding AI Guide> The model of the AI ​​guide is not limited to a fictional character.

[0275] For example, an AI guide may be provided that is generated by learning the personality, behavior, speaking style, etc. of a celebrity. This allows users to enjoy driving while having a conversation with their favorite celebrity, for example. In this case, a business model may be established in which a license fee is paid to the celebrity.

[0276] Similarly, an AI guide may be provided that is generated by learning the personalities, actions, speech patterns, etc. of characters in animations, movies, etc. This allows, for example, a child's favorite character to guide them around town while driving.

[0277] Furthermore, for example, an AI model generated by learning the personalities, behaviors, etc. of historical figures may be provided.

[0278] For example, if the time required for generating output information by the agent 251 is long, the AI ​​guide corresponding to the agent 251 may utter a hesitation such as "Umm" to make the conversation more natural. Also, the AI ​​guide may fill in the gaps by uttering something about the internal processing, such as "I'm looking into it now..."

[0279] For example, it is possible to output only the speech of the AI ​​guide without displaying an image of the AI ​​guide.

[0280] For example, the content of the AI ​​guide's speech may be presented as text information using speech bubbles, subtitles, etc. In this case, output of speech sounds may be omitted.

[0281] The number of AI guides that can operate at one time is not particularly limited. For example, three or more AI guides may operate simultaneously and converse with each other.

[0282] In the above explanation, an example has been shown in which the AI ​​guide is displayed through the windshield 301 using a HUD, but the AI ​​guide may also be displayed through other windows, such as door windows, side windows, or rear windows.

[0283] For example, the AI ​​guide may be displayed on a display inside the vehicle 1, such as the display 302. In this case, a through image captured by a camera capturing an image of the outside of the vehicle may be displayed on the display, and the AI ​​guide may be superimposed on the through image.

[0284] For example, the display position of the AI ​​guide and the output position of the speech sound may be controlled based on the external and internal conditions of the moving body. For example, as described above, the display position of the AI ​​guide and the output position of the speech sound may be controlled based on the user's position or the position of the topic of the AI ​​guide's speech. For example, the display position of the AI ​​guide and the output position of the speech sound may be controlled based on the user's focus position (e.g., the position of the user's gaze, the position where the user is pointing, etc.).

[0285] For example, the AI ​​guide (or the corresponding agent 251) may generate output information based on the situation either outside or inside the mobile object. This may occur, for example, when the AI ​​guide is unable to recognize the situation outside the mobile object or the situation inside the mobile object.

[0286] For example, the AI ​​guide may be set to a standby state. Specifically, for example, a standby AI guide may not appear in the table (for example, an image of the AI ​​guide may not be displayed), but may listen to conversations between the user and other AI guides, understand the flow of the conversation while on standby, and immediately participate in the conversation when it returns from standby. In this case, for example, the user or other AI guide may be able to put the AI ​​guide on standby or return it from standby, or the AI ​​guide may be able to wait or return from standby on its own.

[0287] <Application Examples of the Present Technology> The present technology can be applied to any moving body, as long as it is possible for a user to board the vehicle and converse with an agent while moving. For example, the present technology can be applied to automobiles, buses, trucks, motorbikes, e-bikes, scooters, personal mobility, specific small vehicles (e.g., kick scooters), rail-running trains, airplanes, ships, flying cars, bicycles, etc.

[0288] The present technology can also be applied, for example, to a case where the user himself / herself is a moving body. That is, the present technology can also be applied to a case where the user wears the information processing system 201 and has a conversation with an agent while moving. In this case, the information processing system 201 is realized by, for example, one or more of a wearable device (e.g., a head-mounted display, a glasses-type display, etc.), a mobile information terminal (e.g., a smartphone, a portable game console, a portable music player, etc.), headphones, earphones, etc. Furthermore, the user wearing the information processing system 201 may move by himself / herself or by riding on another moving body.

[0289] <<4. Others>> <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.

[0290] FIG. 11 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0291] In the computer 1000 , a CPU (Central Processing Unit) 1001 , a ROM (Read Only Memory) 1002 , and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004 .

[0292] An input / output interface 1005 is further connected to the bus 1004. An input unit 1006, an output unit 1007, a storage unit 1008, a communication unit 1009, and a drive 1010 are connected to the input / output interface 1005.

[0293] The input unit 1006 includes input switches, buttons, a microphone, an image sensor, etc. The output unit 1007 includes a display, a speaker, etc. The storage unit 1008 includes a hard disk, a non-volatile memory, etc. The communication unit 1009 includes a network interface, etc. The drive 1010 drives removable media 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0294] In the computer 1000 configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program recorded in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0295] The program executed by the computer 1000 (CPU 1001) can be provided by being recorded on a removable medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0296] In the computer 1000, the program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting the removable medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.

[0297] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0298] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0299] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0300] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0301] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0302] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0303] <Examples of Combinations of Configurations> The present technology can also have the following configurations.

[0304] (1) An information processing device comprising: an information processing unit that controls the execution of a plurality of agents, each having different personalities, and each generating output information including utterance content related to a topic corresponding to at least one of an external situation and an internal situation of a moving object; and an output control unit that controls output of the output information generated by each of the agents. (2) The information processing device according to (1), wherein the internal situation of the moving object includes the situation of a user inside the moving object. (3) The information processing device according to (2), wherein the user situation includes the utterance content of the user, and the agent generates the output information including the utterance content corresponding to the utterance content of the user. (4) The information processing device according to (2) or (3), wherein the user situation includes a reaction of the user to the utterance content of the agent. (5) The information processing device according to any of (1) to (4), wherein the internal situation of the moving object includes the utterance content of another agent, and the agent generates the output information including the utterance content corresponding to the utterance content of the other agent. (6) The information processing device according to any one of (1) to (5), wherein the agent activates another agent based on at least one of an external situation and an internal situation of the moving object. (7) The information processing device according to any one of (1) to (6), wherein the output information includes text information indicating the utterance content, and the output control unit controls output of a speech sound of the utterance content based on the text information. (8) The information processing device according to (7), wherein the output control unit controls output of the speech sound using a voice corresponding to the agent that outputs the utterance content. (9) The information processing device according to (8), wherein the output control unit controls output position of the speech sound based on a position of a topic of conversation in the utterance content. (10) The information processing device according to any one of (1) to (9), wherein the output information includes image information including an image of a character corresponding to the agent, and the output control unit controls display of the character based on the image information.(11) The information processing device according to (10), wherein the output control unit controls the display position of the character based on the position of a topic in the utterance content. (12) The information processing device according to any of (1) to (11), wherein the information processing unit selects the output information to be presented to the user based on at least one of a situation outside the moving body, a situation inside the moving body, the user's preferences, and a conversation history between the user and the agent. (13) The information processing device according to (12), wherein the information processing unit selects the output information to be presented to the user based on at least one of a level of interest of the user in a topic corresponding to each piece of the output information and a degree of novelty of a topic corresponding to each piece of the output information for the user. (14) The information processing device according to (12) or (13), wherein the information processing unit learns the user's preferences based on a conversation history between the agent and the user. (15) The information processing device according to any of (1) to (14), wherein the information processing unit recognizes the situation outside and the situation inside the moving body. (16) The information processing device according to any one of (1) to (14), wherein the agent recognizes the external situation and the internal situation of the mobile body. (17) The information processing device according to any one of (1) to (16), wherein the agent generates the output information including utterance content on a topic corresponding to information acquired from the outside. (18) The information processing device according to any one of (1) to (17), wherein the mobile body is a user. (19) An information processing method comprising: a plurality of agents having different personalities each generating output information including utterance content on a topic corresponding to at least one of the external situation and the internal situation of the mobile body; and controlling output of the output information generated by each of the agents. (20) A program causing a computer to execute processing including: controlling execution of a plurality of agents having different personalities each generating output information including utterance content on a topic corresponding to at least one of the external situation and the internal situation of the mobile body; and controlling output of the output information generated by each of the agents.

[0305] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0306] 1 Vehicle, 11 Vehicle control system, 201 Information processing system, 211 Input unit, 212 External data acquisition unit, 213 Internal data acquisition unit, 214 Communication unit, 215 Conversation history DB, 216 User preference DB, 217 Control unit, 218 Voice output unit, 219 Display unit, 231 Information processing unit, 232 Output control unit, 241 Recognition unit, 242 Information generation unit, 243 Information selection unit, 244 Learning unit, 251-1 to 251-n Agent, 261 Voice synthesis unit, 262 Voice output control unit, 263 Display control unit

Claims

1. An information processing device comprising: an information processing unit that controls the execution of a plurality of agents each having different personalities and each generating output information including speech content related to a topic corresponding to at least one of the external and internal conditions of a moving object; and an output control unit that controls the output of the output information generated by each of the agents.

2. The information processing device according to claim 1, wherein the internal state of the mobile object includes the state of the user inside the mobile object.

3. The information processing device according to claim 2, wherein the user's situation includes the content of the user's utterance, and the agent generates the output information including the content of the utterance corresponding to the content of the user's utterance.

4. The information processing device according to claim 2, wherein the user's situation includes the user's reaction to the content of the agent's utterance.

5. The information processing device according to claim 1, wherein the internal situation of the moving body includes the speech content of the other agents, and the agent generates the output information including speech content corresponding to the speech content of the other agents.

6. The information processing device according to claim 1, wherein the agent activates other agents based on at least one of an external situation and an internal situation of the mobile object.

7. The information processing device according to claim 1, wherein the output information includes text information indicating the speech content, and the output control unit controls the output of speech sounds of the speech content based on the text information.

8. The information processing device according to claim 7, wherein the output control unit controls the output of the speech sound by a voice corresponding to the agent that outputs the speech content.

9. The information processing device according to claim 8, wherein the output control unit controls the output position of the speech sound based on the position of the topic of the speech content.

10. An information processing device according to claim 1, wherein the output information includes image information containing an image of a character corresponding to the agent, and the output control unit controls the display of the character based on the image information.

11. The information processing device according to claim 10, wherein the output control unit controls the display position of the character based on the position of the topic of the conversation in the speech content.

12. The information processing device according to claim 1, wherein the information processing unit selects the output information to be presented to the user based on at least one of the external situation of the mobile body, the internal situation of the mobile body, the user's preferences, and the conversation history between the user and the agent.

13. The information processing device according to claim 12, wherein the information processing unit selects the output information to be presented to the user based on at least one of the user's level of interest in the topic corresponding to each piece of output information and the degree of novelty of the topic corresponding to each piece of output information for the user.

14. The information processing device according to claim 12, wherein the information processing unit learns the preferences of the user based on a conversation history between the agent and the user.

15. The information processing device according to claim 1, wherein the information processing unit recognizes the external and internal conditions of the mobile object.

16. The information processing device according to claim 1, wherein the agent recognizes the external and internal conditions of the mobile object.

17. The information processing device according to claim 1, wherein the agent generates the output information including speech content relating to a topic corresponding to information acquired from an external source.

18. The information processing device according to claim 1, wherein the mobile object is a user.

19. An information processing method comprising: a plurality of agents each having different personalities generating output information including speech content on a topic corresponding to at least one of the external and internal conditions of a mobile object; and controlling the output of the output information generated by each of the agents.

20. A program for causing a computer to execute processes including: controlling the execution of a plurality of agents each having different personalities and each generating output information including speech content on a topic corresponding to at least one of the external and internal conditions of a mobile object; and controlling the output of the output information generated by each of the agents.

Citation Information

Patent Citations

  • System and method for context recognition conversation type agent based on machine learning, method, system and program of context recognition journaling method, as well as computer device

    JP2019121360A

  • Agent device, agent control method, and program

    WO2020070878A1