Information processing device, information processing method, and program

The information processing system accurately predicts user actions in moving vehicles by generating and processing texts about user, vehicle, and environment states, using a large language model to minimize computational load.

WO2026074604A1PCT designated stage Publication Date: 2026-04-09PIONEER IP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-01
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict user actions while minimizing computational requirements in information processing systems for moving bodies.

Method used

An information processing apparatus and method that generates texts related to user state, vehicle state, and surroundings, and uses a large language model to predict actions, with optimized input processing to reduce computational load.

Benefits of technology

Accurately predicts user actions with reduced computational resources by generating and processing relevant texts and utilizing a large language model effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024035070_09042026_PF_FP_ABST
    Figure JP2024035070_09042026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises a text generation unit and an input processing unit. The text generation unit performs at least one of: processing for generating a first text related to the state of a user on board a moving body; processing for generating a second text related to the situation of the moving body; and processing for generating a third text related to the situation around the moving body. The input processing unit executes processing for: generating instruction information that is for causing a large-scale language model to generate response information related to a predicted action, which is an action that may be performed by the user, and includes a text generated by the text generation unit from among the first text, the second text, and the third text; and inputting the instruction information into the large-scale language model.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing Method, and Program

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

[0002] In recent years, technologies for supporting users riding on a moving body using information processing technologies have been developed. For example, Patent Document 1 describes an apparatus for predicting a driver's behavior. This apparatus includes an acquisition unit, an estimation unit, a reference unit, and a prediction unit. The acquisition unit acquires a captured image obtained by a camera that captures the interior of the host vehicle. The estimation unit estimates the upper limb movement range of the driver based on the captured image. The reference unit refers to at least one of the traveling information of the host vehicle, the past behavior history of the driver, and the traveling environment information of the host vehicle. The prediction unit predicts the driver's behavior based on the estimation result of the estimation unit and the reference result of the reference unit.

[0003] Japanese Unexamined Patent Application Publication No. 2019-86932

[0004] In order to support a user, it is preferable to be able to accurately predict the actions that the user may perform. On the other hand, it is also necessary to reduce the amount of computation required for this prediction. An example of the object of the present invention is to enable an information processing apparatus to accurately predict the above-described actions and to reduce the amount of computation required for this prediction.

[0005] The invention according to claim 1 is an information processing apparatus including: a text generation unit that performs at least one of a process of generating a first text related to the state of a user riding on a moving body, a process of generating a second text related to the situation of the moving body, and a process of generating a third text related to the situation around the moving body; and an input processing unit that generates instruction information for causing a large language model to generate response information related to a predicted action, which is an action that the user may perform, the instruction information including the text generated by the text generation unit among the first text, the second text, and the third text, and executes a process for inputting the instruction information to the large language model.

[0006] The invention described in claim 10 is an information processing method in which a computer performs at least one of the following: a process of generating a first text relating to the state of a user riding in a mobile vehicle; a process of generating a second text relating to the state of the mobile vehicle; and a process of generating a third text relating to the state of the mobile vehicle's surroundings; and generates instruction information for a large-scale language model to generate response information relating to predictive actions that the user may take, the instruction information including the generated text from the first text, the second text, and the third text; and performs a process of inputting the instruction information into the large-scale language model.

[0007] The invention described in claim 11 is a program that provides a computer with: a text generation unit that performs at least one of the following: a process of generating a first text relating to the state of a user riding in a mobile vehicle; a process of generating a second text relating to the state of the mobile vehicle; and a process of generating a third text relating to the state of the mobile vehicle's surroundings; and an input processing unit that generates instruction information for a large-scale language model to generate response information relating to predictive actions that the user may take, the instruction information including the text generated by the text generation unit from among the first text, the second text, and the third text, and performs a process for inputting the instruction information into the large-scale language model.

[0008] This is a diagram illustrating the usage environment and functional configuration of the information processing device according to this embodiment. This is a diagram illustrating items that may be included in the first text, items that may be included in the second text, and items that may be included in the third text. This is a diagram illustrating an example of multiple candidate predictive actions. This is a diagram illustrating an example of instruction information. This is a diagram illustrating an example of processing performed by the response processing unit. This is a diagram illustrating an example of items displayed on the display. This is a diagram illustrating an example of processing performed by the response processing unit. This is a diagram illustrating an example of processing performed by the response processing unit. This is a diagram illustrating an example of the hardware configuration of the information processing device. This is a flowchart illustrating an example of processing performed by the information processing device according to this embodiment. This is a diagram illustrating the usage environment and functional configuration of the information processing device according to this embodiment. This is a diagram illustrating an example of a table used by the setting unit.

[0009] Embodiments of the present invention will be described below with reference to the drawings. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted as appropriate.

[0010] (First Embodiment) Figure 1 is a diagram illustrating the usage environment and functional configuration of the information processing device 10 according to this embodiment. The information processing device 10 performs processing to predict the actions of a user riding in a mobile vehicle 20, and also performs processing to support the user using the results of this processing. The mobile vehicle 20 is, for example, a car or a motorcycle, but is not limited thereto.

[0011] The information processing device 10 may be mounted on the mobile device 20 or located outside the mobile device 20. For example, the information processing device 10 can be incorporated as a function of a car navigation system or an in-vehicle device. The information processing device 10 may also communicate with an external device, such as a server, as needed. One example of such an external device is a device that performs route suggestion and route search and stores map information. This map information also contains information about facilities such as stores. Furthermore, the information processing device 10 may be realized by a combination of a device mounted on the mobile device 20 and a device outside the mobile device 20.

[0012] The information processing device 10 is used together with the sensor 210, the target device 220, and the model device 30. The sensor 210 and the target device 220 move together with the mobile body 20. The sensor 210 and the target device 220 are, for example, incorporated into the mobile body 20, but may also be incorporated into a portable communication device carried by the user. Alternatively, one of the sensor 210 and the target device 220 may be incorporated into the mobile body 20, and the other into the communication device. In addition, the mobile body 20 may have multiple sensors 210 and multiple target devices 220 incorporated into it. In this case, some of the multiple sensors 210 may be incorporated into the mobile body 20, and the remaining sensors 210 may be incorporated into the communication device. The same applies to the multiple target devices 220.

[0013] Sensor 210 repeatedly generates data used to predict user behavior and transmits the generated data to the information processing device 10. Hereinafter, the information generated by sensor 210 will be referred to as sensor information. Sensor 210 is, for example, at least one of the following. At least some of these examples may also be handled by the control unit that controls various devices of the mobile body 20. - Speedometer of the mobile body 20 - Acceleration sensor of the mobile body 20 - Sensor that detects the open / closed state of the doors and windows of the mobile body 20 - Sensor that detects the illuminated state of the lights (including fog lamps) for the front lighting of the mobile body 20 - Sensor that detects operations performed on the steering wheel of the mobile body 20 - Sensor that detects the current position of the mobile body 20 (e.g., GPS) - Sensor that acquires or detects information indicating at least one of the type of road the mobile body 20 is currently traveling on and traffic conditions (e.g., congestion status such as whether there is traffic congestion or whether there is construction) (e.g., navigation device) - Imaging device that photographs at least one of the surroundings of the mobile body 20, such as the front, rear, and side. - Imaging device that photographs the passenger space of the mobile body 20. - Clock - Thermometer for detecting ambient temperature around the mobile unit 20 - Thermometer for detecting ambient temperature in the passenger compartment of the mobile unit 20 - Microphone for detecting sound in the passenger compartment of the mobile unit 20 - Microphone for detecting ambient sound around the mobile unit 20

[0014] The target device 220 is, for example, at least one of the following: • A device that controls the opening and closing state of the windows of the mobile body 20; • Air conditioning equipment with temperature control function; • Navigation device; • Audio device; • Interactive device. It may also have the function of controlling at least one of the navigation device and the audio device. • Lights (including fog lamps) for the forward illumination of the mobile body 20.

[0015] When the information processing device 10 performs processing to predict user behavior, it may also use data indicating the operating status of the target device 220. Hereinafter, this information will be referred to as operating status data. The operating status data is generated, for example, by the target device 220. The target device 220 repeatedly generates the operating status data and transmits it to the information processing device 10.

[0016] The model device 30 performs processing using large language models (LLMs). The information processing device 10 generates instruction information, such as a prompt, which is input to the model device 30, and transmits this instruction information to the model device 30. The model device 30 inputs this instruction information into the large language model, obtains the response information generated by the large language model, and transmits it to the information processing device 10.

[0017] The information processing device 10 then performs processing using the response information. One example of this processing is to cause the target device 220 to perform a predetermined process. A specific example of this predetermined process will be described later.

[0018] The information processing device 10 may also serve as the model device 30.

[0019] The information processing device 10 includes a text generation unit 110, an input processing unit 120, and a response processing unit 130, and can utilize a storage unit 140. The storage unit 140 stores various types of information used by the information processing device 10. The storage unit 140 may be part of the information processing device 10 or may be located outside the information processing device 10.

[0020] The text generation unit 110 performs at least one of the following processes: generating a first text about the state of the user riding in the mobile body 20; generating a second text about the status of the mobile body 20; and generating a third text about the surrounding environment of the mobile body 20. For example, the text generation unit 110 acquires sensor information and processes this sensor information to generate at least one of the first text, the second text, and the third text.

[0021] As a first example, the text generation unit 110 generates at least one of a first text, a second text, and a third text by processing sensor information using a machine learning model. This machine learning model may be owned by the information processing device 10 or by an external device.

[0022] As a second example, if the sensor information includes a string (e.g., a number), the text generation unit 110 makes at least a part of this string at least a part of the first text. The second example can also be used when generating at least one of the second text and the third text.

[0023] As a third example, if the sensor information indicates the operation of a device mounted on the mobile body 20, the text generation unit 110 generates at least one of the second text and the third text by processing the sensor information according to a rule base.

[0024] As a fourth example, if the sensor information includes an image, the text generation unit 110 detects the user's movement by processing this image and generates first text by processing this movement according to a rule base.

[0025] As a fifth example, if the sensor information includes speech based on the user's utterance, the text generation unit 110 generates the first text by converting this speech into text.

[0026] The input processing unit 120 generates instruction information to be input to the large-scale language model used by the model device 30. This instruction information is used to cause the large-scale language model to generate response information regarding predictive actions that the user may take in the near future, and includes the text generated by the text generation unit 110 from among the first text, second text, and third text.

[0027] The instruction information may also include at least one of the following pieces of information: • The number of predicted actions to be included in the response information • Multiple candidate predicted actions

[0028] If the instruction information includes the number of predicted actions, the large-scale language model used by the model device 30 includes this number of predicted actions in the response information. The number of predicted actions is, for example, between 2 and 5, but is not limited to this. The number of predicted actions is set in advance by the user, for example. This setting may be done before the mobile body 20 moves, or it may be done while the mobile body 20 is moving. Information indicating this number is then stored, for example, in the storage unit 140.

[0029] If the instruction information includes multiple candidate predictive actions, the large-scale language model used by the model device 30 includes a selected predictive action in the response information. The multiple candidate predictive actions are set, for example, based on a rule base. This rule includes, for example, information that associates combinations of sensor information acquired by the information processing device 10 and the content of each sensor information with the multiple candidate predictive actions, and is stored, for example, in the storage unit 140.

[0030] The input processing unit 120 then performs processing to input the instruction information into the large-scale language model. One example of this processing is to send the instruction information to the model device 30. However, if the information processing device 10 also functions as the model device 30, this processing is to input the instruction information into the large-scale language model.

[0031] Furthermore, it is preferable that the text generation unit 110 generates all of the first text, second text, and third text. And it is preferable that the input processing unit 120 includes all of the first text, second text, and third text in the instruction information. Doing so improves the accuracy of the predicted behavior included in the response information generated by the large-scale language model.

[0032] The response processing unit 130 performs processing using the response information. One example of this processing is to cause the target device 220 to perform a predetermined process.

[0033] Figure 2 shows the items that may be included in the first text, the second text, and the third text. Each of the first, second, and third texts provides specific details about these items.

[0034] The information contained in the first text, i.e., the user's state, may include at least one of the following items: • Speech content • Predetermined actions such as yawning or stretching • Schedule

[0035] The schedule may be stored in the storage unit 140 in advance, or the information processing device 10 may obtain it from an external device that stores the user's schedule. The schedule may also indicate the time the user must arrive at their destination. In this case, the schedule is obtained from a navigation device that moves along with the mobile body 20.

[0036] The information contained in the second text, i.e., the status of the moving object, may include at least one of the following items: • Driving speed • Presence or absence of abnormal vehicle behavior • Status of lights • Open / closed status of windows and doors • Operating status of audio equipment (e.g., car stereo) • Operating status of air conditioning equipment • Presence and attributes of passengers • Current location • Type of road currently being traveled on (e.g., highway, toll road, or general road) • Presence and duration of traffic congestion • Time elapsed since departure

[0037] The information contained in the third text, i.e., the conditions surrounding the moving object, may include at least one of the following items. At least one of these items may be obtained from an external server, for example, a server that stores weather information or a server that stores road information. • Current time • Current weather • Current temperature • Road surface conditions • Noise level

[0038] Furthermore, road surface conditions can be determined from factors such as the level of road noise, vibration of the moving vehicle, analysis of camera images mounted on the vehicle that capture the road surface ahead, and reference to a real-time road condition map of the road where the vehicle is located.

[0039] Figure 3 shows an example of several candidate predictive actions that may be included in the instruction information. These candidates include, for example, at least two of the following:

[0040] - Recalculating the route on the navigation system - Obtaining traffic information - Searching for rest stops using the navigation system - Searching for restaurants using the navigation system - Engaging in casual conversation - Muting the audio output of the audio system - Operating the equalizer on the audio system - Activating the bass and treble boosting (loudness) on the audio system - Changing the wiper speed - Changing the temperature setting of the air conditioning system - Turning the hazard lights on or off - Turning the fog lights on or off - Starting or stopping the defogger - Switching between two-wheel drive and four-wheel drive - Searching for troubleshooting solutions

[0041] Figure 4 shows an example of instruction information generated by the input processing unit 120. This instruction information includes an instruction statement requesting the large-scale language model to output the action that the user is expected to take next, along with a first text, i.e., text indicating the "user's state", a second text, i.e., text indicating the "status of the moving object", and a third text, i.e., text indicating the "surroundings" of the moving object. This instruction information further includes the number of predicted actions to be included in the response information, and multiple "candidates" for the predicted actions.

[0042] In addition, an example of three predicted actions included in the response information output from the large language model corresponding to the instruction information shown in FIG. 4 is "1. Headlight lighting, 2. Wiper speed adjustment, 3. Air conditioner temperature adjustment". The order of these predicted actions is in the order of decreasing probability that the user will perform the predicted actions. That is, when the large language model can identify the probability that the predicted action will be performed in addition to the predicted action, the predicted actions included in the response information are preferably arranged in descending order of that probability. The instruction information shown in FIG. 4 also includes information instructing to output the predicted actions in the order of decreasing probability of occurrence.

[0043] Then, the response processing unit 130 causes the display to display display items corresponding to the predicted actions included in the response information. For example, when the response information includes "1. Headlight lighting, 2. Wiper speed adjustment, 3. Air conditioner temperature adjustment", the response processing unit 130 causes the display to display buttons for lighting the headlights, buttons for adjusting the speed of the wipers, and buttons for adjusting the temperature of the air conditioner. Note that the display may be a touch panel.

[0044] FIG. 5 is a diagram for explaining another example of the processing performed by the response processing unit 130. In the example shown in this figure, assume that the model device 30 generates response information including a predicted action of resetting the route in the car navigation device. Then, the response processing unit 130 selects the car navigation device as the target device 220, causes this car navigation device to perform a route re-search process, and causes this car navigation device to perform a process of displaying the result of the re-search on the display. For example, the car navigation device displays the route candidates identified as the result of the re-search together with a map on the display. At this time, the car navigation device may further display display items for operating the car navigation device, such as operation buttons.

[0045] In the example shown in this figure, the display items include a button for returning to the original route, a button for displaying the details of the route identified by re-search, and a button for canceling the re-search process itself. However, the display items may be buttons with other functions.

[0046] As shown in FIG. 6, the display items displayed in FIG. 5 may be shortcut icons of icons that are displayed after at least one icon is selected when starting from the top screen in normal operations. In other words, when the top screen is regarded as the first layer, the response processing unit 130 may be a shortcut icon of an icon located in the second layer and subsequent layers. By doing so, the user does not need to perform the operation from the top screen to reach that icon, and as a result, that icon can be selected immediately.

[0047] In addition, when the target device 220 is a device other than a car navigation device, for example, in the case of an audio device or air conditioning equipment, the icons described with reference to FIG. 6 may be displayed on the display.

[0048] For example, when the response information generated by the model device 30 includes the predicted action of changing the sound quality or volume of the audio device, the response processing unit 130 causes the audio device to display an icon for changing the sound quality or volume on the display as shown in FIG. 7.

[0049] Also, when the response information generated by the model device 30 includes the predicted action of changing the lighting state of the lights mounted on the moving body 20, the response processing unit 130 causes the control unit of the moving body 20 to display an icon for changing the lighting state of the lights on the display as shown in FIG. 8.

[0050] As explained using Figures 5 to 8, when the target device 220 controls a display visible to the user, one example of a predetermined process performed by the response processing unit 130 is to display a display item on the display for controlling equipment mounted on the mobile body 20 (including cases where it is the target device 220 or equipment other than the target device 220). One example of this display item is an icon indicating a button. This icon includes a shortcut icon for controlling equipment mounted on the mobile body 20.

[0051] The buttons as display items are not limited to the examples shown in Figures 5 to 8. For example, these display items may indicate that the user is interacting with a large-scale language model, or that the user is setting a rest stop as a waypoint or destination.

[0052] Another example of a predetermined process performed by the response processing unit 130 is, as shown in Figure 9, the output to an output device such as a display, which includes information indicating the process to be performed by the control unit of the mobile body 20 (for example, a button to switch to four-wheel drive mode) and information indicating the action to be taken by the user (for example, a button to display how to get out of being stuck). This information may also be displayed on the display as an icon such as a button. The screen shown in Figure 9 is displayed when the sensor information estimates that the mobile body 20 is stuck.

[0053] Other examples of information indicating the processing to be performed by the control unit of the mobile unit 20 may include, for example, switching from two-wheel drive to four-wheel drive, turning the automatic driving mode on and off, turning the auto cruise function on and off, and switching the driving mode (sport mode / normal mode).

[0054] Furthermore, information indicating actions that the user should take can also be considered adviceal information. This information may, for example, indicate actions that the user should take regarding the operation of the mobile vehicle 20, more specifically, the operation procedure for at least one of the accelerator, brake, and steering wheel.

[0055] Figure 10 shows an example of the hardware configuration of the information processing device 10. The information processing device 10 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, an input / output interface 1050, and a network interface 1060.

[0056] Bus 1010 is a data transmission path for the processor 1020, memory 1030, storage device 1040, input / output interface 1050, and network interface 1060 to send and receive data to and from each other. However, the method of connecting the processor 1020 and the other components to each other is not limited to bus connection.

[0057] Processor 1020 is a processor implemented using components such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).

[0058] Memory 1030 is a main memory device implemented as RAM (Random Access Memory), etc.

[0059] The storage device 1040 is an auxiliary storage device implemented as a removable media such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or memory card, or as ROM (Read Only Memory). The storage device 1040 stores program modules that implement each function of the information processing device 10 (for example, the text generation unit 110, the input processing unit 120, the response processing unit 130, and the determination unit 150 and setting unit 160, which will be described later). The processor 1020 reads each of these program modules into the memory 1030 and executes them, thereby realizing each function corresponding to that program module. The storage device 1040 may also function as a storage unit 140.

[0060] The input / output interface 1050 is an interface for connecting the information processing device 10 with various input / output devices. For example, the information processing device 10 may communicate with at least one of the sensor 210 and the target device 220 via the input / output interface 1050.

[0061] The network interface 1060 is an interface for connecting the information processing device 10 to a network. This network may be, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The method by which the network interface 1060 connects to the network may be wireless or wired. The information processing device 10 may communicate with at least one of the sensor 210 and the target device 220 via the network interface 1060.

[0062] Figure 11 is a flowchart illustrating an example of the processing performed by the information processing device 10. The information processing device 10 repeatedly performs the processing shown in this figure. First, the information processing device 10 acquires sensor information from the mobile body 20 (step S10). Then, the text generation unit 110 uses this sensor information to generate at least one of the first text, second text, and third text (step S20). Then, the response processing unit 130 generates instruction information and transmits this instruction information to the model device 30 (step S30).

[0063] The model device 30 inputs instruction information into a large-scale language model and obtains response information from this large-scale language model. The model device 30 then transmits this response information to the information processing device 10. The response processing unit 130 of the information processing device 10 obtains this response information (step S40). The response processing unit 130 then uses this response information to perform predetermined processing (step S50).

[0064] In this way, the information processing device 10 generates at least one of three texts: a first text concerning the state of the user riding in the mobile body 20, a second text concerning the status of the mobile body 20, and a third text concerning the surrounding environment of the mobile body 20, and generates instruction information including the generated text. The information processing device 10 then inputs this instruction information into a large-scale language model to obtain response information concerning predicted actions, which are actions that the user may take. Therefore, by using the information processing device 10, the user's actions can be predicted with high accuracy, and the amount of computation required for this prediction can be reduced.

[0065] As shown in Figure 12, the information processing device 10 may further include a determination unit 150.

[0066] For example, the determination unit 150 determines whether at least one of the user riding in the mobile body 20 and the mobile body 20 meets a predetermined condition. If this predetermined condition is met, the input processing unit 120 generates the instruction information described above and performs processing to input this instruction information into the large-scale language model. If this predetermined condition is not met, the input processing unit 120 does not generate the instruction information, and as a result does not perform processing to input the instruction information into the large-scale language model.

[0067] This reduces the number of times the large-scale language model generates response information, thereby further reducing the computational load required for prediction.

[0068] The predetermined conditions may include the user performing a predetermined action. The action performed by the user is identified, for example, by processing an image generated by the imaging device acting as sensor 210. An example of a predetermined action is an action performed when the user is fatigued, such as yawning or stretching a part of the body. However, the predetermined action is not limited to these.

[0069] The specified conditions may include the user uttering a specified statement. The specified statement may be, for example, a statement that suggests at least one of the user's physical and mental states is in a specified state, such as a statement that suggests the user is tired. The specified statement may also include a specific word. Furthermore, the specified statement may be a request to turn the target device 220 on or off or change its settings.

[0070] When the surrounding conditions of the mobile body 20 change significantly, users often change the on / off status or settings of the target device 220. Therefore, the predetermined conditions may include that the change in the surrounding conditions of the mobile body 20 meets a criterion. The criterion here includes, for example, a change in the type of road on which the mobile body 20 is traveling, and at least one of the following: the change in the ambient temperature, brightness, and sound level outside the mobile body 20 is greater than or equal to a criterion value.

[0071] Furthermore, when the mobile device 20 approaches its destination, the user often changes the on / off status or settings of the target device 220. Also, if the mobile device 20 is operating continuously for a predetermined period of time, that is, if the user has not been taking a break for a predetermined period of time, it would be advisable to make some kind of suggestion to the user, for example, via the navigation device. Therefore, the predetermined conditions may include at least one of the following: the distance from the current position of the mobile device 20 to the destination is less than or equal to a predetermined value, and the mobile device 20 is operating continuously for a predetermined period of time.

[0072] The determination unit 150 may also use the text generated by the text generation unit 110 to determine whether predetermined conditions have been met. For example, the determination unit 150 can use the first text to determine whether the user has made an utterance of predetermined content. The determination unit 150 can also use the second text to determine whether the mobile object 20 is approaching its destination. The determination unit 150 can also use the third text to determine whether the surrounding conditions of the mobile object 20 have changed significantly.

[0073] In this case, the text generation unit 110 repeatedly generates at least one of the first text, the second text, and the third text. The determination unit 150 then uses the generated text to determine whether predetermined conditions have been met each time text is generated. If the predetermined conditions are met, the processing from step S30 onwards in Figure 11 is performed.

[0074] Furthermore, as shown in Figure 13, the information processing device 10 may have a setting unit 160 in addition to the configuration shown in Figure 12.

[0075] In the example shown in this figure, the determination unit 150 determines, in place of or in addition to the process described using Figure 12, whether at least one of the following conditions is met: the state of the user riding on the mobile body 20, the state of the mobile body 20, or the state of the surroundings of the mobile body 20. The setting unit 160 then sets the items to be included in the instruction information according to the conditions that have been met. The input processing unit 120 then generates instruction information including the set items and executes processing to input the instruction information into the large-scale language model. For example, the input processing unit 120 generates instruction information that does not include any items other than the set items.

[0076] This reduces the amount of information contained in the instruction data. As a result, the computational complexity of the large-scale language model required for prediction can be further reduced.

[0077] The setting unit 160 includes at least one of the following items to be included in the instruction information: the user's status, the status of the mobile body 20, and the surrounding conditions of the mobile body 20. Specific examples of these are explained using Figure 2.

[0078] Furthermore, the setting unit 160 includes multiple candidate predictive actions in at least one of the items to be included in the instruction information. Specific examples of these are explained using Figure 3.

[0079] The setting unit 160 sets the items to be included in the instruction information using a list of items, such as those shown in Figures 2 and 3. For example, the setting unit 160 sets the items to be included in the instruction information by deleting candidates from this list according to the predetermined conditions that have been met.

[0080] Figure 14 shows an example of a table that associates "predetermined conditions" with "items to be removed from the list." For example, the setting unit 160 sets the items to be included in the instruction information by removing items corresponding to the satisfied predetermined conditions from the lists shown in Figures 2 and 3, according to this table. In Figure 14, the items to be removed for each satisfied predetermined condition are items that can be considered to have a generally weak relevance to the user's state, the state of the mobile body 20, or the surrounding environment of the mobile body 20, among multiple types of current situations and multiple types of predicted behavior candidates, as indicated by the predetermined conditions. In this way, the setting unit 160 can easily set the items to be included in the instruction information.

[0081] The embodiments of the present invention have been described above with reference to the drawings, but these are merely examples of the present invention, and various other configurations can also be adopted.

[0082] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above embodiments can be combined to the extent that their contents do not conflict.

[0083] 10 Information processing device 20 Mobile device 30 Model device 110 Text generation unit 120 Input processing unit 130 Response processing unit 140 Storage unit 150 Judgment unit 160 Setting unit 210 Sensor 220 Target device 230 Display

Claims

1. An information processing device comprising: a text generation unit that performs at least one of the following processes: generating a first text relating to the state of a user riding in a mobile vehicle; generating a second text relating to the status of the mobile vehicle; and generating a third text relating to the surrounding conditions of the mobile vehicle; and an input processing unit that generates instruction information for a large-scale language model to generate response information relating to predictive actions that the user may take, the instruction information including the text generated by the text generation unit from among the first text, the second text, and the third text, and performs a process for inputting the instruction information into the large-scale language model.

2. An information processing device according to claim 1, wherein the instruction information further includes the number of predicted actions to be included in the response information.

3. An information processing device according to claim 1 or 2, wherein the instruction information further includes a plurality of candidates for the predicted action.

4. An information processing apparatus according to any one of claims 1 to 3, wherein the text generation unit acquires sensor information generated by a sensor mounted on the moving body, and uses the sensor information to generate at least one of the first text, the second text, and the third text.

5. An information processing apparatus according to any one of claims 1 to 4, wherein the text generation unit generates all of the first text, the second text, and the third text, and the instruction information includes all of the first text, the second text, and the third text.

6. An information processing apparatus according to any one of claims 1 to 5, comprising a response processing unit that acquires the response information and causes a target device moving together with a moving object to perform a predetermined process using the response information.

7. An information processing apparatus according to claim 6, wherein the target device controls a display visible to the user, and the predetermined process includes displaying a display item for controlling equipment mounted on the mobile body on the display.

8. An information processing device according to claim 7, wherein the display item includes a shortcut icon for controlling a device mounted on the mobile body.

9. An information processing apparatus according to claim 6, wherein the predetermined processing includes causing an output device to output at least one of information indicating processing to be performed by the control unit of the mobile body and information indicating actions to be performed by the user.

10. An information processing method comprising: a computer performing at least one of the following: a process of generating a first text relating to the state of a user riding in a mobile vehicle; a process of generating a second text relating to the status of the mobile vehicle; and a process of generating a third text relating to the status of the mobile vehicle's surroundings; and an instruction information for a large-scale language model to generate response information relating to predictive actions that the user may take, the instruction information including the generated text from the first text, the second text, and the third text; and an instruction information input to the large-scale language model.

11. A program that provides a computer with: a text generation unit that performs at least one of the following: a process for generating a first text relating to the state of a user riding in a mobile vehicle; a process for generating a second text relating to the status of the mobile vehicle; and a process for generating a third text relating to the surroundings of the mobile vehicle; and an input processing unit that generates instruction information for causing a large-scale language model to generate response information relating to predictive actions that the user may take, the instruction information including the text generated by the text generation unit from among the first text, the second text, and the third text, and performs a process for inputting the instruction information into the large-scale language model.

Citation Information

Patent Citations

  • Efficient driver action prediction system based on time fusion of sensor data using deep (bidirectional) recurrent neural network

    JP2018028906A

  • Driver behavior prediction device

    JP2019086932A

  • Dialogue system using knowledge base and language model for automotive systems and application

    JP2024043564A