Information processing equipment, information processing programs, driver assistance systems, and vehicles

The information processing device enhances pedestrian behavior prediction by using a large-scale language model to estimate and verbalize pedestrian behavior, addressing inaccuracy and cost issues in existing models, ensuring accurate and adaptable driving assistance.

JP2026087259APending Publication Date: 2026-05-27DENSO TEN LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DENSO TEN LTD
Filing Date
2024-11-15
Publication Date
2026-05-27

Smart Images

  • Figure 2026087259000001_ABST
    Figure 2026087259000001_ABST
Patent Text Reader

Abstract

We provide technology that enables accurate prediction of human behavior. [Solution] An exemplary information processing device is an information processing device that generates prompts to be input to a large-scale language model, and it acquires image data, detects people in the acquired image data, estimates the behavior of the detected people, acquires a language expression that verbalizes the estimated behavior, and generates the prompt using the acquired language expression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , ,

[0005] , , ,

[0001] The present invention relates to a technique for predicting human behavior using large language models (LLMs).

Background Art

[0002] Conventionally, there is known a technique for predicting whether a vehicle will collide with a pedestrian based on information such as the position and posture of pedestrians around the vehicle, and notifying of the risk of collision (see, for example, Patent Document 1). For example, there is known a technique in which an orbit prediction model (AI (Artificial Intelligence) model) predicts the future orbit of a pedestrian by inputting time-series data of the position and posture of the pedestrian, and the risk of collision with a vehicle is determined based on the prediction.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The walking orbit of a pedestrian is affected by various factors such as, for example, the pedestrian's own situation, road surface conditions, surrounding conditions, or weather conditions. For example, if the walking orbit changes due to factors not considered in an orbit prediction model or the like, the prediction of the walking orbit becomes inaccurate, and a collision may occur.

[0005] <0000For example, one possible solution is to train a trajectory prediction model to address all factors that influence changes in pedestrian trajectories. However, there are a vast number of factors that influence changes in pedestrian trajectories. Therefore, there are concerns that the costs required for design and data collection to accurately predict trajectories will be extremely high. Furthermore, since new factors influencing changes in pedestrian trajectories may arise over time, it will be necessary to redesign or retrain the model each time a new factor emerges, which is a concern as it will increase costs.

[0006] In view of the above points, the present invention aims to provide a technology that enables accurate prediction of human behavior. [Means for solving the problem]

[0007] An exemplary information processing device of the present invention is an information processing device that generates a prompt to be input to a large-scale language model, which acquires image data, detects a person in the acquired image data, estimates the behavior of the detected person, acquires a language expression that verbalizes the estimated behavior, and generates the prompt using the acquired language expression. [Effects of the Invention]

[0008] According to an exemplary information processing device of the present invention, prompts generated based on the estimated behavior of a person present in an image can be input to a large-scale language model. For this purpose, information regarding the behavior of a person present in the image can be obtained from the large-scale language model. The information thus obtained from the large-scale language model can be suitable for predicting the behavior of a person present in an image, and the acquisition of such information can improve the accuracy of the prediction of human behavior. In other words, according to an exemplary version of the present invention, a technology that enables accurate prediction of human behavior can be provided. [Brief explanation of the drawing]

[0009] [Figure 1] Block diagram showing the general configuration of the information processing system (driving assistance system) [Figure 2] Block diagram showing the general configuration of the in-vehicle device. [Figure 3] Diagram showing a language conversion table [Figure 4] Flowchart illustrating the operation of an in-vehicle device [Figure 5] A flowchart illustrating the detailed flow of the process in step S2 in Figure 4. [Figure 6] A diagram illustrating prompts to input into a large-scale language model and the responses from the large-scale language model. [Figure 7] This figure shows an additional language conversion table that converts detailed information about walking patterns into additional language expressions. [Figure 8] A schematic diagram showing a prompt composed of text data and image data. [Modes for carrying out the invention]

[0010] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings.

[0011] <1. Information Processing Systems> Figure 1 is a block diagram showing the schematic configuration of the information processing system SYS according to an embodiment of the present invention. In this embodiment, the information processing system SYS is a driver assistance system that assists in the driving of vehicle 100. Hereinafter, the information processing system SYS will be referred to as the driver assistance system SYS. The driver assistance system SYS may also be understood as a collision prevention system that prevents vehicle 100 from colliding with a person.

[0012] As shown in Figure 1, the driver assistance system SYS comprises a server 1, an in-vehicle device 2, and a camera 3. The in-vehicle device 2 and camera 3 are mounted on the vehicle 100. In other words, the vehicle 100 is equipped with the in-vehicle device 2 and camera 3. A specific example of the vehicle 100 is an automobile.

[0013] Server 1 is located outside the vehicle 100. Server 1 is configured to communicate with the in-vehicle device 2 via a communication network (not shown), such as the Internet. Server 1 is, for example, a cloud server.

[0014] Server 1 is configured to execute processing using a Large-Scale Language Model (LLM) 11. The Large-Scale Language Model 11 is software configured to perform natural language processing according to a model trained using a large amount of text data. In detail, the functionality of the Large-Scale Language Model 11 is achieved when the processor in Server 1 executes processing based on the model information contained in the Large-Scale Language Model 11. The processor in Server 1 includes arithmetic circuits such as a CPU (Central Processing Unit). The model information contained in the Large-Scale Language Model 11 includes the structure and parameters of the Large-Scale Language Model 11, as well as code instructions for executing the Large-Scale Language Model 11. When the Large-Scale Language Model 11 receives prompts such as command statements or question statements generated by the in-vehicle device 2, it outputs answers to those prompts, such as answer statements, to the in-vehicle device 2.

[0015] The in-vehicle device 2 is provided to communicate with the camera 3. This communication may be wired or wireless. The in-vehicle device 2 acquires image data (image data) of images captured by the camera 3. The in-vehicle device 2 processes the acquired image data and performs various processes necessary for driving assistance. The various processes performed by the in-vehicle device 2 include generating prompts to be input to the large-scale language model 11. The various processes performed by the in-vehicle device 2 also include acquiring the response from the large-scale language model 11 to which the prompts have been input. In other words, the in-vehicle device 2 generates prompts to be input to the large-scale language model 11 and acquires the response from the large-scale language model 11 to which the prompts have been input. Furthermore, the various processes performed by the in-vehicle device 2 include driving assistance processing based on the response acquired from the large-scale language model 11. Details of the various processes performed by the in-vehicle device 2 will be described later.

[0016] The camera 3 captures the surroundings of the vehicle 100. In the present embodiment, the camera 3 is used for detecting a person existing in the surroundings of the vehicle 100. The camera 3 is used, for example, for detecting a pedestrian who is about to cross the road. The camera 3 is preferably disposed on the vehicle 100 so as to be able to capture at least the front of the vehicle 100. However, the camera 3 may be plural instead of single, and in addition to the camera 3 that captures the front, a camera 3 that captures the rear or the side may be disposed on the vehicle 100. The camera 3 may be, for example, a camera included in a drive recorder.

[0017] The schematic configuration of the information processing system SYS of the present embodiment configured as a driving support system is as described above. However, the information processing system of the present invention may be other than a driving support system for the vehicle 100. The information processing system may be configured, for example, as a collision prevention system for preventing a collision with a person in an autonomous robot. In this case, the above vehicle 100 may be replaced with an autonomous robot, and the above in-vehicle device 2 may be replaced with a robot-mounted device.

[0018] Also, in the above, the driving support system SYS is configured to include the server 1, but this is merely an example. The driving support system may be configured by the in-vehicle device 2 and the camera 3 provided in the vehicle 100, and may be configured such that the driving support system configured in this way exchanges information with a server provided outside the system.

[0019] <2. Configuration of In-Vehicle Device> The configuration of the in-vehicle device 2 included in the driving support system SYS will be described in detail. FIG. 2 is a block diagram showing a schematic configuration of the in-vehicle device 2 according to an embodiment of the present invention. In FIG. 2, the components necessary for explaining the features of the in-vehicle device 2 according to the embodiment are shown, and the description of general components is omitted.

[0020] As shown in Figure 2, the in-vehicle device 2 includes an information processing device 2a. The information processing device 2a is configured using a computer device and has the functions of generating prompts to be input to the large-scale language model 11 and performing behavioral prediction to predict human behavior.

[0021] The in-vehicle device 2 includes an information processing device 2a, a speaker 2b, and a display device 2c. The information processing device 2a is configured to communicate with both the speaker 2b and the display device 2c. This communication may be wired or wireless. The speaker 2b outputs sound under the control of the information processing device 2a. The display device 2c displays information under the control of the information processing device 2a. The speaker 2b and the display device 2c are positioned in appropriate locations within the vehicle 100. For example, the speaker 2b and the display device 2c are positioned around the driver's seat where the driver of the vehicle 100 sits. The display device 2c may also have an operating function that allows the driver or other occupants of the vehicle 100 to operate the in-vehicle device 2. Such a display device 2c with an operating function may be configured as a touch panel.

[0022] As shown in Figure 2, the information processing device 2a comprises a controller 21, a memory 22, and a communication unit 23. The controller 21, memory 22, and communication unit 23 are connected to each other so as to be able to communicate with one another. This communication may be wired or wireless.

[0023] The controller 21 is a computer device comprising an arithmetic circuit that performs calculation processing. More specifically, the controller 21 includes a processor that performs calculation processing, etc. The processor includes, for example, a CPU. The controller 21 may consist of one processor or multiple processors. If it consists of multiple processors, these processors only need to be connected to each other in a way that allows them to communicate with one another.

[0024] Memory 22 consists of volatile memory and non-volatile memory. Volatile memory is specifically RAM (Random Access Memory). Non-volatile memory is specifically ROM (Read Only Memory). Non-volatile memory may also include flash memory or a hard disk drive, etc. Non-volatile memory stores computer-readable programs (computer programs) 221 and data.

[0025] The communication unit 23 is configured as a communication interface having an interface circuit for connecting to a communication network (not shown), such as the Internet. The in-vehicle device 2 is able to communicate with the server 1 by having the communication unit 23.

[0026] As shown in Figure 2, the controller 21 includes, as its functions, an information acquisition unit 211, a person detection unit 212, a behavioral pattern estimation unit 213, a language expression acquisition unit 214, a prompt generation unit 215, a behavior prediction unit 216, and a driving support unit 217. The functions of the controller 21 are realized by the processor executing calculations according to a program 221 stored in memory 22. The program 221 that realizes the functions of the controller 21 may consist of a single program or multiple programs.

[0027] The program 221 stored in memory 22 may be provided, for example, on a computer-readable non-volatile recording medium. The non-volatile recording medium may be, for example, an optical recording medium (e.g., an optical disc), a magneto-optical recording medium (e.g., a magneto-optical disc), a USB memory, or an SD card, in addition to the non-volatile memory described above. As another example, the program 221 stored in memory 22 may be provided from a program provision server via a communication line such as the Internet (a configuration provided by so-called download).

[0028] Furthermore, in this embodiment, the functions of the controller 21 are realized by software, i.e., by the arithmetic circuit (processor) executing arithmetic processing according to the program, but this is an example, and it may be realized by other methods. At least some of the functions of the controller 21 may be realized using, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). In other words, at least some of the functions of the controller 21 may be realized by hardware using a dedicated IC or the like. Also, at least some of the functions of the controller 21 may be realized by using both software and hardware.

[0029] Furthermore, each of the functional units 211 to 217 provided by the controller 21 is a conceptual component. The function performed by one component may be distributed among multiple components. Alternatively, the functions of multiple components may be integrated into a single component.

[0030] The information acquisition unit 211 acquires information from the memory 22 and other devices that are configured to communicate with the in-vehicle device 2. Other devices configured to communicate with the in-vehicle device 2 include the camera 3 and the server 1. Specifically, the information processing device 2a (in-vehicle device 2) acquires image data from the camera 3. More specifically, the information acquisition unit 211 periodically acquires image data from the camera 3. The information processing device 2a (in-vehicle device 2) also acquires responses (output information) from the large-scale language model 11. More specifically, when a prompt generated by the information processing device 2a (in-vehicle device 2) is input to the large-scale language model 11, the information acquisition unit 211 acquires a response to the prompt from the large-scale language model 11.

[0031] The person detection unit 212 performs a process to detect people in each image data acquired by the information acquisition unit 211, and detects a person if one exists in the image. In other words, the information processing device 2a (in-vehicle device 2) detects people in the acquired image data. The process of detecting people in an image may be a process using a known object detection method, and more specifically, it may be a process using an object detection model, which is an AI model trained by machine learning such as deep learning. In the person detection process using an object detection model, if a person is present in the image when image data is input to the object detection model, detection information indicating the detection of a person is output from the object detection model. The detection information includes position information that indicates the position (coordinates) of the person in the image. In the person detection process using an object detection model, if there are multiple people in the image, detection information for multiple people is acquired.

[0032] Furthermore, in object detection processing using an object detection model, it is possible to distinguish between people to be detected, for example, pedestrians and people riding mobility devices such as bicycles or electric scooters (as separate classes). Taking this into consideration, in this embodiment, the person detection performed by the person detection unit 212 is, as an example, the detection of pedestrians. For this reason, below, people detected by the person detection unit 212 will be referred to as pedestrians. However, the people detected by the person detection unit 212 may include not only pedestrians but also people riding mobility devices, and there is no intention to exclude such forms from the scope of the technical concept of the present invention.

[0033] Furthermore, while all pedestrians detected using the object detection model may be processed by the behavior estimation unit 213 described below, this embodiment does not employ such a configuration. In this embodiment, for the purpose of reducing the subsequent processing burden, among the pedestrians detected using the object detection model, only those pedestrians that meet specific requirements are detected as targets for processing by the behavior estimation unit 213. Pedestrians that meet the specific requirements are those whose trajectory is difficult to predict. The method for detecting pedestrians that meet the specific requirements will be described later.

[0034] The behavioral pattern estimation unit 213 estimates the behavioral patterns of pedestrians (more specifically, pedestrians who meet specific requirements) detected by the person detection unit 212. In other words, the information processing device 2a (in-vehicle device 2) estimates the behavioral patterns of detected people (in a detailed example, pedestrians). The behavioral patterns of pedestrians are, in detail, walking patterns. If the detection targets of the person detection unit 212 include people riding in mobility devices, the behavioral pattern estimation unit 213 will also estimate the behavioral patterns of people riding in mobility devices. The estimation of the behavioral patterns of people riding in mobility devices may be an estimation of the behavioral patterns of people including mobility devices.

[0035] The estimation of walking patterns specifically involves estimating the pedestrian's posture within each image data, and obtaining parameter values ​​for walking pattern parameters from the pedestrian's movement obtained by collecting the estimated pedestrian postures from each image data in a time series.

[0036] For estimating a pedestrian's posture, known posture estimation methods can be used, which involve identifying the position of each joint (joint point) of the multiple joints the pedestrian possesses within an image. Such posture estimation methods may be performed using a posture estimation model, which is an AI model trained using machine learning such as deep learning. In a configuration using a posture estimation model, by inputting image data into the posture estimation model, prediction results of the position of each joint (joint point) of the pedestrian within the image are output.

[0037] The parameter values ​​of the gait pattern parameters can be determined, for example, based on the movement of the pedestrian's joint points in multiple image data collected in a time series. The gait pattern parameters are a concept established to numerically represent the pedestrian's gait pattern, and in this embodiment, they include multiple types of parameters. Examples of multiple types of parameters include forward lean, gaze intensity, arm swing, and leg lift. The types and number of parameters included in the gait pattern parameters can be determined arbitrarily.

[0038] For example, the forward lean is a parameter that represents the degree to which a pedestrian is leaning forward. The larger the forward lean value (parameter value), the more the pedestrian is leaning forward. The forward lean value can be determined, for example, from the relationships between a predetermined set of joint points among several joint points identified by a posture estimation model. The forward lean value can be obtained from a single image data, but it may also be obtained by averaging the forward lean values ​​obtained from each of multiple time-series image data.

[0039] Furthermore, gaze intensity is a parameter that represents the degree to which a pedestrian is paying attention to a particular direction. The higher the gaze intensity value (parameter value), the more the pedestrian is paying attention to that particular direction. The gaze intensity value can be determined, for example, from the movement of joint points in the head over time in multiple time-series image data.

[0040] Furthermore, the arm swing degree is a parameter that represents the magnitude of a pedestrian's arm swing. The larger the arm swing degree value (parameter value), the larger the pedestrian's arm swing. The arm swing degree value can be determined, for example, from the range of motion of the joint points of the arm in the anterior-posterior direction in multiple image data obtained in a time series.

[0041] Furthermore, the leg lift value is a parameter that represents the degree to which a pedestrian lifts their leg. The larger the leg lift value (parameter value), the greater the pedestrian's leg lift. The leg lift value can be determined, for example, from the range of vertical movement of the joint points of the foot in multiple image data obtained in a time series.

[0042] In this embodiment, the parameter values ​​of the gait pattern parameters are calculated from the pose estimation results obtained using a pose estimation model, but this is merely an example. For example, the parameter values ​​of the gait pattern parameters may be obtained using an AI model that outputs gait pattern parameter values ​​based on time-series data of pedestrian images.

[0043] The language expression acquisition unit 214 performs a conversion process to verbalize the pedestrian's walking pattern estimated by the behavior pattern estimation unit 213, and acquires a language expression of the walking pattern. In other words, the information processing device 2a (in-vehicle device 2) acquires a language expression that verbalizes the estimated behavior pattern (in a detailed example, the walking pattern). More specifically, the language expression acquisition unit 214 performs a conversion process to a language expression based on the acquired parameter values ​​for various parameters such as the degree of forward lean. In this embodiment, the language expression acquisition unit 214 performs the conversion process to a language expression using a pre-prepared language conversion table. The language conversion table is stored in memory 22 (see Figure 2).

[0044] Figure 3 shows a language conversion table 222 according to an embodiment of the present invention. The information items in the language conversion table 222 include walking pattern parameters, range, and language expression.

[0045] The item "Gait Pattern Parameters" stores type information for multiple types of parameters included in the gait pattern parameters. This type information includes, for example, forward lean, gaze intensity, arm swing, and leg lift. For each parameter type, the "range" and "verbal expression" are stored.

[0046] The "Range" field stores a set numerical range for classifying the parameter values ​​obtained by the behavioral pattern estimation unit 213. In the example shown in Figure 3, for each type of parameter, three numerical ranges are set: a first range less than 0.2, a second range greater than 0.2 and less than or equal to 0.8, and a third range greater than 0.8. In the example shown in Figure 3, it is assumed that the parameter values ​​for each type of parameter are between 0 and 1. Also, in the example shown in Figure 3, the numerical ranges set for classifying the parameter values ​​are the same for all types of parameters, but this is just an example, and different settings may be used depending on the type of parameter. The number of classifications is not limited to three and may be changed as appropriate.

[0047] The "Linguistic Expression" field stores linguistic expressions appropriate for each group classified under the "Range" field. The appropriate linguistic expression differs depending on the type of parameter and is determined in advance by the designer. Note that in Figure 3, the parts where the linguistic expression is "-" indicate that the pedestrian's walking pattern is estimated to be typical based on the parameter values, and therefore no linguistic expression is specifically created.

[0048] For example, suppose the behavioral pattern estimation unit 213 obtained a forward lean of 0.9, a gaze intensity of 0.9, an arm swing intensity of 0.1, and a leg lift intensity of 0.5. In this case, the language expression acquisition unit 214 obtains the language expressions for the pedestrian's walking pattern as "leaning forward," "looking in a specific direction," and "not swinging arms" through a conversion process using the language conversion table 222. If it is estimated that the pedestrian is looking in a specific direction based on the gaze intensity value, the direction in which the pedestrian is looking may be further estimated and expressed in language, for example, based on the posture estimation result. In this case, instead of "looking in a specific direction," language expressions such as "looking down," "looking at hands," or "looking up" may be obtained.

[0049] The prompt generation unit 215 generates prompts to be input to the large-scale language model 11 using the language expression acquired by the language expression acquisition unit 214. In other words, the information processing device 2a (in-vehicle device 2) generates prompts using the acquired language expression. In this configuration, prompts generated based on the estimated walking patterns of pedestrians present in the image can be input to the large-scale language model 11. For this purpose, information about the behavior of pedestrians present in the image can be acquired from the large-scale language model 11. The information thus obtained from the large-scale language model 11 can be made suitable for predicting pedestrian behavior by appropriately setting walking pattern parameters and expressing the walking patterns in appropriate language. By acquiring information suitable for predicting pedestrian behavior, it is expected that the prediction of pedestrian behavior can be made accurately.

[0050] In detail, the prompt generation unit 215 generates a prompt using the acquired language expression and a prompt generation template stored in memory 22 beforehand. For example, suppose the language expression acquired by the language expression acquisition unit 214 is "bent over". In this case, the prompt generation unit 215 generates a prompt such as, "What state is a pedestrian walking with their head bent over? What kind of behavioral characteristics do such a pedestrian exhibit?"

[0051] Furthermore, it is preferable that the generated prompts are structured in such a way that the answer to "what kind of actions do pedestrians take" can be obtained from the large-scale language model 11. Taking this into consideration, in the example above, the prompt generation template includes a question such as, "What kind of behavioral characteristics do such pedestrians typically exhibit?"

[0052] Furthermore, the generated prompts may include not only questions about walking patterns but also additional questions that inquire about matters related to walking patterns. Examples of additional questions include questions about road conditions (a detailed example being road surface conditions). A specific example is generating a prompt such as, "What conditions can be expected for a pedestrian walking without lifting their feet, and for the road surface in such a case?" By generating such prompts, it is possible to obtain not only the characteristics of the pedestrian's walking but also information about the surrounding roads.

[0053] Furthermore, additional questions may or may not be included in the prompt, depending on the situation. For example, if unfavorable weather information for driving vehicle 100, such as rain, snow, or icy roads, is obtained via a communication network such as the internet, the prompt may be configured to include questions about road conditions only. In such cases, separate prompt generation templates should be prepared for cases that include additional questions and cases that do not.

[0054] The behavior prediction unit 216 predicts pedestrian behavior based on the response from the large-scale language model 11 that receives the prompt generated by the prompt generation unit 215. In other words, the information processing device 2a (in-vehicle device 2) predicts behavior based on the response from the large-scale language model 11 that receives the generated prompt.

[0055] This configuration makes it possible to predict pedestrian behavior by considering the diverse characteristics and conditions of pedestrians and roads, thus improving the accuracy of behavior prediction. In particular, this configuration makes it possible to modify pedestrian trajectory predictions using conventional trajectory prediction models by incorporating pedestrian and road information obtained from the large-scale language model 11. This improves the accuracy of pedestrian behavior prediction. Furthermore, this configuration is cost-effective because it improves the accuracy of pedestrian behavior prediction without requiring the trajectory prediction model to be trained to handle all factors that affect changes in pedestrian trajectories. In addition, with this configuration, as long as the large-scale language model 11 is updated, even if new factors affecting changes in pedestrian trajectories arise over time, accurate behavior prediction can be made without redesigning or retraining the trajectory prediction model.

[0056] In this embodiment, the behavior prediction unit 216 predicts the pedestrian's trajectory using a trajectory prediction model. The trajectory prediction model may be a known, trained AI model. The trajectory prediction model outputs a predicted trajectory, which is the pedestrian's future trajectory, by inputting, for example, time-series data of the pedestrian's coordinates and posture within an image. The pedestrian's posture can be obtained, for example, by the posture estimation model described above.

[0057] When the large-scale language model 11 is used, the behavior prediction unit 216 predicts pedestrian behavior by appropriately correcting the predicted trajectory obtained by the trajectory prediction model based on the response obtained from the large-scale language model 11. When the large-scale language model 11 is not used, the behavior prediction unit 216 predicts pedestrian behavior using the trajectory prediction model. The cases in which the large-scale language model 11 is used and when it is not will be described later.

[0058] The correction of the predicted trajectory described above is performed, for example, as follows: Natural language processing is performed on the response (text data) obtained from the large-scale language model 11 using a small-scale language model to extract feature information (text data) related to the pedestrian's walking motion. Feature information related to walking motion includes, for example, information such as "slowly," "fast," and "at an unstable speed." A conversion table (not shown) that associates correction coefficients with each of these feature information items is prepared in advance and stored in memory 22. The correction coefficients are obtained using the extracted feature information and the conversion table. The predicted trajectory is corrected by performing calculations using the obtained correction coefficients. For example, if "slowly" is extracted as feature information, the predicted trajectory is corrected by calculations using the correction coefficients so that the distance reached by the pedestrian at each time is shortened. Also, for example, if "at an unstable speed" is extracted, multiple predicted trajectories are obtained using multiple correction coefficients, each of which is prepared to determine patterns where the distance reached by the pedestrian at each time is shortened and patterns where it is lengthened.

[0059] As another example, the behavior prediction unit 216 may predict the pedestrian's behavior based on an improved trajectory prediction model, which is an improved version of a known trajectory prediction model, and the responses obtained from the large-scale language model 11. In the improved trajectory prediction model, in addition to time-series data of the pedestrian's coordinates and posture, feature variables related to the pedestrian's walking are input, and a trajectory prediction result that reflects the responses from the large-scale language model 11 is output. The values ​​of the feature variables related to walking can be obtained from the responses of the large-scale language model 11 using feature information (such as "slowly") obtained by the same method as described above, and a pre-prepared conversion means (such as a table).

[0060] The driving support unit 217 executes processing to support the driving of the vehicle 100 based on the pedestrian behavior prediction results from the behavior prediction unit 216. By executing the processing to support driving, driving support by the in-vehicle device 2 is realized. In other words, the in-vehicle device 2 provides driving support based on behavior prediction. In this embodiment, because pedestrian behavior can be accurately predicted as described above, appropriate driving support for the vehicle 100 can be provided.

[0061] The driver assistance provided by the in-vehicle device 2 in response to processing performed by the driver assistance unit 217 may include collision prevention to prevent the vehicle 100 from colliding with a pedestrian. More specifically, the driver assistance provided by the in-vehicle device 2 may include notifying the driver of the vehicle 100 of a risk such as a collision when such a risk is detected based on predictions of pedestrian behavior. The risk notification may be, for example, an audio notification using speaker 2b or a screen notification using display device 2c. Furthermore, the driver assistance provided by the in-vehicle device 2 may include avoiding the risk by using automatic braking or automatic steering when a risk such as a collision is detected based on predictions of pedestrian behavior. In addition, the driver assistance provided by the in-vehicle device 2 may include notifying the driver of the vehicle 100 of the response from the large-scale language model 11 directly using voice or screen display.

[0062] <3. Operation of on-board devices> Next, the operation of the in-vehicle device 2 configured as described above will be explained. Figure 4 is a flowchart illustrating the operation of the in-vehicle device 2. More specifically, Figure 4 is a flowchart illustrating the flow of processing for driving assistance performed by the information processing device 2a provided in the in-vehicle device 2. This flowchart shows the technical content of a computer program (program 221) that enables the computer to implement the driving assistance method of this embodiment. More specifically, the driving assistance method of this embodiment includes a method for generating prompts to be input to the large-scale language model 11 and a behavior prediction method for predicting the behavior of pedestrians (people). That is, the flowchart shown in Figure 4 includes the technical content of a computer program that enables the computer to implement the prompt generation method and the technical content of a computer program that enables the behavior prediction method of this embodiment.

[0063] The above computer program can be stored on various non-volatile recording media readable by a computer and provided (sold, distributed, etc.). The above computer program may consist of a single program, or it may consist of multiple programs working together.

[0064] The process shown in Figure 4 is executed as appropriate when the vehicle 100 is powered on and the onboard device 2 becomes capable of performing driver assistance processing.

[0065] In step S1, the controller 21 (information acquisition unit 211) acquires image data from the camera 3. In other words, program 221 makes the computer function as a means to acquire image data. Once the image data is acquired, the process proceeds to the next step S2. The information acquisition unit 211 acquires image data periodically. Each time image data is acquired, the processing from step S2 onward is performed.

[0066] In step S2, the controller 21 (person detection unit 212) detects pedestrians in the acquired image data. That is, the program 221 causes the computer to function as a means to detect people (pedestrians) in the acquired image data. In this embodiment, the person detection unit 212 does not simply detect pedestrians in the image, but detects them in groups. In detail, the person detection unit 212 distinguishes between pedestrians that meet certain requirements and pedestrians that do not meet those requirements. This will be explained with reference to Figure 5. Figure 5 is a flowchart illustrating the detailed flow of the process in step S2 in Figure 4.

[0067] In step S21, the controller 21 (person detection unit 212) determines whether or not a pedestrian is present in the image data obtained. Specifically, pedestrian detection processing is performed using an object detection model. If it is determined that a pedestrian is present (Yes in step S21), the process proceeds to the next step S22. If it is determined that a pedestrian is not present (No in step S21), the process returns to step S1 as shown in Figure 4.

[0068] In step S21, if it is determined that a pedestrian is present, there may be multiple pedestrians. If there are multiple pedestrians, the process described below will be performed for each pedestrian.

[0069] In step S22, the controller 21 (person detection unit 212) predicts the future trajectory of the pedestrian. The trajectory prediction can be performed, for example, using the trajectory prediction model described above. Note that trajectory prediction requires time-series data (image data) of the pedestrian. For this reason, in detail, the trajectory prediction is performed after securing a predetermined amount (time) of time-series data. Once the trajectory prediction is complete, the process proceeds to the next step S23.

[0070] In step S23, the controller 21 (person detection unit 212) calculates the degree of deviation between the measured trajectory and the predicted trajectory. The measured trajectory is the trajectory the pedestrian actually traveled from the start time in the predicted trajectory until a predetermined time has elapsed. The degree of deviation is obtained, for example, by comparing the pedestrian's actual coordinate position at the point when the predetermined time has elapsed from the start time in the predicted trajectory with the coordinate position in the predicted trajectory, and calculating the difference (difference, etc.) between them. Once the degree of deviation is calculated, the process proceeds to the next step S24.

[0071] In step S24, the controller 21 (human detection unit 212) determines whether the calculated deviation is greater than a preset threshold. If the deviation is greater than a preset value (Yes in step S24), the process proceeds to the next step S25. If the deviation is less than or equal to a preset value (No in step S24), the process proceeds to step S26.

[0072] In step S25, the controller 21 (person detection unit 212) determines that the pedestrians detected in step S21 are pedestrians that meet specific requirements (hereinafter referred to as "specific pedestrians"). Once the processing in step S25 is complete, the process proceeds to step S3 in Figure 4. As mentioned above, if there are multiple pedestrians detected in step S21, the processing from step S22 onwards is performed for each pedestrian, so there may be multiple specific pedestrians. If there are multiple specific pedestrians, the processing from step S3 onwards in Figure 4 is executed for each detected specific pedestrian.

[0073] In step S26, the controller 21 (person detection unit 212) determines that the pedestrian detected in step S21 is a general pedestrian who does not meet specific requirements (hereinafter referred to as a general pedestrian). Once the processing in step S26 is complete, the process proceeds to step S3 in Figure 4. Note that, as with specific pedestrians, there may be multiple general pedestrians. If there are multiple general pedestrians, the processing from step S3 onwards in Figure 4 is executed for each detected general pedestrian.

[0074] As can be seen from the above, in this embodiment, a person (pedestrian) whose degree of deviation between the predicted trajectory based on acquired image data and the actual trajectory is greater than a predetermined set value is determined to be a person (pedestrian) that meets specific requirements. With this configuration, it is possible to perform pedestrian behavior prediction using the large-scale language model 11, focusing on pedestrians for whom trajectory prediction is difficult. As a result, it is possible to accurately predict pedestrian behavior while suppressing an increased processing load on the in-vehicle device 2 (information processing device 2a).

[0075] Returning to Figure 4, in step S3, the controller 21 (behavior estimation unit 213) determines whether the pedestrian detected in step S2 is a specific pedestrian. If it is a specific pedestrian (Yes in step S3), the process proceeds to the next step S4. If it is not a specific pedestrian (No in step S3), the process proceeds to step S9.

[0076] In step S4, the controller 21 (behavior estimation unit 213) estimates the walking pattern of a specific pedestrian based on the image data. That is, program 221 makes the computer function as a means to estimate the behavior (walking pattern) of the detected person (specific pedestrian). As described above, the walking pattern is estimated using, for example, a posture estimation model, and as the final processing result, parameter values ​​of walking pattern parameters (see, for example, Figure 3) are obtained. Specifically, values ​​such as the degree of forward lean, the degree of gaze, the degree of arm swing, and the degree of leg lift are obtained. Once the parameter values ​​of the walking pattern parameters are obtained, the process proceeds to the next step S5.

[0077] In step S5, the controller 21 (language expression acquisition unit 214) acquires a language expression that verbalizes the walking pattern of the pedestrian (specific pedestrian) obtained in the previous step S4. In other words, program 221 makes the computer function as a means to acquire a language expression that verbalizes the estimated behavior pattern (walking pattern). As described above, the verbalization of the walking pattern is performed using a language conversion table 222 (see Figure 3) which is a table that shows the relationship between the parameter values ​​of the walking pattern parameters and the language expression. Once the language expression of the walking pattern is acquired, the process proceeds to the next step S6.

[0078] In step S6, the controller 21 (prompt generation unit 215) generates a prompt using the language expression acquired in step S5. In other words, program 221 causes the computer to function as a means to generate a prompt using the acquired language expression. As described above, prompt generation is performed using the acquired language expression and a prompt generation template stored in memory 22. Once the prompt is generated, the process proceeds to the next step S7.

[0079] In step S7, the controller 21 (prompt generation unit 215) sends the generated prompt (text data) to the server 1 via a communication network such as the Internet. Once the prompt transmission is complete, the process proceeds to the next step S8.

[0080] Server 1, upon receiving the prompt, inputs the prompt into the Large-Scale Language Model (LLM) 11 and executes processing using the LLM 11. The LLM 11, upon receiving the prompt, outputs a response (text data) to the prompt. Server 1 then transmits the response output from the LLM 11 to the in-vehicle device 2 (information processing device 2a).

[0081] In step S8, the controller 21 (information acquisition unit 211) obtains a response from the large-scale language model (LLM) 11. Once the response from the large-scale language model 11 is obtained, the process proceeds to the next step S9.

[0082] Now, referring to Figure 6, we will explain the flow from step S5 to step S8 described above with a specific example. Figure 6 is an example of a prompt to be input to the large-scale language model 11 and a response from the large-scale language model 11.

[0083] In the example shown in Figure 6, the language expression acquisition process in step S5 acquires the language expressions "lean forward," "look at your hands," and "don't swing your arms." Then, in step S6, the prompt generation unit 215 uses these three language expressions to generate the prompt shown in Figure 6 (step S6). The generated prompt is sent to the server 1 (step S7).

[0084] Server 1, which generates a prompt, uses the large-scale language model 11 to generate a response to the prompt. When the large-scale language model 11 is given a prompt with the keywords "leaning forward," "looking at hands," and "not swinging arms," ​​it outputs as a response that there is a possibility of "walking while using a smartphone" and the behavioral characteristics of a pedestrian who is walking while using a smartphone (see Figure 6). The characteristics of a pedestrian who is walking while using a smartphone include "walking slowly" and "not looking around." The response from the large-scale language model 11 is acquired by the information acquisition unit 211 (step S8).

[0085] In step S9, the controller 21 (behavior prediction unit 216) predicts the pedestrian's behavior. Specifically, if the pedestrian whose behavior is being predicted is a specific pedestrian (see Figure 5), the behavior prediction unit 216 predicts the behavior based on the response (output) of the large-scale language model 11 to the prompt input. If the pedestrian whose behavior is being predicted is a general pedestrian (see Figure 5), the behavior prediction unit 216 does not use the response of the large-scale language model 11, but instead uses the predicted trajectory of the trajectory prediction model as the pedestrian's behavior prediction result. In other words, program 221 makes the computer function as a means to perform behavior prediction based on the response of the large-scale language model 11 to the prompt input. Also, program 221 makes the computer function as a means to perform behavior prediction without using the response of the large-scale language model 11 to the prompt input. Once the pedestrian's behavior prediction is complete, the process proceeds to the next step S10.

[0086] Furthermore, when predicting behavior based on the responses of the large-scale language model 11, as described above, the pedestrian characteristic information obtained from the responses of the large-scale language model 11 is used to correct the predicted trajectory obtained by the trajectory prediction model, and the corrected predicted trajectory is used as the pedestrian behavior prediction result. For example, if a response like the one shown in Figure 6 is obtained, "slow walking" is acquired as the pedestrian characteristic information. In accordance with the characteristic information of "slow walking," the predicted trajectory obtained by the trajectory prediction model is corrected so that the distance traveled by the pedestrian per unit time is shortened.

[0087] In step S10, the controller 21 (driving support unit 217) performs processing to support driving based on the behavior prediction results obtained in step S9. As described above, depending on the processing performed by the driving support unit 217, voice notifications from the speaker 2b and screen displays from the display device 2c are appropriately performed for purposes such as collision prevention.

[0088] <4. Variation> [4-1. First variation] In the embodiments described above, the system simply generates a prompt by obtaining a linguistic expression that verbalizes the estimated walking pattern of the pedestrian. However, when generating this prompt, it is also possible to obtain further detailed information about the estimated walking pattern (behavioral pattern) and generate a prompt using an additional linguistic expression that verbalizes this detailed information. In such a configuration, since the amount of information about the pedestrian's walking pattern input to the large-scale linguistic model 11 increases, it is expected that the information about the pedestrian obtained from the large-scale linguistic model 11 (information representing behavioral characteristics, etc.) will become more accurate.

[0089] Detailed information regarding a pedestrian's walking style may, for example, be an expression that describes the degree of their walking style (such as the degree of leaning forward). The degree of the walking style may be provided by, for example, a comparison with other pedestrians present around the pedestrian whose behavior was estimated (a specific pedestrian), or by a comparison with the behavior of pedestrians in general. For example, if the pedestrian's walking style is described as "leaning forward," the detailed information may be information that describes the degree of "leaning forward" compared to surrounding pedestrians or pedestrians in general.

[0090] Figure 7 shows the additional language conversion table 223, which converts detailed information about walking patterns into additional language expressions. The information items in the additional language conversion table 223 include indicators, ranges, and language expressions.

[0091] The item "Indicator" corresponds to detailed information about walking patterns, and more specifically, it is the standard deviation (σ) of various parameter values ​​(such as forward lean and gaze intensity) that represent the walking patterns of a pedestrian whose walking pattern has been estimated (specific pedestrian). Whenever a standard deviation (σ) is calculated for the various parameter values ​​representing walking patterns (such as forward lean and gaze intensity), the calculated standard deviation (σ) is applied to the item "Indicator".

[0092] The standard score (σ) is calculated using, for example, data from pedestrians whose walking patterns have been estimated (specific pedestrians) and statistical data obtained by collecting data from a large number of pedestrians in advance. The statistical data is stored in memory 22. This configuration assumes the "comparison with pedestrians in general" described above. If the configuration assumes the "comparison with the surroundings" described above, the standard score (σ) should be calculated using data from pedestrians whose walking patterns have been estimated (specific pedestrians) and data from other pedestrians other than the specific pedestrians. In this case, it is necessary to estimate the walking patterns of pedestrians other than the specific pedestrians as well.

[0093] The "Range" field stores a set numerical range for classifying the calculated standard score (σ). In the example shown in Figure 7, the numerical ranges are set as follows: Range 1 (less than 30), Range 2 (30 or greater but less than 40), Range 3 (40 or greater but less than 60), Range 4 (60 or greater but less than 70), and Range 5 (70 or greater).

[0094] The "Linguistic Expression" field stores linguistic expressions appropriate for each group classified under the "Scope" field. The appropriate linguistic expressions for each group are determined in advance by the designer. Note that in Figure 3, the parts where the linguistic expression is "-" indicate that the pedestrian's behavior is estimated to be typical based on the standard score (σ), and therefore no specific linguistic expression is needed.

[0095] For example, suppose the value of the leg lift is small, and the language translation table 222 (see Figure 3) retrieves "Do not lift your feet." In this scenario, suppose the standard score obtained from the leg lift value is "35." In this case, the additional language translation table 223 retrieves "More than a normal pedestrian." As a result, the prompt generation uses "More than a normal pedestrian," which is obtained from the language translation table 222 with "Do not lift your feet" added. In this case, the prompt might be, for example, "What conditions might a pedestrian who is walking with less leg lift than a normal pedestrian be in?" The large-scale language model 11 might respond to this with, for example, "They might be walking on an icy surface, which could cause them to walk slowly or fall."

[0096] For example, suppose the value for the degree of leg lift is high, and the language translation table 222 (see Figure 3) retrieves "lift your leg". In this scenario, suppose the standard score obtained from the value of the degree of leg lift is "75". In this case, the additional language translation table 223 will retrieve "prominently higher than a normal pedestrian". As a result, the prompt will use "lift your leg more than a normal pedestrian", which is obtained from the language translation table 222 with "do not lift your leg" added. In this case, the prompt might be, for example, "What condition might a pedestrian who is walking with their leg lifted more than a normal pedestrian be in?" The large-scale language model 11 might respond to this with, for example, "This pedestrian may not be used to walking on snowy roads. Be careful not to fall."

[0097] [4-2. Second variation] In the embodiments described above, the prompt consisted only of text data. However, the large-scale language model 11 may be a multimodal LLM that can handle data such as images in addition to text data. If the large-scale language model 11 is a multimodal LLM, the prompt generation configuration may include image data that reflects the estimated behavioral patterns. In such a configuration, the amount of information about the pedestrian's walking patterns input to the large-scale language model 11 increases, and it is expected that the information about the pedestrian obtained from the large-scale language model 11 (information representing behavioral characteristics, etc.) will become more accurate.

[0098] Figure 8 schematically illustrates a prompt composed of text data and image data. In the example shown in Figure 8, it is assumed that the linguistic expression "Don't swing your arms" is obtained to describe a walking style. In such a case, as illustrated in Figure 8, the prompt may be generated by including image data of a pedestrian walking without swinging their arms. An example of a response from the large-scale language model 11 in the example shown in Figure 8 would be, "The pedestrian may be walking slowly because they are pulling a large load. Also, they may not be able to check their surroundings sufficiently."

[0099] In the example shown in Figure 8, the image is a still image, but a moving image may be used instead. Also, if there are multiple image data that reflect the acquired linguistic expression (in the example shown in Figure 8, "Don't swing your arms"), the image data used for the prompt may be selected from any of those images. However, if there is an image that particularly reflects the state of the acquired linguistic expression, it is preferable to use the data of such an image for the prompt. For example, if a linguistic expression such as "Lean forward" is obtained from the value of the forward lean, it is preferable to use the image data with a particularly large forward lean value for the prompt.

[0100] [4-3. Third Variation] When generating a prompt, the system may include at least one of the following information: the current location information and the time of the device (information processing device 2a, in-vehicle device 2). With this configuration, when the large-scale language model 11 generates a response to a prompt, it becomes easier to access environmental information (weather information, congestion information, etc.) surrounding the pedestrian whose behavior is to be predicted from an external device. As a result, it is expected that the information about the pedestrian obtained from the large-scale language model 11 (information representing behavioral characteristics, etc.) will be more accurate.

[0101] An example of a prompt that includes current location information and time is: "What condition might a pedestrian be in if they are walking with their feet down and looking down? The current time is 7:00 AM, and the location is Hyogo Ward, Kobe City." Given such a prompt, the large-scale language model 11 is expected to respond with something like, "The pedestrian may be walking on an icy surface. The pedestrian is likely walking slowly and is at high risk of falling."

[0102] Whether or not to include the current location information and current time in the prompt may be configured by the user, or it may be configured so that the information processing device 2a automatically selects it. As an example of the latter case, the information processing device 2a may be configured to include the current location information in the prompt only when it has obtained weather information unfavorable for driving the vehicle 100, such as rain, snow, or icy roads, using weather information obtained via the Internet or the like.

[0103] <5. Things to keep in mind> The various technical features disclosed in the embodiments for carrying out the invention as described herein can be modified in various ways without departing from the spirit of the technical creation. Furthermore, the multiple embodiments and modifications disclosed in the embodiments for carrying out the invention as described herein may be combined to the extent possible. [Explanation of Symbols]

[0104] 1. Server 2...In-vehicle device 2a.. Information Processing Device 3. Camera 11. Large-scale language models 100...vehicles 221...Program SYS... Driver assistance system

Claims

1. An information processing device that generates prompts to be input to a large-scale language model, Acquire image data, The acquired image data detects people within the image, The behavioral patterns of the detected person are estimated, Obtain a linguistic expression that verbalizes the estimated behavioral pattern, An information processing device that generates the prompt using the acquired language expression.

2. The information processing apparatus according to claim 1, wherein, when generating the prompt, it further obtains detailed information regarding the estimated behavioral pattern and generates the prompt using an additional linguistic expression that verbalizes the detailed information.

3. The information processing apparatus according to claim 1, wherein when generating the prompt, the apparatus includes image data that reflects the estimated behavioral pattern when generating the prompt.

4. The information processing apparatus according to claim 1, wherein when generating the prompt, it includes at least one of the current location information and the current time information of the device when generating the prompt.

5. An information processing device that performs behavioral prediction to predict human behavior, Acquire image data, The acquired image data detects people within the image, The behavioral patterns of the detected person are estimated, Obtain a linguistic expression that verbalizes the estimated behavioral pattern, Using the acquired language representation, a prompt is generated to be input to the large-scale language model. An information processing device that performs the behavioral prediction based on the response of the large-scale language model to which the generated prompt has been input.

6. The person in the aforementioned image is a person who meets certain requirements, The information processing device according to claim 5, which determines that a person who has a greater degree of deviation between the predicted trajectory based on the acquired image data and the actual trajectory than a predetermined set value is a person who meets the specific requirements.

7. A program that causes a computer to execute a method for generating prompts to input into a large-scale language model, The aforementioned computer, Acquiring image data and To detect a person in the image in the acquired image data, To estimate the behavioral patterns of the person detected, To obtain a linguistic expression that verbalizes the estimated behavioral pattern, The process involves generating the prompt using the acquired language expression, An information processing program that functions as a means to execute a certain action.

8. A program that causes a computer to execute a behavior prediction method that predicts human behavior, The aforementioned computer, Acquiring image data and To detect a person in the image in the acquired image data, To estimate the behavioral patterns of the person detected, To obtain a linguistic expression that verbalizes the estimated behavioral pattern, Using the acquired language representation, generate prompts to be input to a large-scale language model, The generated prompt is used to input the response of the large-scale language model, and the behavior prediction is performed based on that response. An information processing program that functions as a means to execute a certain action.

9. A driver assistance system that assists in driving a vehicle, A server that performs processing using a large-scale language model, An in-vehicle device that generates prompts to be input to the large-scale language model, and obtains the response of the large-scale language model to which the prompts have been input, A camera that photographs the area around the aforementioned vehicle, Equipped with, The in-vehicle device is Acquire image data from the aforementioned camera, The acquired image data detects people within the image, The behavioral patterns of the detected person are estimated, Obtain a linguistic expression that verbalizes the estimated behavioral pattern, The prompt is generated using the acquired language expression. Based on the response of the large-scale language model to the generated prompt, the behavior of the person is predicted. A driver assistance system that provides driving support based on predictions of the aforementioned actions.

10. A camera to film the surroundings, An in-vehicle device provided to communicate with the aforementioned camera, Equipped with, The in-vehicle device is Acquire image data from the aforementioned camera, The acquired image data detects people within the image, The behavioral patterns of the detected person are estimated, Obtain a linguistic expression that verbalizes the estimated behavioral pattern, Using the acquired language representation, a prompt is generated to be input to the large-scale language model. Based on the response of the large-scale language model to the generated prompt, the behavior of the person is predicted. A vehicle that provides driving assistance based on predictions of the aforementioned actions.