Information processing system and computer readable medium

By using large-scale language models and proxy memories in information processing systems, the problem of inaccurate simulation and prediction of human behavior in existing technologies has been solved, achieving high-precision behavior simulation applicable to fields such as urban planning, congestion simulation, recommendation systems, and sales forecasting.

CN121745141APending Publication Date: 2026-03-27TOYOTA JIDOSHA KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies have not yet been able to effectively improve the accuracy of simulations when simulating and predicting human behavior, especially in areas such as urban planning, congestion simulation, recommendation systems, and sales forecasting.

Method used

An information processing system comprising a control unit and a storage unit is adopted. The behavior is generated using a first large-scale language model, evaluated using a second large-scale language model, and updated using the first large-scale language model. Combined with agent memory and role data, a high-precision simulation of human behavior is achieved.

Benefits of technology

It improves the accuracy of human behavior simulation, ensuring that the generated behavior is consistent with the character data and historical information, adapts to the simulation environment, and enhances the realism and accuracy of the simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745141A_ABST
    Figure CN121745141A_ABST
Patent Text Reader

Abstract

The invention provides an information processing system and a computer readable medium, which improve simulation of human behavior in terms of accuracy and / or computational efficiency. A command for causing an information processing system to generate an agent is stored, the agent having: an agent memory for storing history information relating to the agent; a first large-scale language model conditioned on character data representing a character of the agent; a second large-scale language model; and a control module that executes a process of causing the first large-scale language model to generate one or more behaviors on the basis of the history information; causing the second large-scale language model to output an evaluation for the one or more behaviors generated by the first large-scale language model; the first large-scale language model is made to update one or more behaviors on the basis of the evaluation output by the second large-scale language model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an information processing system, and more particularly to an information processing system for simulating behavior of a person. BACKGROUND

[0002] There is known a technique of predicting behavior of a person based on historical information. For example, in Patent Literature 1, a behavior prediction device that predicts behavior of a person using a prediction model learned based on behavior history information of the person is disclosed.

[0003] Patent Literature 1: Japanese Patent No. 7476984

[0004] Since simulation and prediction of behavior of a person are very useful in fields such as city planning, congestion simulation, recommendation system, sales prediction, and personalized activity, there is a demand for a technique of improving simulation of behavior of a person. SUMMARY

[0005] The information processing system in the present disclosure has a control section and a storage section that stores a command that causes the information processing system to generate an agent when executed by the control section, the agent having an agent storage that stores historical information related to the agent, a first large-scale language model conditioned on role data that characterizes a role of the agent, a second large-scale language model, and a control module that performs processing of causing the first large-scale language model to generate one or more behaviors based on the historical information, causing the second large-scale language model to output an evaluation for the one or more behaviors generated by the first large-scale language model, and causing the first large-scale language model to update the one or more behaviors based on the evaluation output by the second large-scale language model.

[0006] The non-transitory computer-readable medium in the present disclosure stores a command that causes the information processing system to generate an agent when executed by the information processing system, the agent having an agent storage that stores historical information related to the agent, a first large-scale language model conditioned on role data that characterizes a role of the agent, a second large-scale language model, and a control module that performs processing of causing the first large-scale language model to generate one or more behaviors based on the historical information, causing the second large-scale language model to output an evaluation for the one or more behaviors generated by the first large-scale language model, and causing the first large-scale language model to update the one or more behaviors based on the evaluation output by the second large-scale language model.

[0007] The information processing system of the present disclosure can improve the simulation accuracy of behavior of a person. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 is a diagram showing a configuration example of an information processing system.

[0009] Figure 2 is a diagram showing one example of an agent.

[0010] Figure 3 is a flowchart showing one example of a process for simulating a person's behavior. DETAILED DESCRIPTION

[0011] Figure 1 is a diagram showing a configuration example of an information processing system 10 in one embodiment. The information processing system 10 includes a control section 12 and a storage section 14. The control section 12 includes one or more processors, and the storage section 14 includes one or more storage units. For ease of understanding, the information processing system 10 illustrated in Figure 1 is simplified, but can include one or more network interfaces, one or more communication interfaces, one or more input interfaces, and / or various other constituent elements equivalent thereto. For example, the storage section 14 stores instruction information 16 including a command that, if executed by the control section 12, installs one or more functions of the information processing system 10. The instruction information 16 can be supplied to the information processing system 10 through a non-transitory computer readable medium, and can also be received via a network (not illustrated) such as the Internet, a mobile communication network, an ad hoc network, a local area network (LAN), a metropolitan area network (MAN), or any combination thereof.

[0012] The control section 12 can include one or more processors, one or more dedicated circuits, or a combination thereof. As for the processor, there are included a general-purpose processor (e.g., a CPU (Central Processing Unit)), a dedicated processor optimized for a specific purpose (e.g., a GPU (Graphic Processing Unit)), or any combination thereof. As examples of the dedicated circuit, there are included an FPGA (Field Programable Gate Array), an ASIC (Application Specific Integrated Circuit). The control section 12 executes information processing for controlling actions performed by the information processing system 10.

[0013] The storage unit 14 may include, for example, one or more semiconductor memories, one or more magnetic memories, one or more optical memories, or any combination thereof. The storage unit 14 can function as a main storage device, an auxiliary storage device, a cache memory, or any combination thereof. Examples of semiconductor memories include RAM (Random Access Memory) and ROM (Read-Only Memory). Examples of RAM include SRAM (Static RAM) and DRAM (Dynamic RAM). Examples of ROM include EEPROM (Electrically Erasable Programmable ROM). The storage unit 14 stores commands and information used in operations performed by the information processing system 10.

[0014] The storage unit 14 stores the instruction information 16, which is then executed by the control unit 12 to install one or more functions of the information processing system 10. The instruction information 16 is sometimes provided to the information processing system 10 via a non-transitory computer-readable medium such as an optical disc or an SSD (Solid State Drive). Alternatively, the instruction information 16 may also be received via a network (not shown) consisting of the Internet, mobile communication networks, ad hoc networks, local area networks (LANs), metropolitan area networks (MANs), or any combination thereof.

[0015] Information processing system 10 uses artificial intelligence to simulate or predict human behavior. More specifically, information processing system 10 generates a software agent (hereinafter referred to as "agent") that simulates the behavior of one or more people based on role data representing the role desired by the agent. Furthermore, information processing system 10 can be configured to simulate or predict human behavior based on an agent memory that stores historical information representing one or more behaviors performed by the agent in the past.

[0016] Figure 2 It means through Figure 1 A simplified diagram of the agent 100 instantiated from the information processing system 10. For ease of explanation, Figure 2 The agent 100 is illustrated by representing independent functional units in blocks according to their functions, where each independent functional unit is installed by the agent 100 via software. Furthermore, Figure 2 The functions shown can be appropriately combined and / or divided.

[0017] The agent 100 includes a control module 102, an agent memory 104, a first large language model (LLM) 106, a second LLM 108, character data 110, one or more lower control modules 112, and a narrative module 114.

[0018] The agent 100 engages in a dialogue with a simulated environment 200. The simulated environment 200 can be installed by the information processing system 10 or by another system external to the information processing system 10. The simulated environment 200 is a virtual environment in which the agent 100 interacts in order to perform various behaviors or tasks. For example, the simulated environment 200 can also be a simulation of a street or city in which the agent 100 resides. In some implementations, multiple agents interact with the simulated environment 200, whereby the multiple agents are able to interact with each other. The interaction with the simulated environment 200 can be implemented via an application programming interface (API) provided by the simulated environment 200.

[0019] The control module 102 is configured to provide overall control of the agent 100. Specifically, the control module 102 is configured to control or orchestrate operations performed by the first LLM 106, the second LLM 108, the one or more lower control modules 112, and the narrative module 114 in order to generate one or more behaviors. The control module 102 can perform control by instructing the first LLM 106, the second LLM 108, the one or more lower control modules 112, and the narrative module 114 individually in a control sequence, whereby the components of the agent 100 are instructed in a manner that performs a particular task. Such commands can be provided according to a particular syntax, API, and / or natural language form. For example, the control module 102 can control the first LLM 106 and the second LLM 108 by issuing one or more natural language prompts.

[0020] The character data 110 specifies or characterizes a person simulated by the agent 100. That is, the character data 110 encapsulates the characteristics of the person simulated by the agent 100, which the agent 100 can simulate in a realistic manner. The character data 110 can include one or more parameters characterizing a character of the agent 100 and / or one or more natural language narratives related to the character of the agent 100. For example, the character data 110 can include one or more parameters characterizing a personality, physical attributes, interests, and / or background of the person simulated by the agent 100. Here, the personality can be specified in one or more parameters or scores corresponding to one or more personality attributes. The personality can also be specified in an alternative or additional manner according to a natural language narrative perspective.

[0021] The character data 110 is sometimes generated based on historical information obtained from a human population. For example, such historical information can also be acquired via a system that monitors the activities of a human population. In other examples, historical information can also be obtained from historical information generated by the agent 100 itself. Based on this historical information, an induction of personality attributes is generated in natural language form. Using this induction, the LLM is used to generate a candidate character that is consistent with the historical information. For example, the LLM is instructed in a manner that generates multiple candidate characters and provides a score indicating the match with the historical information for each candidate character. Depending on the situation, a set of selectable characteristics (a set of different occupations, etc.) can also be provided to the LLM in order to improve the diversity of the candidate characters. If an appropriate character is selected, the selected character is stored in the character data 110.

[0022] The agent memory 104 stores historical information of the agent 100. The historical information provides a record of past actions and experiences of the agent 100. The agent memory 104 is supplemented in real time as the agent 100 interacts with the simulated environment 200. The agent memory 104 can also be initialized based on historical information generated by the agent 100 in previous simulations in an alternative or additional manner. The agent memory 104 functions as a long-term record of the behavior, observations, feelings, and thoughts of the agent 100, and thus characterizes the internal state of the agent 100.

[0023] The first LLM 106 is configured to function as a high-level planner LLM that generates a daily plan that includes one or more behaviors to be performed by the agent 100. Specifically, the first LLM 106 generates a daily plan that is consistent with a character that is associated with the agent 100, given the character data 110. In order to ensure that the daily plan is consistent with the history of the agent 100, the first LLM 106 can also refer to the agent memory 104. Typically, the first LLM 106 can generate a daily plan in units of one day (i.e., once per day). The daily plan can specify one or more behaviors to be performed in one-hour time slots, along with the location at which each behavior is to be performed, and subtasks that are associated with each behavior.

[0024] The first LLM 106 outputs the daily plan to the second LLM 108. The second LLM 108 is configured to function as a critic LLM that evaluates the daily plan generated by the first LLM 106. More specifically, the second LLM 108 is configured to determine whether the daily plan generated by the first LLM 106 is realistic and consistent with the character defined by the character data 110 and / or the history information stored in the agent memory 104. In addition, the second LLM 108 determines whether the daily plan is consistent with one or more external contextual information such as the simulation environment 200, cultural norms, and / or habits of people. The second LLM 108 outputs the evaluation of the daily plan to the first LLM 106. In the evaluation, a score for the daily plan and / or feedback in natural language related to the daily plan are included.

[0025] If the feedback is received, the first LLM 106 updates or changes the daily plan as necessary. By improving the daily plan based on the feedback from the second LLM 108, the first LLM 106 generates a realistic daily plan that is consistent with the internal state of the agent 100 represented by the history information stored in the agent memory 104, given the character data 110. In this way, the agent 100 is able to generate realistic human behavior.

[0026] The first LLM 106 and the second LLM 108 can be installed using a pre-trained LLM such as GPT-4 (Generative Pre-trained Transformer 4) made by OpenAI, Inc. of San Francisco, California, USA, can be installed by fine-tuning the pre-trained LLM, or can be installed by training the LLM specifically for use as the first LLM 106 and / or the second LLM 108.

[0027] The daily plan is generated, reviewed, and updated as necessary, after which the daily plan is executed in the simulated environment 200. Execution of the daily plan is delegated to one or more lower-level control modules 112. The one or more lower-level control modules 112 are configured to perform one or more low-level tasks required to execute the daily plan. For example, the one or more lower-level control modules 112 include a controller specialized to predict the optimal method of moving between two locations within the simulated environment 200. The one or more lower-level control modules 112 can include one or more LLMs trained to perform a specific low-level task, and / or one or more models based on behavior trees, reinforcement learning (RL), neural networks, etc. Generally, the one or more lower-level control modules 112 do not consider the role of the agent 100, but rather execute low-level tasks based on the defined plan. Thus, by delegating low-level tasks required to execute the daily plan to the one or more lower-level control modules 112, the agent 100 is able to simulate human behavior with improved computational efficiency.

[0028] Execution of the daily plan in the simulated environment 200 can include several steps. Initially, the agent 100 can make one or more observations of the simulated environment 200. These observations are sometimes made by the one or more lower-level control modules 112 that control the interaction between the agent 100 and the simulated environment 200. In some embodiments, these observations can also be visual observations translated into natural language by the narration submodule 114. In such embodiments, the narration submodule 114 can be implemented by a visual language model (VLM) configured to augment visual observations made by the one or more lower-level control modules 112 with commentary provided to the first LLM 106 and / or the second LLM 108 in the form of natural language text.

[0029] Upon receiving the observation results from the lower-level control modules 112 and / or the narration submodule 114, the first LLM 106 determines whether a change to one or more behaviors that form the daily plan is required. For example, based on the observation results, the first LLM 106 infers the current mood of the agent 100, the physical state of the agent 100 (e.g., hunger, fatigue), and determines whether the daily plan needs to be revised. For example, if the first LLM 106 determines that the agent 100 is hungry and there is no meal scheduled in the daily plan for several hours, the first LLM 106 determines to go to a restaurant within the simulated environment 200 and revises the daily plan accordingly. In another example, the first LLM 106 determines that the daily plan needs to be revised if one or more observation results indicate that the current location of the agent 100 within the simulated environment 200 is raining and the daily plan includes one or more outdoor activities.

[0030] In addition, the one or more observations of the simulation environment 200 can be used to update the history information stored in the agent memory 104 so that the history information correctly reflects the experiences of the agent 100 in the simulation environment 200. In this way, the experiences of the agent 100 in the simulation environment 200 can be reflected in the generation of future daily plans by the first LLM 106, thereby enabling a more correct simulation of human behavior.

[0031] In certain embodiments, the agent 100 can provide interfaces (not shown), such as APIs, that can initiate questioning of the first LLM 106 via one or more natural language prompts. For example, a human operator can use the interfaces to ask the agent 100 why a particular behavior was performed or is being performed in the simulation environment 200. In this way, the human operator can gain insights into the role of the agent 100. Such insights can be beneficial for urban planning related to the street or city being simulated by the simulation environment 200.

[0032] Figure 3 is a flowchart representing a sequence 300 for simulating human activities in one embodiment. For the sake of clarity, in the following description, it is assumed that the sequence 300 is executed by the control module 102. However, the sequence 300 can also be executed by any constituent element of the agent 100 or a combination of any constituent elements of the agent 100. Furthermore, here, the sequence 300 is described with reference to the simulation environment 200 described above. Figure 2 One example of the control sequence described above is explained.

[0033] In 302, the control module 102 causes the first LLM 106 to generate one or more behaviors based on the history information stored in the agent memory 104. As described above, the one or more behaviors can constitute a daily plan. Here, since the first LLM 106 performs condition setting based on the role data 110, the one or more behaviors reflect the role of the agent 100 and past experiences.

[0034] In 304, the control module 102 causes the second LLM 108 to output an evaluation for the one or more behaviors generated by the first LLM 106. As described above, the evaluation can include a score for the one or more behaviors and / or feedback related to the one or more behaviors in natural language form. The evaluation indicates whether the one or more behaviors are realistic or consistent with the role defined by the role data 110 and / or the history information stored in the agent memory 104.

[0035] In 306, the control module 102 causes the first LLM 106 to update the one or more behaviors based on the evaluation output by the second LLM 108. In this way, the one or more behaviors are realistic with respect to the agent defined by the agent data 110 and are guaranteed to be consistent with the internal state of the agent 100 as characterized by the history information stored in the agent memory 104.

[0036] In 308, the control module 102 causes the one or more behaviors to be executed in the simulated environment 200. As described above, execution of the daily plan can be achieved by the one or more lower control modules 112.

[0037] In 310, the control module 102 causes the first LLM 106 to update the one or more behaviors based on one or more observations of the simulated environment 200. As described above, these observations can be made by the one or more lower control modules 112 that control the interaction between the agent 100 and the simulated environment 200. In this way, the experience of the agent 100 in the simulated environment 200 can be used to improve the one or more behaviors generated by the first LLM 106.

[0038] Embodiments of the present disclosure have been described with reference to the drawings and embodiments, but various modifications and alterations can be made by those skilled in the art based on the description. Therefore, such modifications and alterations are included in the scope of the present disclosure.

Claims

1. An information processing system comprising a control unit and a storage unit, wherein, The storage unit stores the commands that cause the information processing system to generate agents when executed by the control unit. The agent has: A proxy memory that stores historical information related to the proxy; The first large-scale language model is conditioned on role data that represents the role of the agent; The second large-scale language model; and Control module, The control module performs the following processing: The first large-scale language model generates more than one behavior based on the historical information; The second large-scale language model outputs an evaluation for one or more behaviors generated by the first large-scale language model; as well as The first large-scale language model updates more than one behavior based on the evaluation output by the second large-scale language model.

2. The information processing system according to claim 1, wherein, The control module enables more than one behavior to be executed in the simulation environment. The first large-scale language model updates historical information based on more than one observation obtained from the simulation environment.

3. The information processing system according to claim 1, wherein, The agent has a description submodule, which is configured to transform one or more observations obtained from the simulation environment into a natural language description of one or more observations. The control module enables the first large-scale language model to update one or more behaviors of the first large-scale language model based on more than one observed natural language description.

4. The information processing system according to claim 2, wherein, The agent has one or more subordinate control modules configured to perform more than one action in the simulation environment.

5. A non-transitory computer-readable medium for storing commands that, when executed by an information processing system, cause the information processing system to generate an agent, wherein, The agent has: A proxy memory that stores historical information related to the proxy; The first large-scale language model is conditioned on role data that represents the role of the agent; The second large-scale language model; and Control module, The control module performs the following processing: The first large-scale language model generates more than one behavior based on the historical information; The second large-scale language model outputs an evaluation for one or more behaviors generated by the first large-scale language model; as well as The first large-scale language model updates more than one behavior based on the evaluation output by the second large-scale language model.