Large-scale intelligent agent complex interaction behavior simulation system
By constructing a large-scale intelligent agent complex interaction behavior simulation system, and using a large language model to drive intelligent agents to autonomously generate credible social interaction behaviors in a virtual environment, the problems of information omission and redundant storage in existing technologies are solved, and efficient and reliable complex interaction and decision-making efficiency are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-07
AI Technical Summary
Existing generative agents suffer from information omissions and redundant information storage in virtual environments, resulting in low decision-making efficiency and difficulty in achieving complex interactions that do not conform to social logic.
A large-scale intelligent agent complex interaction behavior simulation system is constructed, including a virtual environment construction module, an intelligent agent architecture module, and a user platform module. It uses a large language model to drive intelligent agents to autonomously generate credible social interaction behaviors in highly semantic historical scenarios. It achieves environmental perception, experience storage, and behavior planning through the collaboration of intelligent agent profiling unit and multiple cognitive units.
It improves the decision-making efficiency of intelligent agents in environmental perception, realizes efficient, reliable and socially logical complex interactions, and enhances user experience and system ease of operation.
Smart Images

Figure CN121809516A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the field of computer simulation technology, and in particular to a simulation system for complex interactive behaviors of large-scale intelligent agents. Background Technology
[0002] With rapid social development and continuous technological progress, the methods of disseminating traditional culture are undergoing profound changes. However, current dissemination practices still face many limitations. First, the forms of dissemination are relatively singular and lack interactivity. The "2023 Report on the Dissemination Index of Excellent Traditional Chinese Culture" shows that 78% of respondents believe that current cultural dissemination mainly relies on static displays (such as museum exhibitions and promotional videos), lacking immersive experiences and user participation. Statistics from the China Internet Network Information Center (CNNIC) indicate that the average daily usage time of users on traditional culture-related apps is only 15 minutes, significantly lower than the 50 minutes for entertainment apps, reflecting the difficulty of traditional dissemination models in effectively maintaining users' long-term attention.
[0003] In recent years, with the development of artificial intelligence technology, research has shown that LLMs can not only act as knowledge-based agents in solving complex tasks, but also combine with intelligent agents to build generative intelligent agent architectures, endowing them with perception, decision-making, and reflection capabilities, thereby achieving autonomous behavior and complex interactions. These generative intelligent agents can not only simulate real social interactions in virtual spaces, but also, due to their highly realistic behavioral performance, may become an important potential tool for promoting cultural dissemination.
[0004] However, despite the promising applications of multi-agent systems, current research still faces several key challenges. On the one hand, generative agents often miss information during environmental perception, affecting their comprehensive understanding of scene events; on the other hand, their memory storage mechanisms are prone to generating a large amount of redundant information, thereby reducing decision-making efficiency and accuracy. Therefore, how to achieve efficient, reliable, and socially logical interaction between agents in a virtual environment remains a core problem that urgently needs to be solved. Summary of the Invention
[0005] To address the aforementioned technical issues, embodiments of this application propose a large-scale intelligent agent complex interaction behavior simulation system. This system aims to construct a three-layer system integrating a virtual environment, a humanoid intelligent agent architecture, and a user interaction platform. By driving the intelligent agent through a large language model, it autonomously generates credible social interaction behaviors in highly semantic historical scenarios, thereby improving the decision-making efficiency of the intelligent agent in environmental perception and achieving efficient, credible, and socially logical complex interactions.
[0006] To achieve the above objectives, embodiments of this application propose a large-scale intelligent agent complex interaction behavior simulation system, the system comprising: a virtual environment construction module, an intelligent agent architecture module, and a user platform module; The virtual environment construction module is used to build a structured virtual space that fits the target application scenario; the structured virtual space includes functional area divisions and interactive semantic information; The agent architecture module is used to drive generative large language models and includes an agent profiling unit and multiple cognitive units. The agent profiling unit includes basic information dimensions, psychological state dimensions, time awareness dimensions, and social awareness dimensions to provide the agent with cognitive and behavioral capabilities. The cognitive units work together to realize the agent's perception of the environment, experience storage, behavior planning, and semantic interaction among multiple agents. The user platform module provides a visual interface that allows users to configure agent attributes and simulation parameters, and displays the interaction process and results of multiple agents in real time.
[0007] To achieve the above objectives, embodiments of this application also propose a method for simulating complex interactive behaviors of large-scale intelligent agents, the method comprising: Construct a structured virtual space that fits the target application scenario; the structured virtual space includes functional area divisions and interactive semantic information; Generative large language model drives the intelligent agent's profile unit and multiple cognitive units; the intelligent agent profile unit includes basic information dimension, psychological state dimension, time awareness dimension and social awareness dimension to provide the intelligent agent with cognitive and behavioral capabilities; the cognitive units work together to realize the intelligent agent's perception of the environment, experience storage, behavior planning and semantic interaction among multiple intelligent agents; It provides a visual operation interface, allowing users to configure agent attributes and simulation parameters, and display the interaction process and results of multiple agents in real time.
[0008] To achieve the above objectives, embodiments of this application also propose an electronic device, including: a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to execute the instructions such that the electronic device can implement a method for simulating complex interactive behaviors of a large-scale intelligent agent as described above.
[0009] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of a method for simulating complex interactive behaviors of large-scale intelligent agents as described above.
[0010] This application proposes a large-scale intelligent agent complex interaction behavior system, comprising: a virtual environment construction module, an intelligent agent architecture module, and a user platform module; the virtual environment construction module is used to construct a structured virtual space that fits the target application scenario. Since the structured virtual space includes functional area divisions and interactive semantic information, it can provide the intelligent agent with the ability to perceive its surrounding environment; the intelligent agent architecture module is used for generative large language model driving and includes an intelligent agent profiling unit and multiple cognitive units; since each cognitive unit collaboratively realizes the intelligent agent's perception of the environment, experience storage, behavior planning, and semantic interaction among multiple intelligent agents, it can support the complete decision-making process of the intelligent agent from environmental perception to behavior output, thus endowing it with… It more closely resembles human behavior and adaptability; the user platform module provides a visual operation interface, supports users in configuring agent attributes and simulation parameters, and displays the interaction process and results of multiple agents in real time, thus improving the user-friendliness and ease of operation of the system; based on this, the large-scale intelligent agent complex interaction behavior system provided in this application can construct a three-layer system integrating a virtual environment, a humanoid intelligent agent architecture, and a user interaction platform. Through a large language model, it drives intelligent agents to autonomously generate credible social interaction behaviors in highly semantic historical scenarios, thereby improving the decision-making efficiency of intelligent agents in environmental perception and realizing efficient, credible, and socially logical complex interactions. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0012] Figure 1 This is a structural diagram of a large-scale intelligent agent complex interaction behavior simulation system provided in one embodiment of this application; Figure 2 This is a schematic diagram of a scene creation process provided in one embodiment of this application; Figure 3 This is a schematic diagram of a restaurant's first-floor scene and area division provided in one embodiment of this application; Figure 4 This is a schematic diagram of the second floor scene and area division of a restaurant provided in one embodiment of this application; Figure 5 This is a schematic diagram of a map semanticization method provided in one embodiment of this application; Figure 6This is an example diagram of an intelligent agent profile provided in one embodiment of this application; Figure 7 This is a schematic diagram of the structure of a sensing unit provided in one embodiment of this application; Figure 8 This is an example diagram of fine-tuning data provided in one embodiment of this application; Figure 9 This is a schematic diagram of the structure of a memory unit provided in one embodiment of this application; Figure 10 This is a schematic diagram of the structure of a reflection unit provided in one embodiment of this application; Figure 11 This is a decision flowchart provided in one embodiment of this application; Figure 12 This is an action execution flowchart provided in one embodiment of this application; Figure 13 This is a system flowchart provided in one embodiment of this application; Figure 14 This is a flowchart of a user platform module provided in one embodiment of this application; Figure 15 This is a schematic diagram of a system simulation interface provided in one embodiment of this application; Figure 16 This is a schematic diagram of a character basic information viewing and modification interface provided in one embodiment of this application; Figure 17 This is a schematic diagram of an interface for viewing and modifying agent relationships provided in one embodiment of this application; Figure 18 This is a schematic diagram of a simulation parameter modification and viewpoint switching interface provided in one embodiment of this application; Figure 19 This is a schematic diagram of a simulated display interface provided in one embodiment of this application; Figure 20 This is a flowchart of a method for simulating complex interactive behaviors of large-scale intelligent agents provided in another embodiment of this application; Figure 21 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0014] This application proposes a large-scale intelligent agent complex interaction behavior system for Tang Dynasty society. It aims to leverage the powerful planning and reasoning capabilities of a large model to design a reasonable intelligent agent simulation framework, driving the agents to autonomously perceive, make action decisions, and engage in social interaction. This allows the agents to autonomously generate credible behaviors such as interactions within a virtual environment simulating Tang Dynasty society, achieving large-scale information dissemination and intelligent agent interaction within the environment. The following description uses Tang Dynasty society as an example to illustrate this application's embodiments.
[0015] like Figure 1 As shown, Figure 1 This is a structural diagram of a large-scale intelligent agent complex interaction behavior simulation system proposed in one embodiment of this application. The system includes: a virtual environment construction module 110, an intelligent agent architecture module 120, and a user platform module 130.
[0016] The virtual environment construction module 110 is used to construct a structured virtual space that fits the target application scenario. This structured virtual space includes functional area divisions and interactive semantic information.
[0017] Understandably, in order to recreate vivid scenes of Tang Dynasty society, this application uses typical Tang Dynasty social venues, such as restaurants, as a virtual environment, and designs and implements a two-dimensional simulation scene. For example... Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the scene creation process provided in one embodiment of this application. This scene not only includes meticulous spatial layout but also takes into account unique cultural elements and social customs of the Tang Dynasty, making the entire simulated environment more closely resemble historical reality.
[0018] Specifically, the restaurant's interior is divided into multiple functional areas, such as the counter, front hall, stage, main hall, and private rooms. Each area is equipped with corresponding interactive objects, such as dining chairs and notice boards. These objects not only have visual effects but, more importantly, possess semantic-level interactive logic, enabling intelligent agents to make reasonable behavioral choices based on their roles and needs. For example, an intelligent agent can view notice boards with poem themes, participate in performances, or enjoy food. Six intelligent agents are deployed in this environment, each endowed with unique personality traits and backgrounds, including but not limited to roles such as poet, merchant, and performer. These diverse role settings simulate complex interactive behaviors in real historical contexts.
[0019] To ensure that these intelligent agents can clearly perceive their surroundings and have a clear understanding of their environment, this application proposes an innovative map semanticization method. The core of this method lies in converting a visualized two-dimensional map into a structured natural language description, enabling the intelligent agent to accurately understand spatial layout, scene elements, and event context based on textual information. For example, the intelligent agent will know how to move from one room to another, or how to react appropriately upon seeing a specific notice, based on the description. See the following embodiments for details.
[0020] For example, this application uses the Tilemap plugin in Unity to draw the restaurant scene. The scene creation process mainly includes scene drawing, animation production, and UI design. For details, please refer to... Figure 2 And the following text.
[0021] In the scene rendering section, the map assets were first decomposed into a series of appropriately sized tiles to accommodate different levels of detail. Then, using the Tilemap plugin in the Unity engine, four independent but interconnected layers were drawn: foreground, midground, background, and entity layers. This layered design not only enhances the map's visual depth and richness but also facilitates subsequent physical property settings. Finally, by integrating these four layers, a complete and functional virtual map scene was constructed. Based on this, detailed attribute settings were further applied to the map, particularly configuring key physical properties such as colliders.
[0022] In terms of animation production, continuous four-dimensional character assets are used as the basic resources. In the specific implementation process, the character's walking, standing, and other motion assets are first imported into Unity's Animator component, and the corresponding image sequence is played in a loop according to a preset frame rate to generate the dynamic effect of character movement. This application created eight basic animations, including walking forward, walking backward, standing still forward, and standing still backward, and used the BlendTree plugin and C# scripts to control the animation state, realizing the function of automatically switching the corresponding animation based on the character's current state.
[0023] In the design and implementation of the user interface (UI), this application adopts a material segmentation method similar to that used in map drawing. Then, using UI components provided by Unity, such as TextMeshPro, Button, and Canvas, the relevant interfaces of the system were drawn. Furthermore, this application developed a series of scripts to control the display and hiding of UI elements, user input responses, and data transmission functions, ensuring that users can efficiently operate the system and obtain the information they need through an intuitive and easy-to-use interface.
[0024] like Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram of a restaurant's first-floor scene and area division provided in one embodiment of this application. Figure 4 This is a schematic diagram illustrating the second-floor scene and area division of a restaurant as provided in one embodiment of this application. The restaurant scene consists of two sub-scenes: the first floor and the second floor. Internally, it is divided into multiple areas and equipped with various interactive objects. The first floor encompasses functional areas such as a counter, lobby, stage, main hall, and private rooms. This area creates a lively and vibrant atmosphere, equipped with interactive objects such as dining chairs and notice boards to support dining and performances. The second floor is designed as a space integrating dining and accommodation, including guest rooms, balconies, and dining areas. This area features multiple guest rooms and open-plan seating with views, equipped with interactive objects such as beds and dining chairs to ensure the needs of intelligent agents in different situations are met. The objects in the restaurant scene not only have visual representations but also possess semantic-level interactive logic, enabling intelligent agents to make reasonable behavioral choices based on their roles and needs. For example, an intelligent agent can view notices with poem themes and dine in this area.
[0025] However, unlike humans who can quickly perceive their surroundings, identify their location, and form an intuitive understanding of space through their visual system, intelligent agents cannot directly rely on visual information to acquire scene content. Therefore, this application requires that intelligent agents actively collect information related to their position and state in a virtual environment, thereby constructing a basic understanding of the surrounding environment. This serves as the foundation for the agent's behavioral decisions and is a key prerequisite for achieving contextualized and rationalized actions. To this end, this paper proposes a map semanticization method that enables map data to be understood by the intelligent agent and thereby form an understanding of the environment, as detailed in the following embodiments.
[0026] In one possible embodiment, the virtual environment construction module is used to convert a visualized space into a structured natural language description using a map semanticization method. The structured natural language description includes spatial attribution, functional attributes, and dynamic state information, while the map semanticization method includes map digitization and map initialization.
[0027] like Figure 5 As shown, Figure 5 This is a schematic diagram of a map semanticization method provided in one embodiment of this application. The method includes two steps: map digitization and map initialization. For example, in the map digitization stage, the restaurant scene is constructed using a tile map format, consisting of several regularly arranged square tiles, each tile occupying a unique two-dimensional coordinate. This forms the spatial framework of the scene.
[0028] In one possible embodiment, map digitization includes: a set of scenarios for the target application scenario. Each scene in the program is divided into functional areas. The system performs semantic recognition on the divided functional areas and interactive objects to obtain a set of regions. , and item collection , And assign a unique identifier code to each of them. , Establish a mapping table from codes to regions or objects. Among them, the mapping table This can be expressed by the following formula (1): (1); Subsequently, by semantic recognition of functional areas and interactive objects, all tile coordinates are traversed. Based on the functional area or interactive object it belongs to, a corresponding encoding value is assigned to it, thereby forming a digital map. For example, a structured virtual space is divided into tiles, and the coordinates of each tile are... By assigning coded values representing functional areas or interactive objects, a digital map is obtained. Among them, digital maps This can be expressed by the following formula (2): (2); Through the above process, the original scene is transformed into a two-dimensional grid composed of digital codes, realizing a structured transformation from a "physical map" to a "digital map".
[0029] However, simply digitizing the map is insufficient to support an agent's semantic understanding of the environment. An agent cannot obtain crucial information such as the functional attributes and interaction status of a current tile solely from digital encoding. Therefore, map initialization is also necessary to further integrate the encoding mapping table with the digital map, injecting structured semantic information into each tile.
[0030] In one possible embodiment, the map initialization includes: traversing all coordinate points. From digital maps Extract the encoding information at the corresponding position. Through the mapping table Analyze each coordinate point Region type and interactive objects; based on coordinate points The region type and interactive objects are used to construct a structured dictionary data containing multi-level semantic information. and obtain tuples , and event information collection , Among them, dictionary data The fields include world Scene Functional areas Interactive objects and dynamic event states And expressed by the following formula (3): (3).
[0031] Ultimately, each tile coordinate corresponds to a data structure containing rich semantic information, forming a complete mapping relationship from coordinates to semantics, for example: ={ "world": "Tang Dynasty Town", "sector": "first floor of the restaurant" "area":"counter", "game_object":"ledger", "event":"Idle"} The fields “world,” “sector,” “area,” and “game_object” describe the region affiliation and functional attributes of the coordinates at a granular level from macro to micro; the “event” field records the current state information of the object. In this way, the map not only possesses spatial structural information but also integrates rich semantic features, providing a data foundation for intelligent agents to achieve efficient perception, reasoning, and behavior generation in simulated environments.
[0032] The intelligent agent architecture module 120 is used for driving generative large language models and includes an intelligent agent profiling unit and multiple cognitive units. The intelligent agent profiling unit includes basic information dimensions, psychological state dimensions, time awareness dimensions, and social awareness dimensions to provide the intelligent agent with cognitive and behavioral capabilities. The cognitive units work together to realize the intelligent agent's perception of the environment, experience storage, behavior planning, and semantic interaction among multiple intelligent agents.
[0033] To achieve a highly realistic simulation of individual behavior in Tang Dynasty society, this paper constructs a well-structured and fully functional humanoid intelligent agent architecture. This architecture integrates concepts from cognitive science and artificial intelligence, driven by a large language model, and focuses on the agent's personalized modeling and environmental interaction capabilities. The humanoid intelligent agent comprises six core modules: internal state (i.e., the agent's profile module), perception unit, memory unit, reflection unit, planning unit, and dialogue unit. Through the coordinated operation of these six modules, the agent can achieve a complete closed loop from environmental perception to behavior generation, possessing key capabilities such as role recognition, dynamic memory management, experiential inductive reasoning, and goal-oriented decision-making. This provides a foundation for constructing multi-agent systems with social relationship networks.
[0034] As an example, such as Figure 6 As shown, Figure 6 This is an example diagram of an intelligent agent profile provided in one embodiment of this application. In the intelligent agent profile module, each intelligent agent is given a comprehensive profile that includes multiple dimensions such as personality traits, social background, and professional identity. This profile system provides the intelligent agent with clear identity recognition and behavioral boundaries, enabling it to generate situation-appropriate behavioral responses based on its own role, thereby significantly improving the realism and consistency of the simulation.
[0035] For example, this application can construct an agent's consciousness model from four dimensions: basic information, psychological state, time awareness, and social awareness, aiming to comprehensively characterize the agent's cognitive structure and behavioral logic. The basic information dimension describes the agent's identity attributes and external characteristics, including name, age, gender, identity positioning, and appearance description. This information constitutes the basic framework of the agent's profile and is the primary basis for identifying and distinguishing different agents. The psychological state dimension focuses on the agent's emotions and personality traits, specifically covering current mood and personality. This dimension reflects the dynamic changes in the agent's internal emotions and personality traits, providing psychological support for behavior prediction and response mechanisms. Among them, personality traits can be modeled using the MBTI personality assessment system to enhance the ability to characterize individual differences. The time awareness dimension reflects the agent's cognition and grasp of its own state on the timeline, including three tense levels: past, present, and future: (a) Past: The agent's description and summary of its recent state. (b) Present: Capturing the agent's real-time state, including ongoing events, spatial location, current time, and action state (e.g., "eating," "conversing"); (c) Future: Involving goal setting and action plans, reflecting the agent's expectations and planning capabilities for future development. The social awareness dimension is used to model the agent's social relationship network and key factors influencing social behavior, such as social desires and social status. Interpersonal relationships are organized in a graph structure: (a) Nodes represent the agent's impressions of others, including basic information, descriptions, and attitudes; (b) Edges represent the type and strength of the relationship between the two, forming a dynamically updated social cognitive graph. This graph not only helps to understand the agent's role and influence in a group but also provides structured support for simulating complex social behaviors.
[0036] In one possible embodiment, the multiple cognitive units include a perception unit, which is used for an attention-score-based event filtering mechanism to collect, identify, and preliminarily process events in the environment. The event filtering mechanism includes an event perception phase and an event evaluation phase.
[0037] For example, such as Figure 7 As shown, Figure 7 This is a schematic diagram of a sensing unit provided in one embodiment of this application. As a core component for interaction between an intelligent agent and its external environment, the sensing unit performs the functions of information acquisition and preliminary processing, similar to the human sensory system (such as vision and hearing). This module is responsible for identifying and analyzing the state changes and related events of other intelligent agents and interactive objects in the surrounding environment, and for conducting preliminary assessments of these events to determine whether they warrant further attention.
[0038] This application proposes an event screening mechanism based on attention scores. This mechanism comprehensively considers three dimensions to score events. Overall, the event screening mechanism consists of two main parts: event perception and event evaluation.
[0039] In the event perception phase: the current agent Acquire all event information in the current scene. This event information includes relevant event information from interactive objects and other intelligent agents present; the event information includes the actor responsible for the event, the action performed by the actor, the object affected by the action, the location of the event, and the timestamp of the event. For example, in the event perception phase, the current intelligent agent... It will acquire relevant event information from all interactive objects in the current scene and other intelligent agents present. For any event Its structure can be represented in the form of a quintuple, including the subject S, predicate P, object O, location L, and time T of the event, specifically as follows: ; Based on the information source type of the event information, a description of the event information is generated; where the information source type includes events of interactive objects and events of other intelligent agents.
[0040] For example, after obtaining the five-tuple of an event, the system automatically generates a natural language description to facilitate event filtering. When describing events of interactive objects, this application only generates the corresponding description information when the object is in an idle state, for example: Event Description = At 12:00 noon on April 20, 2025, chair number 1 in the lobby on the first floor of the restaurant was unused.
[0041] In describing the perceived intelligent agent When discussing the event, this article will begin... The description of an event consists of two parts: basic information such as age, gender, appearance, and behavior, which can be obtained through direct observation, and relationship information, which can be obtained by querying social networks.
[0042]
[0043] Finally, based on the basic information obtained through observation and the relational information obtained through querying, this application will generate... Event descriptions, for example: Event Description: At 12:00 PM on April 20, 2025, in the lobby of a restaurant, a tall and handsome middle-aged man is talking to someone. The man's name is {...}, age is {...}, and identity is {...}. Your attitude towards him is {...}, and your relationship is {...}.
[0044] It is worth noting that in this application, the content of the dialogue between agents is not limited to the perception of the two parties involved; other agents within a certain range of the scene can also receive the specific content of the conversation, even if they are not directly involved in the exchange. In particular, when an agent performs a broadcast-like action such as shouting or performing, the system will globally broadcast the information carried by that action within the current scene, ensuring that all agents can perceive and respond to the relevant content. This mechanism aims to simulate the information dissemination effect of sound propagation in real-world environments, enhancing the system's expressive capabilities in fine-grained spaces.
[0045] In the event evaluation phase: based on the attention scoring model, information source type, and event description information, all event information is prioritized, and the first event information with a higher preset priority is selected. The attention scoring model is used to evaluate the multidimensional attributes of events, and it is constructed using knowledge distillation and efficient parameter fine-tuning techniques, and trained using event score data.
[0046] In one possible embodiment, the multidimensional attributes include perceptual saliency, task relevance, and emotional intensity; the prioritization of all event information based on the attention scoring model, information source type, and event information description information includes: if the information source type is an event of an interactive object, then the attention scoring model is controlled to evaluate the event information description information using task relevance to obtain the priority of the event information; if the information source type is an event of another intelligent agent, then the attention scoring model is controlled to evaluate the event information description information using perceptual saliency, task relevance, and emotional intensity to obtain the priority of the event information.
[0047] Understandably, since the amount of information acquired through perception is usually large, and not all information is equally important or worthy of further processing, it is necessary to introduce an attention scoring model to evaluate and prioritize perceived events, selecting those with the highest processing value. For events involving interactive objects, this application only considers task relevance. For events involving people, this paper will evaluate events from three aspects: Perceptual salience: refers to the degree to which certain information in the environment is more likely to attract attention due to its obvious features or prominent relationships. Information with strong salience is given priority for the agent to pay attention to and process. Task relevance: indicates the degree of association between the current event and the agent's goals and decisions. Events with strong relevance have a greater impact on behavioral choices and should be given priority. Emotional intensity: refers to the intensity of the emotions (such as fear, surprise, anger) evoked by the event. The stronger the emotion, the easier it is to attract attention and influence memory and decision-making.
[0048] While large language models can comprehensively evaluate dimensions such as perceptual saliency, task relevance, and emotional intensity during the construction of attention scoring models, their computational cost is high, and limitations in inference speed and generation stability often lead to convergent scoring results, making it difficult to accurately reflect subtle differences between events. To address these issues, this application combines knowledge distillation and efficient parameter fine-tuning techniques to construct a lightweight attention scoring model for efficient and accurate scoring of the multidimensional attributes of events.
[0049] First, leveraging the capabilities of a large language model, this paper generates over 2000 structured data samples. Each sample contains basic information about the agent, an event description, the agent's impression of the event's protagonist, and initial ratings for the event across three dimensions: perceptual saliency, task relevance, and emotional intensity. To further improve the accuracy and consistency of the ratings, the model-generated scores are manually reviewed and corrected, ultimately yielding the event score data.
[0050] Based on this, such as Figure 8 As shown, Figure 8 This is an example diagram of fine-tuning data provided in one embodiment of this application. This application utilizes LoRA technology to fine-tune the Qwen-0.6B-Insruct to obtain an attention score model. During model training, this application employs the LLAMAFACTORY framework, using event score data to fine-tune an existing lightweight model, successfully constructing an attention score model, which is then integrated into the system's perception unit. This model is used to quantitatively evaluate input events across multiple cognitive dimensions and, based on this, select events with higher priority as the basis for subsequent processing.
[0051] In one possible embodiment, the multiple cognitive units further include memory units connected to the perceptual units. The memory units include different levels of memory, including perceptual memory, short-term memory, working memory, and long-term memory.
[0052] For example, the memory unit records important events experienced by the agent and their behavioral feedback in natural language, constructing a structured experience knowledge base. Combined with a retrieval enhancement mechanism based on relevance and importance, this module can efficiently extract the memory fragments most relevant to the current state in complex situations, assisting the agent in making more context-coherent decisions.
[0053] Understandably, research on memory mechanisms has always been a hot topic in cognitive science and artificial intelligence. The Atkinson-Schifflin model provides a basic framework for constructing memory units for intelligent agents, emphasizing the hierarchical structure of memory, mainly comprising three subsystems: perceptual memory, short-term memory, and long-term memory. In the field of artificial intelligence, simulating human memory mechanisms is crucial for creating human-like intelligent agents capable of complex interactions and generating reliable behavior over the long term. Related technologies store the memories of human-like intelligent agents in the form of a tree, accessing tree nodes to retrieve memories relevant to the current event, and reflecting and summarizing through the memories in each child node, using the generated more abstract memories as the parent nodes of these nodes. However, such memory units suffer from problems such as incomplete retrieval, low accuracy, and storage redundancy.
[0054] Therefore, as Figure 9 As shown, Figure 9 This is a schematic diagram of a memory unit provided in one embodiment of this application. Based on cognitive science theories such as the Atkinson-Schifflin model, this application designs a hierarchical intelligent agent memory unit, comprising three memory levels: perceptual memory, short-term memory, and long-term memory. It also designs methods for memory transfer and forgetting. Perceptual memory is used to acquire the current task context, obtain first event information from the perceptual unit, and determine the attention width threshold corresponding to the current task context. Simultaneously, an attention bottleneck mechanism is employed to filter out second event information from the first event information that exceeds the attention width threshold corresponding to the current task context. Perceptual memory is located on the far left and is responsible for temporarily storing environmental information and transferring some information to short-term memory. Short-term memory further processes and integrates information and decides whether to transfer it to long-term memory based on activation factors. The cylindrical structure on the right represents long-term memory, used to store persistent knowledge and experience, and related memories can be retrieved and added to working memory using RAG technology.
[0055] In the memory unit designed in this application, perceptual memory is generated by the perceptual unit and represents all event information acquired by the agent during a single perception process. This information includes, but is not limited to: event triples (subject-action-object), event description, time and place of occurrence, and the impression of the relationship between the agent and the event subject. The formal structure of perceptual memory is shown below: {"event": ("Li Zhi", "conversation", "Li Wenbo"), "time": "2025-12-1 12:00" "sp": "Tang Dynasty Town: Restaurant, First Floor: Counter", "des": "At 12:00 on December 1, 2025, a slender and gentle young woman was talking to someone behind the counter on the first floor of the restaurant." "relationship": Your impression of them is: name unknown, age unknown, identity is restaurant owner. Your attitude towards them is friendly, and your relationship is that you have just met. As mentioned earlier, the vast amount of raw information generated by the sensory unit needs to be processed through a filtering mechanism before it can enter the deeper memory system (i.e., the conversion from sensory memory to short-term memory). In this process, this study introduces the Attention Bottleneck Mechanism to simulate the role of selective attention in the human cognitive system. This theory states that sensory information must pass through a limited-capacity selective attention channel to have a chance of entering the short-term memory system.
[0056] Specifically, the agent's attentional resources have a width limit, which can be dynamically adjusted under different task contexts. This "width" is equivalent to a bottleneck in information filtering, determining the upper limit of the number of events allowed to enter short-term memory at the current moment. Only high-priority events with scores above a threshold can pass through this bottleneck and enter the next stage of processing, while other low-priority events are discarded or not processed further.
[0057] Short-term memory is used to send event information from the second event information that is relevant to the decision-making task faced in the current task situation to working memory, so that working memory can support the execution of the decision-making task, while based on activation factors. The third event information is selected from the second event information and transferred to long-term memory so that long-term memory can store and retrieve the third event information; among them, the activating factor Indicates the first The degree of activation of a memory, when the activating factor Above the activation threshold Then the first One memory is transferred to long-term memory; Among them, the activation factor is calculated. Formula (4) is as follows: (4); in, This is the current timestamp. It is memory The total number of times it was perceived. It is the repetitive perception coefficient. It is memory Total number of searches It is the repetitive perception coefficient. That is, the current timestamp with the timestamp of the event Time difference, This represents the attenuation rate.
[0058] In this application, short-term memory is configured as a queue of finite length, storing dynamic attribute information of the character, including but not limited to real-time changing state data such as character coordinates, location, and ongoing events. Furthermore, as an information relay station for the entire memory unit, short-term memory primarily sources data from perceptual memory units and carries richer semantic and contextual information than perceptual memory. These memories have a high probability of being related to the decision-making task currently faced by the agent and may be activated into working memory for further processing. Simultaneously, short-term memory is also the sole input source for long-term memory (LTM). After screening and evaluation, some key memories will be encoded and transferred to the long-term memory system to achieve long-term knowledge retention.
[0059] The following is a typical example of short-term memory, which adds attention scores, activators, and more detailed event descriptions to the original perceptual memory: {"event": ("Li Zhi", "conversation", "Li Wenbo"), "time": "2025-12-1 12:00" "sp": "Tang Dynasty Town: Restaurant, First Floor: Counter", "des": "At 12:00 on December 1, 2025, a slender and gentle young woman was talking to someone behind the counter on the first floor of the restaurant." "relationship": Your impression of them is: name unknown, age unknown, identity is restaurant owner. Your attitude towards them is friendly, and your relationship is that you have just met. "score" : { "s": 0.73, "c" : 0.88, "e": 0.25 } "chat": "Sir, you're joking." Activation Factor: 1.45 } In the above formula (4), the calculation of the activation factor is affected by the following three dimensions: First, the number of perceptions and repeatability: whenever the sensory unit recognizes a new event... Then, the system will search its short-term memory database for similar events. And calculate the semantic cosine similarity. If Greater than a certain threshold If the memory is positive, the existing memory entry is updated and its activation factor is increased; otherwise, it is added to the list as a new memory, and its activation factor is initialized as its attention score. Second, retrieval frequency and usage intensity: whenever a memory is retrieved... i When a memory is retrieved into working memory to respond to an event, its activation factor is increased. Third, the time decay mechanism: memories decay naturally over time, and we use an exponential decay function to simulate this process.
[0060] In one possible embodiment, long-term memory stores third event information by employing vector database technology for storing and retrieving the third event information.
[0061] Understandably, information in long-term memory does not decay or disappear rapidly. This information is relatively persistent, although it may be modified by other information or become temporarily inaccessible.
[0062] In this application, long-term memory stores static attributes such as name, personality, identity and appearance, interpersonal relationship diagram of the agent, and experience gained by the agent after reflection.
[0063] Long-term memory is not forgotten; therefore, this study introduces vector database technology to achieve efficient storage and semantic retrieval of long-term memory content. The core advantage of vector databases lies in their ability to transform unstructured data (such as text) into high-dimensional numerical vectors with semantic features, thereby supporting efficient retrieval based on semantic similarity. Compared to traditional databases, this method demonstrates stronger capabilities in handling complex semantic information retrieval tasks, particularly excelling in scenarios requiring an understanding of potential relationships between data.
[0064] In this implementation, FAISS is chosen as the vector database for long-term memory. FAISS is an efficient similarity search library specifically designed for handling high-dimensional vector search problems, particularly suitable for semantic retrieval tasks in large-scale data scenarios. Its core idea is to accelerate the near nearest neighbor search process by constructing an efficient index structure. This study uses the IVF (Inverted File Index) structure provided by FAISS to efficiently store and retrieve text-type memory data. The IVF index first clusters the entire vector space into k clusters. Each newly added vector is assigned to the nearest cluster center, and similarity search is performed only within that cluster, thus significantly reducing search complexity.
[0065] The set consists of all memory vectors in the vector database. , ,in, Indicates the first Each memory entry A dimensional semantic vector can be divided into several parts using an index structure. A subset; when a new memory entry When long-term memory needs to be incorporated, the semantic embedding vector is extracted using the BERT model and expressed by the following formula (5): ( 5); New memory entries The corresponding semantic embedding vectors are stored in a vector database so that they can be used to query text. corresponding vector Retrieves text matching the query. The corresponding set of memory entries Specifically, it is represented by the following formula (6): (6).
[0066] Understandably, by introducing the FAISS vector database and IVF indexing mechanism, this application achieves efficient storage and semantic retrieval of long-term memory content, providing support for intelligent agents to invoke historical experience in complex tasks.
[0067] In one possible embodiment, the multiple cognitive units further include a reflection unit, which is used to perform structured updates on information in the vector database. The structured updates include dynamic updates of agent attribute states and reflective optimization of memory content. In the dynamic updates of agent attribute states: a retrieval-enhanced generation technique is used to extract all memory entries related to the current object from the vector database, and all memory entries related to the current object are used as context and input into the large language model; the large language model is guided to update and output the agent's internal state through a preset structured prompt word template. Understandably, relying solely on persistent memory storage mechanisms will inevitably lead to information redundancy, contradictions, and inefficiency in long-term memory, thus affecting the quality of subsequent reasoning and decision-making. Therefore, effectively integrating, refining, and optimizing memory is one of the key challenges in building intelligent agents with continuous learning capabilities. Figure 10 As shown, Figure 10 This is a schematic diagram of the structure of a reflection unit proposed in one embodiment of this application. This application proposes a reflection mechanism based on semantic clustering, aiming to simulate the information integration and abstraction generation behavior during memory consolidation. This mechanism is activated when the agent performs a "rest" behavior, performing structured updates to the information in its long-term memory bank, specifically including two aspects: dynamic updating of the agent's attribute states and reflective optimization of the memory content.
[0068] In attribute updates, the system employs retrieval-enhanced generation technology to extract all memory entries related to the currently viewed object from the memory bank and input them as context into the large language model. Subsequently, combined with the character's basic attribute information (such as name, identity, age, personality, etc.), structured cue words are constructed to guide the model in updating interpersonal relationships, plans, and other aspects.
[0069] Here is a sample prompt template for updating interpersonal relationships: Prompt = Your name is {...}, your identity is {...}, your age is {...}, your personality is {...}, and you describe yourself as {...}. You have lived a day and remember some things about {...} as follows: {...}. These things may affect your relationship with him / her. Your original relationship with him / her was as follows: {...} (The relationship may not exist; if the relationship does not exist, the character and he / she are strangers). Now, please consider how these things will affect your impression of him / her and how your relationship will change (of course, it is also possible that there will be no change). Please output the results in the following JSON format: {{"des":"Impression","relationship":"Relationship","strength":"Relationship Strength"}}. Please note that events may not have any effect on impression or relationship; in this case, output the unchanged aspects as is. Please note that impressions should not include specific events. Please note that an impression is essentially an answer to what kind of person someone is, including basic information. Please note that relationships should be summarized in one word. If you cannot summarize the current relationship in one accurate word, please add some brief descriptions.
[0070] In the reflection and optimization of memory content: using the semantic embedding vectors of long-term memory items, the cosine similarity between memory fragments is calculated to construct a semantic association graph with memory texts as nodes and semantic similarity as edge weights. The Louvain community detection algorithm was used to analyze the semantic association graph. Community segmentation identifies semantically highly related memory clusters; each community represents a set of memory fragments with inherent semantic consistency, helping to reveal underlying thematic structures and experiential patterns. Specific implementation includes: mapping semantic association maps... Each node in the graph is initially an independent community; the semantic association graph is then used to... Each node in Move to the community of the neighboring node and select the operation that increases the network module degree Q the most: (7); in, The change in modularity Q, For semantic association graph The total number of all edges in the array. The total weight of the edges in the current community. This represents the total weight of all edges connecting the current community to the outside world. For nodes The degree, For nodes The number of edges connected to other nodes within the current community; Treat each community as a new node and construct an updated semantic relationship graph. Semantic association graph The edge weights in the semantic association graph are The sum of the weights of all edges between corresponding communities This can be expressed by the following formula (8): (8); in, and Belongs to semantic association graph Different communities in China For semantic association graph Middle node and nodes Edge weights between them; Repeat the above steps until a stable community segmentation result is obtained. For each identified community, summarize the memory texts it contains, and guide the large language model to generate question and answer information using prompt words. Merge the question and answer information into reflective memory entries, and generate corresponding semantic embedding vectors using the BERT model. Calculate the cosine similarity between the new memory and nodes in the existing memory graph to determine whether to replace the original node, thereby updating the memory information in the vector database. Specifically, if the semantic similarity between a new memory and an existing node exceeds a set threshold, they are considered to have high semantic overlap, and the new memory replaces the original node; otherwise, insert it as an independent node into the graph, establish connections between related nodes, and update the database index. The attention score of the new memory is a weighted sum of the attention scores of related memories.
[0071] The following are prompt word templates for summarizing experiences: Prompt = Your name is {...}, your identity is {...}, your age is {...}, and your personality is {...}. You have lived a day and remembered some related events as follows: {...}. Based on the above information, please ask some questions, answer these questions, and explain why you answered them. Be sure to stay true to your role and not exceed the character's cognitive abilities. Please note that you must strictly adhere to the following JSON format: {[{"question":"question","answer":"answer","reason":"reason for answering","memories":"memories on which memories led to your answer"}, ...]}. Please note that you can ask multiple questions, but you must be able to answer each question correctly based on the information provided and give a reason.
[0072] In one possible embodiment, the agent architecture module further includes a planning unit, which is the core component for the agent to achieve autonomous decision-making and behavior execution. Its functional structure can be divided into two sub-modules: decision generation and action execution. Based on comprehensive sensory input, memory information, and role state, this module drives the agent to make behavioral choices that conform to contextual logic, and transforms these choices into specific executable instructions for the client.
[0073] Combination Figure 11 and Figure 12 As shown, Figure 11 This is a decision flowchart provided in one embodiment of this application. Figure 12 This is a flowchart of action execution provided in one embodiment of this application. The planning unit is used to: take the highest priority first event information output by the perception unit as the current event to be processed; generate a semantic embedding vector for the current event to be processed using the BERT model; in short-term memory, retrieve relevant memory fragments by calculating the cosine similarity between the semantic embedding vector and the memory entries stored in short-term memory, and using the Top-K strategy; in long-term memory, perform a similarity search based on the semantic embedding vector of the current event to be processed to obtain relevant historical memories in the vector database; after the retrieval is completed, process the memory entries... Sort the entries to obtain the memory entries. corresponding sequence : (9); in, , , For weight parameters, This is the normalization function; Representing memory entries Freshness of time, Indicates a memory entry. Indicates attention score; Memory Entries Time freshness This can be expressed by the following formula (10): (10); in, Indicates the current timestamp. The timestamp representing the memory entry. It is a constant used to avoid the denominator being zero; Decision information is generated based on the memory entries of the top-K sequences as reference information; the decision information includes ignoring, talking, or non-continuous actions.
[0074] For example, ultimately, the Top-K highest-scoring memories are selected as reference information for subsequent decision-making. This mechanism not only improves the accuracy and context adaptability of decision-making but also ensures the system's response speed and efficiency. Subsequently, the system integrates the role's basic attribute information (such as name, gender, identity, age, personality, current mood, recent self-evaluation, etc.), environment (scene, time, goal, ongoing task), perceived external events, and related memories into a unified prompt word, and submits it to a large language model for comprehensive reasoning to output specific behavioral decisions. Based on the model's output, behavioral decisions are categorized into the following three types: (a) Ignore: If the agent determines that the current event is not important or does not require intervention, it maintains the current state and does not make any explicit action response. (b) Converse: When the decision result is "converse," the system transmits relevant information to the dialogue unit, generates a natural language response based on the agent's information, and achieves coherent and personalized interpersonal interaction. (c) Non-continuous actions: These include actions such as "eating," "checking notices," and "performing," which do not require continuous interaction with other agents. These behaviors will be further decomposed into multiple steps. After the system updates the character's action status, it encapsulates the corresponding instruction sequence into a data structure that the client can parse, stores it in the action queue, and sends it to the client for execution.
[0075] In one possible embodiment, the agent architecture module further includes a dialogue unit, which is used to realize natural language interaction between agents through a dialogue group mechanism. The natural language interaction includes a dialogue start phase, a dialogue progress phase, and a dialogue end phase. In the dialogue start phase, an agent sends a dialogue request to a target agent, which decides whether to accept based on its current behavior. In the dialogue progress phase, a dialogue group instance is created. This instance records member information, dialogue history, and relationship status, and generates natural language responses that conform to the role settings using a sequential dialogue mechanism. In the dialogue end phase, when any agent exits the dialogue, the dialogue history is summarized, the summarized information is stored in the agent's long-term memory, and the mood and relationship strength of the relevant agents are updated. If all members exit, the dialogue group resources are released, and the corresponding data structure is deleted.
[0076] For example, the Conversation Group mechanism is used to maintain structured information such as basic information of participating members, relationship status, historical dialogue content, and related memories during conversations. Through this mechanism, the system can effectively support contextually coherent and semantically consistent dialogue interactions between multiple roles.
[0077] The dialogue behavior of an intelligent agent can be divided into three stages: dialogue start, dialogue progress, and dialogue end.
[0078] When a certain intelligent agent Decide to work with other intelligent agents When engaging in conversation, a dialogue request will be sent to the target agent. At this point, the target agent receives this event in its perception unit and, based on the behavior type output by the current decision module,... Determine whether to accept this dialogue request. If If the agent chooses to "converse," the dialogue process officially begins; otherwise, it terminates. If the agent decides to accept the dialogue request, the system creates a dialogue group instance. The system initializes its information, including registering the identities and basic attributes of dialogue members, loading and recording the relationship state between members, initializing the dialogue history and including the initial trigger statement, etc. After the dialogue officially begins, the system adopts a sequential dialogue mechanism, where each agent generates only one reply at a time, taking turns speaking. In each round of dialogue, the information of the current speaker... Personality traits Mood state Self-description Dialogue with History Relationship Network and long-term memory related to the current topic This content will be integrated into the prompts and then fed into a large language model to generate natural language responses that match the character's role. Let the current speaking agent be... Its set of states is represented as: ; This information will be fed into the large language model as part of the input prompt. Generate reply content with reasoning : ; When an agent is in a conversational state but its decision result is no longer "conversation", it is considered that the agent has voluntarily withdrawn from the dialogue. At this time, the system performs the following operations: removes the member's information from the dialogue group; and calls the large language model to analyze the content of this conversation. Summarize and extract key information. .
[0079] ; The summarized information K is stored as a new memory entry in the long-term memory bank. The moods of the relevant characters are updated based on the dialogue content. Relationship strength Internal state attributes: ; ; in, and These represent update functions for mood and relationship status, respectively.
[0080] If there are no more members in a chat group, the system determines that the chat has ended, releases the group resources, and deletes the corresponding data structure.
[0081] User platform module 130 provides a visual operation interface, allowing users to configure agent attributes and simulation parameters, and display the interaction process and results of multiple agents in real time.
[0082] To enhance the system's user-friendliness and ease of operation, this application designs a comprehensive management interface with complete functions and a clear structure, enabling users to efficiently perform various operations and intuitively view relevant data. For example... Figure 13 As shown, Figure 13 This application provides a system flowchart. The web-based homepage serves as the system login interface. This page not only briefly explains the system's core functions and application scenarios but also provides a user login entry point and a quick navigation path to the simulation environment.
[0083] like Figure 14 As shown, Figure 14 This is a flowchart of a user platform module provided in this application. On the homepage, the system clearly defines its functions through concise text descriptions and provides intuitive operation guidance to help users quickly understand the basic usage of the platform. Users can conveniently access multiple functional modules through the top navigation bar, including user login and agent behavior simulation. Particularly in the simulation module, users can complete a series of operations such as agent parameter configuration, simulation operation, and result analysis.
[0084] like Figure 15 As shown, Figure 15 This is a schematic diagram of a system simulation interface provided in this application. Currently, as the system is still in the testing phase, only a single administrator account (Admin) has been set up, and the registration function is not yet available. The simulation interface is the core interactive area of the user platform. In this interface, users can view and edit agent information, set simulation parameters, and view simulation results.
[0085] The simulation interface adopts a partitioned layout, divided into three main parts: a navigation area, a function area, and a display area, to improve user operation efficiency and interface interaction experience. The function area integrates three core functions: agent information management, simulation parameter configuration, and perspective switching. For example, Figure 16 As shown, Figure 16 This is a schematic diagram of an interface for viewing and modifying basic role information provided in this application. In terms of agent information management, users can view the basic information of a specific agent by clicking on it, including key attributes such as role background and behavioral patterns. Simultaneously, users can modify the information of selected agents according to their own needs, thereby achieving personalized configuration.
[0086] In addition, such as Figure 17 As shown, Figure 17 This diagram illustrates an interface for viewing and modifying agent relationships provided in this application. Users can also view the relationships between agents by clicking the "Relationship Graph" button in the upper right corner. This function not only supports the visualization of the interaction network between agents but also allows users to directly adjust the corresponding connections in the relationship graph, enabling flexible modification of the relationship structure between agents according to experimental needs.
[0087] like Figure 18 As shown, Figure 18This diagram illustrates a simulation parameter modification and perspective switching interface provided in this application. Regarding simulation parameter configuration, users can set basic simulation parameters, such as the number of agents and the number of simulation rounds, in the function area to meet the needs of different experimental scenarios. Simultaneously, the system supports multi-view switching, allowing users to observe the simulation process from different angles, enhancing the comprehensiveness and flexibility of simulation analysis. This interface design not only enhances the system's functionality and interactivity but also provides users with a clearly structured and intuitive platform, facilitating in-depth research into the behavioral patterns of agents and their interaction mechanisms.
[0088] like Figure 19 As shown, Figure 19 This is a schematic diagram of a simulation display interface provided in this application. The simulation display area, as the core component of the simulation interface, is mainly used to display the simulation process and final results in real time. When the user clicks the "Start" button in the upper right corner of the interface, the system sends a simulation request to the server and starts the simulation task. After receiving the request, the server performs calculations according to the set logic, generates the simulation video result, and returns it to the web frontend for playback and display. During the simulation, the system dynamically displays a progress bar in the workspace to visually reflect the current task's execution progress, allowing the user to monitor the simulation status in real time. Once the progress bar is fully loaded, the system automatically enters the result display stage, and the simulation video will play in the display area, allowing the user to intuitively view and analyze the simulation effect.
[0089] This application proposes a large-scale intelligent agent complex interaction behavior system, comprising: a virtual environment construction module, an intelligent agent architecture module, and a user platform module; the virtual environment construction module is used to construct a structured virtual space that fits the target application scenario. Since the structured virtual space includes functional area divisions and interactive semantic information, it can provide the intelligent agent with the ability to perceive its surrounding environment; the intelligent agent architecture module is used for generative large language model driving and includes an intelligent agent profiling unit and multiple cognitive units; since each cognitive unit collaboratively realizes the intelligent agent's perception of the environment, experience storage, behavior planning, and semantic interaction among multiple intelligent agents, it can support the complete decision-making process of the intelligent agent from environmental perception to behavior output, thus endowing it with… It more closely resembles human behavior and adaptability; the user platform module provides a visual operation interface, supports users in configuring agent attributes and simulation parameters, and displays the interaction process and results of multiple agents in real time, thus improving the user-friendliness and ease of operation of the system; based on this, the large-scale intelligent agent complex interaction behavior system provided in this application can construct a three-layer system integrating a virtual environment, a humanoid intelligent agent architecture, and a user interaction platform. Through a large language model, it drives intelligent agents to autonomously generate credible social interaction behaviors in highly semantic historical scenarios, thereby improving the decision-making efficiency of intelligent agents in environmental perception and realizing efficient, credible, and socially logical complex interactions.
[0090] like Figure 20 As shown, one embodiment of this application proposes a method for simulating complex interactive behaviors of large-scale intelligent agents, which can be applied to the large-scale intelligent agent complex interactive behavior system described above. The method includes: Step 2010: Construct a structured virtual space that fits the target application scenario; wherein, the structured virtual space includes functional area divisions and interactive semantic information; Step 2020: Generative large language model driving the intelligent agent's agent profiling unit and multiple cognitive units; wherein, the agent profiling unit includes basic information dimension, psychological state dimension, time awareness dimension and social awareness dimension, to provide the intelligent agent with cognitive and behavioral capabilities; each cognitive unit works together to realize the intelligent agent's perception of the environment, experience storage, behavior planning and semantic interaction between multiple intelligent agents; Step 2030: Provide a visual operation interface to allow users to configure agent attributes and simulation parameters, and display the interaction process and results of multiple agents in real time.
[0091] The following is a detailed description of the implementation details of a method for simulating complex interactive behaviors of large-scale intelligent agents proposed in this embodiment. The following content is only for the convenience of understanding and is not necessary for implementing this solution.
[0092] The steps described above are for clarity only. In implementation, they can be combined into one step, or some steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0093] Another embodiment of this application provides an electronic device, such as Figure 21 As shown, it includes a processor 211 and a memory 212. The memory 212 stores instructions that the processor 211 can execute. When the processor 211 is configured to execute the instructions, the electronic device can realize a method for simulating complex interactive behaviors of a large-scale intelligent agent as described in the above method embodiment.
[0094] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0095] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory, on the other hand, is used to store data used by the processor during operation.
[0096] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a method for simulating complex interactive behaviors of a large-scale intelligent agent as described in the above method embodiments.
[0097] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0098] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A system for simulating complex interactive behaviors of large-scale intelligent agents, characterized in that, include: Virtual environment construction module, intelligent agent architecture module, and user platform module; The virtual environment construction module is used to build a structured virtual space that fits the target application scenario; the structured virtual space includes functional area divisions and interactive semantic information; The agent architecture module is used to drive generative large language models and includes an agent profiling unit and multiple cognitive units. The agent profiling unit includes basic information dimensions, psychological state dimensions, time awareness dimensions, and social awareness dimensions to provide the agent with cognitive and behavioral capabilities. The cognitive units work together to realize the agent's perception of the environment, experience storage, behavior planning, and semantic interaction among multiple agents. The user platform module provides a visual interface that allows users to configure agent attributes and simulation parameters, and displays the interaction process and results of multiple agents in real time.
2. The system according to claim 1, characterized in that, The virtual environment construction module is used to convert visualized space into a structured natural language description using map semantics. The structured natural language description includes spatial attribution, functional attributes, and dynamic state information. The map semantics method includes map digitization and map initialization.
3. The system according to claim 2, characterized in that, The digitization of maps includes: A set of scenarios for the target application scenario Each scene in the program is divided into functional areas. The system performs semantic recognition on the divided functional areas and interactive objects to obtain a set of regions. and item collection , , And assign a unique identifier code to each of them. , Establish a mapping table from codes to regions or objects. Among them, the mapping table This can be expressed by the following formula (1): (1); The structured virtual space is divided into tiles, and the coordinates of each tile are... By assigning coded values representing functional areas or interactive objects, a digital map is obtained. Among them, digital maps This can be expressed by the following formula (2): (2); The map initialization includes: Traverse all coordinate points From digital maps Extract the encoding information at the corresponding position. Through the mapping table Analyze each coordinate point The region type and interactive objects; Based on each coordinate point The region type and interactive objects are used to construct a structured dictionary data containing multi-level semantic information. and obtain tuples , and event information collection , Among them, dictionary data The fields include world Scene Functional areas Interactive objects and dynamic event states And expressed by the following formula (3): (3)。 4. The system according to claim 1, characterized in that, Multiple cognitive units include a perception unit, which is used for an attention-score-based event filtering mechanism and is responsible for collecting, identifying, and initially processing events in the environment; the event filtering mechanism includes an event perception phase and an event evaluation phase; During the event perception phase: Current intelligent agent Acquire all event information in the current scene; where all event information includes relevant event information of interactive objects and other intelligent agents present; the event information includes the actor of the event, the action performed by the actor, the object affected by the action, the location of the event, and the timestamp of the event. Based on the information source type of the event information, a description of the event information is generated; where the information source type includes events of interactive objects and events of other intelligent agents; During the event evaluation phase: Based on the attention scoring model, information source type, and event information description, all event information is prioritized and the first event information with a higher than preset priority is selected. The attention scoring model is used to evaluate the multidimensional attributes of the event and is constructed using knowledge distillation and efficient parameter fine-tuning techniques, and is trained using event score data.
5. The system according to claim 4, characterized in that, Multidimensional attributes include perceptual saliency, task relevance, and emotional intensity; The description information based on the attention score model, information source type, and event information prioritizes all event information, including: If the information source type is an event of an interactive object, the attention scoring model uses task relevance to evaluate the descriptive information of the event information to obtain the priority of the event information. If the information source type is an event of other agents, the attention scoring model uses perceptual salience, task relevance, and emotional intensity to evaluate the descriptive information of the event information and obtain the priority of the event information.
6. The system according to claim 4, characterized in that, Multiple cognitive units also include memory units, which are connected to sensory units. Memory units include different levels of memory, including sensory memory, short-term memory, working memory, and long-term memory. Perceptual memory is used to acquire the current task context, obtain first event information from the perceptual unit, and determine the attention width threshold corresponding to the current task context; at the same time, an attention bottleneck mechanism is used to filter out second event information that is higher than the attention width threshold corresponding to the current task context from the first event information; Short-term memory is used to send event information from the second event information that is relevant to the decision-making task faced in the current task situation to working memory, so that working memory can support the execution of the decision-making task, while based on activation factors. The third event information is selected from the second event information and transferred to long-term memory so that long-term memory can store and retrieve the third event information; among them, the activating factor Indicates the first The degree of activation of a memory, when the activating factor Above the activation threshold Then the first One memory is transferred to long-term memory; Among them, the activation factor is calculated. Formula (4) is as follows: (4); in, This is the current timestamp. It is memory The total number of times it was perceived. It is the repetitive perception coefficient. It is memory Total number of searches It is the repetitive perception coefficient. That is, the current timestamp with the timestamp of the event Time difference, This represents the attenuation rate.
7. The system according to claim 6, characterized in that, The long-term memory stores information about the third event, including: Vector database technology is used to store and retrieve information about third events; The set consists of all memory vectors in the vector database. , ,in, Indicates the first One memory entry A dimensional semantic vector can be divided into several parts using an index structure. A subset; When new memory entries When long-term memory needs to be incorporated, the semantic embedding vector is extracted using the BERT model and expressed by the following formula (5): (5); New memory entries The corresponding semantic embedding vectors are stored in a vector database so that they can be used to query text. corresponding vector Retrieves text matching the query. The corresponding set of memory entries Specifically, it is represented by the following formula (6): (6)。 8. The system according to claim 7, characterized in that, Multiple cognitive units also include reflection units, which are used to perform structured updates on information in the vector database. The structured updates include dynamic updates of agent attribute states and reflective optimization of memory content. In the dynamic updating of agent attribute states: The retrieval-enhanced generation technique is used to extract all memory entries related to the current object from the vector database, and these memory entries are used as context and input into the large language model. By using pre-set structured prompt word templates, the large language model is guided to update and output the agent's internal state; In the process of reflecting on and optimizing the content to be memorized: By utilizing the semantic embedding vectors of long-term memory entries, cosine similarity is calculated between memory fragments to construct a semantic association graph with memory texts as nodes and semantic similarity as edge weights. ; The Louvain community detection algorithm is used to analyze the semantic association graph. Community segmentation was performed to identify semantically highly related memory clusters; Each community represents a set of memory fragments with inherent semantic consistency, specifically implemented as follows: semantic association graph Each node in the system initially forms an independent community; semantic association graph Each node in Move to the community of the neighboring node and select the operation that increases the network module degree Q the most: (7); in, The change in modularity Q, For semantic association graph The total number of all edges in the array. The total weight of the edges in the current community. This represents the total weight of all edges connecting the current community to the outside world. For nodes The degree, For nodes The number of edges connected to other nodes within the current community; Treat each community as a new node and construct an updated semantic relationship graph. Semantic Relationship Graph The edge weights in the semantic association graph are The sum of the weights of all edges between corresponding communities This can be expressed by the following formula (8): (8); in, and Belongs to semantic association graph Different communities in China For semantic association graph Middle node and nodes Edge weights between them; Repeat the above steps until a stable community segmentation result is obtained. At the same time, for each identified community, summarize the memory text contained therein, and guide the large language model to generate question and answer information through prompt words. The question and answer information are merged into reflective memory entries, and corresponding semantic embedding vectors are generated by encoding them through the BERT model. The cosine similarity between the vector embedding vector and the existing memory graph nodes is calculated to determine whether to replace the original nodes, thereby updating the memory information in the vector database.
9. The system according to claim 8, characterized in that, The intelligent agent architecture module also includes a planning unit, which is used for: The highest priority first event information output by the sensing unit is taken as the current event to be processed; The BERT model is used to generate semantic embedding vectors for the current event to be processed. In short-term memory, the Top-K strategy is used to retrieve relevant memory fragments by calculating the cosine similarity between the semantic embedding vector and the memory entries stored in short-term memory. In long-term memory, similarity search is performed based on the semantic embedding vector of the current event to be processed in order to obtain relevant historical memories in the vector database; After the retrieval is completed, the memory entries are... Sort the entries to obtain the memory entries. corresponding sequence : (9); in, , , For weight parameters, This is the normalization function; Representing memory entries Freshness of time, Indicates a memory entry. Indicates attention score; Memory Entries Time freshness This can be expressed by the following formula (10): (10); in, Indicates the current timestamp. The timestamp representing the memory entry. It is a constant used to avoid the denominator being zero; Decision information is generated based on the memory entries of the top-K sequences as reference information; the decision information includes ignoring, talking, or non-continuous actions.
10. The system according to claim 9, characterized in that, The agent architecture module also includes a dialogue unit, which is used to realize natural language interaction between agents through a dialogue group mechanism; the natural language interaction includes a dialogue start phase, a dialogue process phase, and a dialogue end phase. In the initial stage of the dialogue, the intelligent agent sends a dialogue request to the target intelligent agent, and the target intelligent agent decides whether to accept it based on its current behavior. During the dialogue phase, a dialogue group instance is created; the dialogue group instance is used to record member information, dialogue history and relationship status, and uses a sequential dialogue mechanism to generate natural language responses that conform to the role settings. During the dialogue termination phase, when any agent leaves the dialogue, the dialogue history is summarized, the summary information is stored in the agent's long-term memory, and the mood and relationship strength of the relevant agents are updated; if all members leave, the dialogue group resources are released and the corresponding data structures are deleted.