Intelligent agent simulation decision-making method and device based on large language model

Through the agent simulation decision-making method based on large language model, the problem of insufficient flexibility, accuracy and transparency of agent decision-making in the prior art is solved, and an efficient, flexible and explainable decision-making process is achieved.

CN120069585AActive Publication Date: 2025-05-30GUANGDONG NORMAL UNIV WEIZHI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510043053.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-30
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing intelligent simulation decision-making technology has problems such as rule design limitations, data requirements and feature selection problems, low development efficiency, black boxing of decision-making and insufficient multi-source information processing capabilities.

Method used

Adopt the agent simulation decision-making method based on the large language model, and serialize the agent's basic attribute information and real-time environment information by obtaining the agent's basic attribute information and the process of real-time environment information and passing it to the decision module. The decision module uses a large language model to transform data into natural language descriptions, combining historical decision records and behavioral libraries to generate decision results, including action plans and decision explanations.

Benefits of technology

It enhances the flexibility and adaptability of agent decision-making, improves the accuracy and transparency of decision-making, reduces development efficiency, and can efficiently integrate and process multi-dimensional data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069585A_ABST
    Figure CN120069585A_ABST
Patent Text Reader

Abstract

The invention provides an agent simulation decision-making method and device based on a large language model, and the method comprises the steps: obtaining the basic attribute information of an agent and the environment information of a real-time disaster rescue scene, carrying out the serialization processing of the information, and transmitting the information to a decision-making module. And the decision module converts the serialized data into natural language description by using a large language model, generates a decision result containing an action scheme and decision explanation in combination with historical decision records and a behavior library, transmits the decision result to the execution module for execution, and records the decision result in a message historical record database. According to the method, the flexibility and adaptability of agent decision making are enhanced, so that the agent can flexibly cope with changeable disaster rescue situations, and accurate and efficient decision making in a complex or unknown environment is ensured. And meanwhile, multi-dimensional data are efficiently integrated, information limitation is avoided, and the decision accuracy is improved. The decision-making process is transparent and explainable, and credibility and understanding are enhanced And the development efficiency is improved by utilizing a large language model, and higher timeliness is brought to actual system development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent agent simulation decision-making, and particularly to an intelligent agent simulation decision-making method based on a large language model, an intelligent agent simulation decision-making device based on a large language model, an electronic device, and a computer-readable medium. Background Art

[0002] In the field of intelligent agent simulation decision-making, the existing technologies can be mainly classified into the following methods: (1) Rule-based decision-making system Such systems rely on decision rules written manually, usually constructed through expert systems or knowledge bases. These rules reason according to different situations of environmental states and agent attributes, and give corresponding decisions. For example, in some simple simulation environments, decisions can be made according to pre-set rules (such as "if the water depth exceeds 50 cm, the intelligent agent should choose to move to a high place").

[0003] (2) Traditional machine learning methods Including decision trees, support vector machines (SVM), linear regression, etc. These methods predict the behavior of intelligent agents by training models. These methods usually rely on a large amount of training data and feature selection, and can model certain variables in the environment, but they may encounter performance bottlenecks for complex and changing environments.

[0004] (3) Reinforcement learning (RL) Reinforcement learning methods are trained through the interaction between intelligent agents and the environment. The intelligent agent learns the optimal strategy according to the feedback reward signal (such as reward value or penalty value) of the environment. Reinforcement learning has achieved good results in many complex decision-making problems, but it usually requires a large amount of training time, and for some high-dimensional and diverse environments, the training process may be very complex and the effect is limited.

[0005] (4) Deep learning (such as deep reinforcement learning) Feature extraction and decision prediction are carried out through neural network models (such as convolutional neural network CNN, recurrent neural network RNN, etc.). These methods can process more complex high-dimensional input data and have strong generalization ability, but they also face problems such as strong data dependence, large computational amount, and unstable training.

[0006] The above existing technologies have the following disadvantages: (1) Limitations in rule design Rule-based systems rely on experts to manually write decision rules, which are usually static and limited. When faced with dynamic, unknown, or highly complex environments, the coverage of the rules is insufficient, and the agent may not be able to make reasonable decisions. Moreover, the number of rules grows exponentially with the increase in environmental complexity, which limits the scalability and flexibility of the system.

[0007] (2) Data requirements and feature selection problems Traditional machine learning methods usually require a large amount of labeled data to train the model, and this data must fully represent all possible environmental states. However, in many practical applications, it is very difficult and costly to obtain high-quality labeled data. At the same time, how to extract the most effective features from complex raw data depends on human experience and design, and in the face of new environments, existing features may not be able to effectively capture key information.

[0008] (3) Low development efficiency Although methods such as reinforcement learning can improve decision-making quality in multiple interactions, they usually require a long training process, and a large number of samples and computing resources are often required during the training process. This is difficult to provide fast and flexible decision support for real-time decision-making or dynamically changing environments.

[0009] (4) The black box nature of decision-making and lack of interpretability Most existing machine learning and deep learning methods, such as deep neural networks and reinforcement learning, are often "black box" models. Although they can generate relatively good decision results, they lack interpretability of the decision-making process. Especially in important fields such as security and ethics, users or decision-makers need to understand the decision-making basis of the agent in order to conduct supervision or adjustment. Existing technologies cannot clearly explain why the agent makes a specific decision, reducing the user's trust and acceptability.

[0010] (5) Unable to effectively process multi-source information and complex relationships In multi-agent systems or when dealing with information that requires processing complex multi-dimensional, spatio-temporal associations, traditional methods often have difficulty comprehensively processing a large amount of heterogeneous data and changing decision-making factors. While deep learning-based systems can handle more complex input data, there are still significant challenges for the complex temporal relationships and context dependencies between the agent and the environment. Summary of the Invention

[0011] In view of the above problems, the present invention is proposed to provide an intelligent agent simulation decision-making method based on a large language model and a corresponding intelligent agent simulation decision-making device based on a large language model, an electronic device, and a computer-readable medium that overcome the above problems or at least partially solve the above problems.

[0012] The present invention discloses an intelligent agent simulation decision-making method based on a large language model, and the method includes: Obtain the basic attribute information of the intelligent agent, and collect the current environmental information of the disaster rescue scenario in real time through multiple data sources; Perform serialization processing on the basic attribute information of the intelligent agent and the current environmental information to obtain serialized basic attribute data and serialized environmental information data, and transfer the serialized basic attribute data and serialized environmental information data to the decision-making module; Convert the serialized basic attribute data and serialized environmental information data into natural language descriptions of basic attributes and natural language descriptions of environmental information through the large language model of the decision-making module, and generate a decision result according to the natural language descriptions of basic attributes and environmental information, in combination with historical decision records and the behavior library; The decision result includes an action plan and a decision explanation; Transmit the decision result returned by the decision-making module to the execution module for execution, and record the decision result in the message history database.

[0013] Optionally, The behavior library includes standard operation procedures and standard behavior patterns; The standard operation procedures include evacuation route selection, rescue supply distribution, and search for emergency shelters; The standard behavior patterns include behavior classification, behavior pattern mapping set for different environmental states, and behavior decision priority rules.

[0014] Optionally, the method further includes: The execution module returns the execution result to the decision-making module, and the decision-making module performs dynamic learning and optimizes the large language model based on the execution result to ensure that the decision matches the current environmental state.

[0015] Optionally, The current environmental information of the disaster rescue scenario includes one or more of the occurrence of natural disasters, rescue resource situation, rescue personnel situation, traffic conditions, geographical information, and weather and climate information; If the disaster rescue scenario is a storm surge scenario, the environmental information includes one or more of the change of water level depth, traffic conditions, topographic map, population distribution, rescue resource location, shelter location, shelter occupancy, presence of rescue personnel, environmental temperature, and environmental humidity.

[0016] Optionally, The historical decision records are recorded in the message history database and include an action plan and a decision explanation.

[0017] Optionally, the method further includes: The large model makes decisions regularly at preset decision time intervals and / or determines whether to adjust the decision based on changes in environmental information.

[0018] The present invention also discloses an intelligent agent simulation decision-making device based on a large language model, and the device includes: An information real-time collection module, configured to obtain the basic attribute information of the intelligent agent and collect the current environmental information of the disaster rescue scenario in real time through multiple data sources; A serialization processing module, configured to perform serialization processing on the basic attribute information of the intelligent agent and the current environmental information to obtain basic attribute serialized data and environmental information serialized data, and transmit the basic attribute serialized data and the environmental information serialized data to the decision-making module; A decision result generation module, configured to convert the basic attribute serialized data and the environmental information serialized data into a natural language description of the basic attributes and a natural language description of the environmental information through the large language model of the decision-making module, and generate a decision result according to the natural language description of the basic attributes and the natural language description of the environmental information, in combination with historical decision records and a behavior library; the decision result includes an action plan and a decision explanation; A decision result transmission and recording module, configured to transmit the decision result returned by the decision-making module to the execution module for execution, and record the decision result in the message history database.

[0019] Optionally, The behavior library includes standard operation procedures and standard behavior patterns; The standard operation procedures include evacuation route selection, rescue material distribution, and search for emergency shelters; The standard behavior patterns include behavior classification, behavior pattern mapping set for different environmental states, and behavior decision priority rules.

[0020] Optionally, the device further includes: A decision model dynamic update module, configured to the execution module returns the execution result to the decision-making module, and the decision-making module performs dynamic learning and optimization of the large language model based on the execution result to ensure that the decision matches the current environmental state.

[0021] Optionally, The current environmental information of the disaster rescue scenario includes one or more of the occurrence of natural disasters, rescue resource situation, rescue personnel situation, traffic conditions, geographical information, and weather and climate information; If the disaster rescue scenario is a storm surge scenario, the environmental information includes one or more of the change in water level depth, traffic conditions, topographic map, population distribution, rescue resource location, shelter location, shelter occupancy, whether there are rescue personnel, environmental temperature, and environmental humidity.

[0022] Optionally, Historical decision records are recorded in the message history database, including action plans and decision explanations.

[0023] Optionally, the device further includes: A timing dynamic decision module, configured to make the large language model make decisions regularly at a preset decision time interval and / or determine whether to adjust the decision according to changes in environmental information.

[0024] The present invention also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; When the processor is used to execute the program stored on the memory, it implements the intelligent agent simulation decision method based on the large language model as described in the present invention.

[0025] The present invention also discloses one or more computer-readable media, on which instructions are stored. When executed by one or more processors, the instructions cause the processors to execute the intelligent agent simulation decision method based on the large language model as described in the present invention.

[0026] The present invention has the following advantages: In the intelligent agent simulation decision method based on the large language model of the present invention, by obtaining the basic attribute information of the intelligent agent and the environmental information of the real-time disaster rescue scenario, these information are serialized and transmitted to the decision module. The decision module uses the large language model to convert these serialized data into natural language descriptions, and combines historical decision records and behavior libraries to generate decision results including action plans and decision explanations. Then, the decision results are transmitted to the execution module for execution and recorded in the message history database. This method greatly enhances the flexibility and adaptability of the intelligent agent's decision-making, enabling the intelligent agent to flexibly respond to changing disaster rescue situations and task requirements, ensuring accurate and efficient decision-making in complex or unknown environments. Secondly, by efficiently integrating and processing multi-dimensional data, the limitations of insufficient information processing or information silos are avoided, and the accuracy of decision-making is improved. Furthermore, the transparency and interpretability of the decision-making process make the decision-making basis of the intelligent agent clearly traceable, enhancing the credibility and understandability of the decision-making. Finally, by utilizing the powerful computing power of the large language model, this method significantly improves the development efficiency of the intelligent agent's decision-making, bringing higher timeliness to the actual system development. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of the steps of an intelligent agent simulation decision method based on a large language model provided by an embodiment of the present invention; Figure 2It is a flowchart of intelligent agent simulation decision-making based on a large language model provided by an embodiment of the present invention; Figure 3 It is a block diagram of the structure of an intelligent agent simulation decision-making device based on a large language model provided by an embodiment of the present invention; Figure 4 It is a block diagram of an electronic device provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of a computer-readable medium provided by an embodiment of the present invention. Detailed implementation manners

[0028] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0029] Refer to Figure 1 , which shows a flowchart of the steps of an intelligent agent simulation decision-making method based on a large language model provided in an embodiment of the present invention. Specifically, it may include the following steps: Step 101, obtain the basic attribute information of the intelligent agent, and collect the current environmental information of the disaster rescue scenario in real time through multiple data sources; Step 102, perform serialization processing on the basic attribute information of the intelligent agent and the current environmental information to obtain basic attribute serialized data and environmental information serialized data, and transfer the basic attribute serialized data and environmental information serialized data to the decision-making module; Step 103, use the large language model of the decision-making module to convert the basic attribute serialized data and environmental information serialized data into natural language descriptions of basic attributes and natural language descriptions of environmental information, and generate a decision result based on the natural language descriptions of basic attributes and environmental information, combined with historical decision records and the behavior library; the decision result includes an action plan and a decision explanation; Step 104, transmit the decision result returned by the decision-making module to the execution module for execution, and record the decision result in the message history database.

[0030] The intelligent agent simulation decision-making method based on a large language model of the present invention aims to improve the decision-making efficiency and accuracy of intelligent agents in disaster rescue scenarios through intelligent means. This method integrates the powerful natural language processing ability of the large language model, realizing the effective integration, serialization processing, and intelligent decision generation of the basic attributes of intelligent agents and real-time environmental information.

[0031] Specifically, refer to Figure 2 , the specific process of the intelligent agent simulation decision-making method based on a large language model of the present invention is as follows: 1. The intelligent agent decision-making starts The decision-making process starts from this step, marking the initiation of the agent's preparation for decision-making. At this point, the agent is ready to collect relevant information from multiple data sources for subsequent decision-making.

[0032] 2. Serialization of Basic Information Parameters Before making a decision, the agent needs to process its basic information, such as height, weight, gender, age, personality, health status, education level, etc. This information will be serialized into JSON format for subsequent processing and analysis by the system.

[0033] 3. Serialization of Environmental Information Parameters The agent's environmental perception module provides various types of information in the current environment, including external conditions such as the occurrence of natural disasters (e.g., the depth of a flood), the presence of rescue personnel, environmental temperature, humidity, etc. This information is also serialized into JSON format to ensure effective data transmission and processing.

[0034] 4. Transmission of Serialized Data to the LLM This invention enables the agent to automatically adjust its decision-making strategy according to the changing environment and different task requirements through the large language model (LLM), thereby enhancing the flexibility and adaptability of decision-making.

[0035] Once the basic information and environmental information are serialized, they will be transmitted to the large language model (LLM), such as ChatGPT, Spark Large Language Model. The LLM will process this data and convert this information into natural language descriptions. Specifically, it includes: (1) Natural Language Description of Basic Attributes: The LLM converts the agent's basic information into an easy-to-understand natural language description. This conversion ensures that the data can be accurately interpreted through the language model and provides support for subsequent decision-making.

[0036] (2) Natural Language Description of Environmental Information: The environmental information will also be converted into a natural language description according to a preset format. This conversion provides a clearer environmental background for the LLM to help it comprehensively understand the current situation.

[0037] This invention serializes information and conducts natural language descriptions through the large language model, enabling unified and effective processing of various heterogeneous information. At the same time, various types of information (such as the agent's basic information, environmental information, etc.) can be converted into an easy-to-understand and operable format, thereby improving the decision-making quality of the agent in a complex environment.

[0038] 5. Make a Decision Based on the Behavior Library and Historical Records and Explain the Reasons After receiving the natural language descriptions of the basic information and environmental information, the LLM combines the standard operation processes, standard behavior patterns, and historical decision records stored in the behavior library to conduct decision analysis and give a reasonable action plan. The LLM not only provides action decisions but also generates natural language explanations for the decisions, detailing "why this decision is made", thus greatly enhancing the transparency and credibility of the decision-making process and improving the acceptability and controllability of the agent in human-machine collaboration. These action decisions and their explanations will be recorded in the message history database for subsequent analysis and feedback.

[0039] The LLM utilizes the profound language understanding and generation capabilities cultivated during training on a large amount of text data, and has rich common sense knowledge, domain expertise, and factual data. Therefore, agents based on the LLM usually only need very few samples to perform well in new tasks. Their excellent generalization ability enables them to perform well in situations they have never encountered before. The present invention optimizes the decision-making process through large language models, enabling the agent to quickly generate decisions at a relatively low computational cost and improving the efficiency of the development process of the agent's decision-making module.

[0040] 6. Transmission of decision results to the execution module After the decision is completed, the decision results generated by the LLM will be transmitted to the execution module of the agent. The execution module is responsible for implementing specific actions according to the decisions of the LLM, such as performing rescue tasks, adjusting behavior strategies, etc.

[0041] 7. End The process ends, marking the completion of a single decision-making process of the agent. Subsequently, new decision-making processes may be initiated according to new environmental changes or task requirements.

[0042] In an embodiment of the present invention, The behavior library includes standard operation processes and standard behavior patterns; The standard operation processes include evacuation route selection, rescue supply distribution, and search for emergency shelters; The standard behavior patterns include behavior classification, behavior pattern mapping set for different environmental states, and behavior decision priority rules.

[0043] The method of the present invention includes a behavior library designed for disaster rescue scenarios, which contains standard operation processes such as evacuation route selection, rescue supply distribution, and search for emergency shelters. The core of the behavior library lies in integrating and summarizing domain-specific behavior patterns. Taking the storm surge scenario as an example, the behavior library can be summarized based on research literature: by consulting research papers in the fields of disaster management, disaster psychology, and sociology, and summarizing the behavior patterns of people in storm surges.

[0044] The content of the domain-specific behavior library should cover the following aspects: Behavior classification: Refine behaviors into different types. For example: Individual behaviors: Escape alone, call for help, etc.

[0045] Group behaviors: Cooperative actions (such as family members evacuating together), congestion phenomena.

[0046] Behavior mapping of environmental factors: Set corresponding behavior patterns for different environmental states. For example: When the water level is below the knees, people tend to walk.

[0047] When the water level is above the waist, people tend to wait in place or use rescue equipment.

[0048] Decision priority rules: Provide priority rules for behavior selection for the agent. For example: The shelter preferably selects targets with high accessibility and sufficient capacity.

[0049] Preferably avoid dangerous areas (such as flooded bridges).

[0050] Through the domain-specific behavior library, the method of the present invention can improve the decision-making quality and action efficiency of the agent in disaster scenarios. For example, when a flood comes, it can quickly decide to send a humanoid agent to the nearest shelter or escape exit.

[0051] In an embodiment of the present invention, the method further includes: The execution module returns the execution result to the decision-making module, and the decision-making module performs dynamic learning and optimizes the large language model based on the execution result to ensure that the decision matches the current environmental state.

[0052] The present invention introduces a closed-loop feedback mechanism. The execution module can return the execution result to the decision-making module, and the update interval is short, enabling the LLM to perform dynamic learning and optimization based on new data to ensure that the decision matches the current environmental state.

[0053] In an embodiment of the present invention, The current environmental information in the disaster rescue scenario includes one or more of the occurrence of natural disasters, rescue resource situation, rescue personnel situation, traffic conditions, geographical information, and weather and climate information; If the disaster rescue scenario is a storm surge scenario, the environmental information includes one or more of the water level depth change situation, traffic conditions, topographic map, population distribution, rescue resource location, shelter location, shelter occupancy, whether there are rescue personnel, environmental temperature, and environmental humidity.

[0054] The present invention emphasizes how to integrate and process multi-dimensional information from different sources. The current environmental information in the disaster rescue scenario can include, but is not limited to, one or more of the occurrence of natural disasters, rescue resources, rescue personnel, traffic conditions, geographical information, weather and climate information, such as the processing flow of environmental monitoring data, real-time sensor data (such as water level sensors), and geographical information (such as topographic maps, building distributions). The present invention uses serialization technology (such as JSON format) to standardize multi-dimensional data, enabling different formats of data to be effectively integrated in the system and further analyzed by the LLM. After serialization, the present invention converts the data into natural language descriptions for semantic understanding and comprehensive decision-making by the large language model. For example: "The flood depth in the current area is 1.2 meters, the traffic congestion index is 75%, and the distance to the nearest shelter is 500 meters." Through multi-dimensional data processing, the present invention can provide more comprehensive and in-depth decision-making support. Through dynamic data updates and real-time feedback, the intelligent agent can also quickly respond to environmental changes in the disaster scenario. If the disaster rescue scenario is a storm surge scenario, the environmental information can include, but is not limited to, one or more of the changes in water level depth, traffic conditions, topographic maps, population distribution, rescue resource locations, shelter locations, shelter occupancy, the presence of rescue personnel, environmental temperature, and environmental humidity. In the storm surge scenario, the intelligent agent not only considers the current water level but also combines traffic conditions, population distribution, and the location of rescue resources to make more reasonable evacuation and rescue decisions. For example, the sensor real-time monitors that the current water level depth is 1.2 meters, the water depth data of the environmental perception module of the intelligent agent is 1.2 meters, the traffic congestion index is 75%, and Shelter A is full. The data is transmitted to the LLM through serialization and converted into a natural language description. The LLM combines historical records and the behavior library to decide to go to Shelter B.

[0055] In one embodiment of the present invention, the method further includes: The large model makes decisions regularly at a preset decision time interval and / or determines whether to adjust the decision according to changes in environmental information.

[0056] The present invention focuses on evacuation and rescue in emergency disaster scenarios, requiring the intelligent agent to be able to quickly respond to real-time changing environmental information (such as flood spread, shelter occupancy, etc.). Therefore, real-time performance is the core goal of the technical solution. And the decision time interval of the LLM intelligent agent is relatively short, such as it can be set to 5 seconds. Therefore, the LLM can make decisions for the intelligent agent in real-time and frequently. And through the real-time acquisition and dynamic update mechanism of multi-dimensional information such as sensor data and geographical information, the intelligent agent can continuously receive the latest data on environmental changes and adjust its decisions.

[0057] The present invention has the following advantages: 1. Enhance the flexibility and adaptability of the intelligent agent's decision-making Traditional intelligent agent decision-making systems usually rely on rules or predefined models, lacking the flexibility and adaptability to handle complex dynamic environments. In contrast, the present invention utilizes an LLM to adjust the decision-making model according to real-time environmental changes, execution results, and historical data, enabling the intelligent agent to flexibly respond to changing situations and task requirements, and ensuring that the intelligent agent makes more accurate and efficient decisions in complex or unknown environments.

[0058] 2. Efficient integration and processing of multi-dimensional data The present invention serializes the basic information and environmental information of the intelligent agent into standardized data (such as JSON format) and performs natural language processing through an LLM, enabling efficient integration of multi-dimensional information from different sources. This approach can effectively avoid the limitations of insufficient information processing or information silos in traditional methods, providing a more comprehensive and accurate decision-making basis. Compared with existing methods, the present invention demonstrates stronger capabilities in processing complex and heterogeneous data, being able to more efficiently fuse and analyze multi-source information, and improving the accuracy of decision-making.

[0059] 3. Transparency and interpretability of the decision-making process Traditional decision-making systems (especially those based on deep learning) often suffer from the "black box" problem, that is, the decision results lack explanation and users cannot understand why the intelligent agent makes a specific choice. In contrast, the present invention provides transparency and traceability of the decision-making process through natural language explanations generated by an LLM. The intelligent agent can not only give the execution actions but also explain the basis for its decisions, making the decision-making process more credible and understandable. This transparency is of great value especially in critical fields (such as public safety, emergency evacuation, etc.).

[0060] 4. Improvement in the development efficiency of intelligent agent decision-making The decision-making process of the present invention, through the powerful computing power and parallel processing ability of the LLM, can quickly process a large amount of data and make immediate decisions. This shows higher efficiency in actual system development compared to traditional rule-based systems or deep learning models that require long training times.

[0061] It should be noted that for method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0062] Refer to Figure 3, showing a structural block diagram of an intelligent agent simulation decision-making device provided in an embodiment of the present invention, which may specifically include the following modules: The real-time information collection module 301 is used to obtain the basic attribute information of the intelligent agent and collect the current environmental information of the disaster rescue scenario in real time through multiple data sources; The serialization processing module 302 is used to perform serialization processing on the basic attribute information of the intelligent agent and the current environmental information to obtain the serialized basic attribute data and the serialized environmental information data, and transmit the serialized basic attribute data and the serialized environmental information data to the decision-making module; The decision result generation module 303 is used to convert the serialized basic attribute data and the serialized environmental information data into the natural language description of the basic attributes and the natural language description of the environmental information through the large language model of the decision-making module, and generate a decision result according to the natural language description of the basic attributes and the natural language description of the environmental information, in combination with the historical decision record and the behavior library; the decision result includes an action plan and a decision explanation; The decision result transmission and recording module 304 is used to transmit the decision result returned by the decision-making module to the execution module for execution, and record the decision result in the message history database.

[0063] Optionally, The behavior library includes standard operation procedures and standard behavior patterns; The standard operation procedures include evacuation route selection, rescue supply distribution, and search for emergency shelters; The standard behavior patterns include behavior classification, behavior pattern mapping set for different environmental states, and behavior decision priority rules.

[0064] Optionally, the device further includes: The decision model dynamic update module is used for the execution module to return the execution result to the decision-making module, and the decision-making module performs dynamic learning and optimizes the large language model based on the execution result to ensure that the decision matches the current environmental state.

[0065] Optionally, The current environmental information of the disaster rescue scenario includes one or more of the occurrence of natural disasters, rescue resource situation, rescue personnel situation, traffic conditions, geographical information, and weather and climate information; If the disaster rescue scenario is a storm surge scenario, the environmental information includes one or more of the change of water level depth, traffic conditions, topographic map, population distribution, location of rescue resources, location of shelters, occupancy of shelters, presence of rescue personnel, environmental temperature, and environmental humidity.

[0066] Optionally, The historical decision record is recorded in the message history database and includes an action plan and a decision explanation.

[0067] Optionally, the device further includes: A timing dynamic decision-making module, configured to make decisions for the large model at regular intervals according to a preset decision time interval, and / or determine whether to adjust the decision according to changes in environmental information.

[0068] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, please refer to the partial description of the method embodiment.

[0069] In addition, an embodiment of the present invention further provides an electronic device, as Figure 4 shown, including a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete communication with each other through the communication bus 404. The memory 403 is used to store a computer program; The processor 401, when executing the program stored on the memory 403, implements the intelligent agent simulation decision-making method based on a large language model as described in the above embodiments.

[0070] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0071] The communication interface is used for communication between the above terminal and other devices.

[0072] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0073] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0074] As Figure 5 As shown, in another embodiment provided by the present invention, a computer-readable storage medium 501 is further provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is made to execute the intelligent agent simulation decision-making method based on the large language model described in the above embodiment.

[0075] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, the computer is made to execute the intelligent agent simulation decision-making method based on the large language model described in the above embodiment.

[0076] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0077] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0078] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.

[0079] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. An agent simulation decision-making method based on a large language model, characterized in that: The method comprises: Obtain basic attribute information of the intelligent agent and collect the current environmental information of the disaster relief scene in real time through multiple data sources; Serialize the basic attribute information of the intelligent agent and the current environment information to obtain basic attribute serialization data and environment information serialization data, and pass the basic attribute serialization data and environment information serialization data to the decision module; The basic attribute serialization data and the environmental information serialization data are converted into the basic attribute natural language description and the environmental information natural language description through the large language model of the decision module, and the decision results are generated based on the basic attribute natural language description and the environmental information natural language description, combined with the historical decision records and the behavior library; the decision results include action plans and decision explanations; The decision result returned by the decision module is transmitted to the execution module for execution, and the decision result is recorded in the message history database.

2. The method according to claim 1, characterized in that The behavior library contains standard operating procedures and standard behavior patterns; Standard operating procedures include evacuation route selection, relief material distribution, and search for emergency shelters; The standard behavior pattern includes behavior classification, behavior pattern mapping set for different environmental states, and behavior decision priority rules.

3. The method according to claim 1, characterized in that The method further comprises: The execution module returns the execution results to the decision module, and the decision module dynamically learns and optimizes the large language model based on the execution results to ensure that the decision matches the current environment state.

4. The method according to claim 1, characterized in that: The current environmental information of the disaster rescue scene includes one or more of the occurrence of natural disasters, rescue resources, rescue personnel, traffic conditions, geographic information, and weather and climate information; If the disaster relief scenario is a storm surge scenario, the environmental information includes one or more of water level depth changes, traffic conditions, topographic maps, crowd distribution, rescue resource locations, shelter locations, shelter occupancy, whether there are rescue personnel, ambient temperature, and ambient humidity.

5. The method according to claim 1, characterized in that Historical decision records are recorded in the message history database, including action plans and decision explanations.

6. The method according to claim 1, characterized in that The method further comprises: The large model makes decisions regularly according to a preset decision time interval, and / or determines whether to adjust the decision according to changes in environmental information.

7. An intelligent agent simulation decision-making device based on a large language model, characterized in that: The device comprises: The real-time information collection module is used to obtain the basic attribute information of the intelligent agent and collect the current environmental information of the disaster relief scene in real time through multiple data sources; A serialization processing module is used to serialize the basic attribute information of the intelligent agent and the current environment information to obtain basic attribute serialization data and environment information serialization data, and pass the basic attribute serialization data and environment information serialization data to the decision module; A decision result generation module is used to convert the basic attribute serialization data and the environmental information serialization data into the basic attribute natural language description and the environmental information natural language description through the large language model of the decision module, and generate a decision result based on the basic attribute natural language description and the environmental information natural language description, combined with the historical decision records and the behavior library; the decision result includes an action plan and a decision explanation; The decision result transmission and recording module is used to transmit the decision result returned by the decision module to the execution module for execution, and record the decision result in the message history record database.

8. The device according to claim 7, characterized in that The behavior library contains standard operating procedures and standard behavior patterns; Standard operating procedures include evacuation route selection, relief material distribution, and search for emergency shelters; The standard behavior pattern includes behavior classification, behavior pattern mapping set for different environmental states, and behavior decision priority rules.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the intelligent agent simulation decision-making method based on a large language model as described in any one of claims 1 to 6 when executing the program stored in the memory.

10. One or more computer-readable media having instructions stored thereon, which, when executed by one or more processors, enable the processors to execute the intelligent agent simulation decision-making method based on a large language model as described in any one of claims 1-6.

Citation Information

Patent Citations

  • City emergency evacuation simulation system based on multi intelligent agent

    CN101515309A

  • Coal mine accident simulating method and system based on multi-intelligent agent

    CN102508995A

  • Emergency decision generation method, device and equipment based on deep reinforcement learning

    CN116029389A

  • Intelligent agent interactive emergency decision-making method for flood disaster in urban complex geographic scene

    CN116128322A

  • Multi-agent-based disaster rescue process monitoring method and system

    CN117892973A

Cited By

  • Intelligent agent travel scheme generation and dynamic updating method based on large language model

    CN121279563A

  • Accident rescue method, device and equipment

    CN122266174A