Game-oriented multi-agent decision model architecture method and device, equipment and medium
Through the game-oriented multi-agent decision model architecture method, combined with knowledge rules and data learning, an agent decision-making framework in complex environments is built, solving the problems of decision-making complexity and high computing resource requirements in the existing methods, and achieving efficient confrontational game decision-making.
Patent Information
- Application Number
- CN202411976436.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
The existing game agent decision-making methods have problems such as large engineering volume, complex rule design, high computing resource requirements, excessive state space during training and slow convergence speed in complex environments and dynamic changes.
Using a game-oriented multi-agent decision model architecture method, a decision-making framework for mixed knowledge rules and data learning is constructed by combining hierarchical conclusion ideas with inclusive structures. The framework includes knowledge rules and strategies to guide the training direction of reinforcement learning algorithms, design task decisions, and construct the micro-operation movements of the agent through multi-agent proximity strategy optimization algorithms. Finally, the architecture of the multi-intelligent decision model is verified through game confrontation.
It realizes efficient response to confrontational games in complex wide-area environments, improves decision-making ability and game level, and solves problems such as complex rule design and high computing resource requirements in existing methods.
Smart Images

Figure CN120044981A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned cluster intelligent decision-making, and in particular, relates to a method and device, equipment, and medium for a multi-agent decision-making model architecture for games. Background Art
[0002] In the field of game agents, the methods for constructing agents can be mainly divided into two categories: knowledge-rule-based and data-learning-based. The knowledge-rule-based method relies on expert knowledge to manually construct decision rules and action instructions, enabling the agent to exhibit corresponding behaviors in different situations. This method can directly reflect the knowledge system of the agent through formal logical rules, with strong interpretability and a certain degree of reliability. The data-learning-based method continuously optimizes decision-making strategies through techniques such as reinforcement learning by continuously interacting with the environment. This method can automatically learn the optimal strategy from data and can better handle complex environments and ever-changing adversarial scenarios.
[0003] However, the knowledge-rule-based method faces problems such as huge engineering quantities, complex rule design, and difficulty in coping with dynamically changing environments. At the same time, the portability of the rules is poor. Once the environment changes, the re-design and debugging of the rules are cumbersome and time-consuming. On the other hand, the data-learning-based method also faces challenges such as high computational resource requirements, excessive state spaces during the training process, and slow convergence speeds, especially when facing complex decision-making problems, often requiring a large amount of computation and time.
[0004] Therefore, one or more methods are needed to solve the above problems.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] Embodiments of the present invention provide a method and device, equipment, and medium for a multi-agent decision-making model architecture for games, thereby at least to some extent overcoming one or more problems caused by the limitations and defects of related technologies.
[0007] According to one aspect of the present disclosure, there is provided a method for a multi-agent decision-making model architecture for games, including:
[0008] Constructing a multi-intelligent decision-making framework by combining the hierarchical binding idea with an inclusive structure according to a hybrid system of knowledge rules and data learning to generate a multi-intelligent decision-making framework;
[0009] Designing a task decision by guiding the training direction of the reinforcement learning algorithm in the action according to the operator type through knowledge rule strategies;
[0010] Based on the multi-agent proximal policy optimization algorithm, construct the micro-operations of the agent through data learning, and design action decisions;
[0011] Verify the multi-intelligent decision-making framework introducing the task decision and action decision through game confrontation, and complete the architecture of the multi-agent decision-making model.
[0012] In an exemplary embodiment of the present disclosure, the multi-intelligent decision-making framework includes a battlefield environment, a perception system, a decision-making agent system, a weapon control system, and an operator.
[0013] In an exemplary embodiment of the present disclosure, construct the multi-intelligent decision-making framework by combining the hierarchical conclusion idea with the inclusive structure, and further include:
[0014] Based on the human-machine user interface, the operator perceives the battlefield environment, our operators, and enemy operators through the perception system, and generates situation judgment and situation information;
[0015] Based on the human-machine user interface, the operator makes decisions on the situation judgment and situation information to generate a behavioral strategy decision result;
[0016] Based on the human-machine user interface, the operator analyzes and executes the behavioral strategy decision result to generate an action execution sequence;
[0017] Construct the multi-intelligent decision-making framework based on the situation judgment and situation information, behavioral strategy decision result, and action execution sequence.
[0018] In an exemplary embodiment of the present disclosure, guide the training direction of the reinforcement learning algorithm in actions through knowledge rule strategies, including:
[0019] When the enemy operator is not visible in the situation judgment and situation information, set the training direction according to the knowledge rules to generate a search and detection task strategy;
[0020] When the enemy operator is visible in the situation judgment and situation information, set the training direction according to the knowledge rules to generate a rule assignment task strategy;
[0021] According to the operator platform type, allocate the rules of the rule assignment task strategy to generate the task decision rules for the unmanned aerial vehicle cluster, the unmanned surface vehicle cluster, and the unmanned underwater vehicle cluster.
[0022] In an exemplary embodiment of the present disclosure, construct the micro-operations of the agent through data learning, including:
[0023] Based on the agent type, construct the state space of the agent through data learning to generate a multi-agent state space;
[0024] Based on the agent unit micro-operation interface type, construct the action space of the agent through data learning to generate a multi-agent action space;
[0025] Based on the situation judgment situation information, design the dense report through data learning to generate a reward function.
[0026] In an exemplary embodiment of the present disclosure, verify the multi-agent decision-making framework introducing the task decision and action decision through game confrontation, including:
[0027] Construct a model by introducing the task decision and action decision into the decision agent system to generate a multi-agent decision model;
[0028] Adopt a red-blue game model, and perform game training by setting the expert opponent strategy to generate a win-loss result;
[0029] Generate the winning rate of the multi-agent decision model by counting the win-loss result, and evaluate and verify the multi-agent decision model according to the winning rate of the multi-agent decision model.
[0030] In an exemplary embodiment of the present disclosure, verify the multi-agent decision-making framework introducing the task decision and action decision through game confrontation, including:
[0031] When the evaluation result of the multi-agent decision model does not meet the preset multi-agent decision-making ability, generate an iterative decision-making system through iterative training of the decision agent system, and complete the architecture of the multi-agent decision model based on the iterative decision-making system;
[0032] When the evaluation result of the multi-agent decision model meets the preset multi-agent decision-making ability, complete the architecture of the multi-agent decision model.
[0033] In one aspect of the present disclosure, provide a multi-agent decision model architecture device for games, including:
[0034] A decision framework construction module, which is used to construct a multi-agent decision framework by combining the hierarchical conclusion idea and the inclusive structure according to the hybrid system of knowledge rules and data learning;
[0035] A task decision design module, which is used to design the task decision by guiding the training direction of the reinforcement learning algorithm in the action through the knowledge rule strategy;
[0036] An action decision-making design module for constructing the micro-operations of an agent through data learning and designing action decisions;
[0037] A verification and optimization module for verifying the multi-agent decision-making framework introducing the task decision and action decision through game confrontation and completing the architecture of the multi-agent decision-making model.
[0038] In one aspect of the present disclosure, there is provided an electronic device, including:
[0039] A processor; and
[0040] A memory storing computer-readable instructions thereon, and when the computer-readable instructions are executed by the processor, the method according to any one of the above is implemented.
[0041] In one aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program thereon, and when the computer program is executed by a processor, the method according to any one of the above is implemented.
[0042] The beneficial effects brought by the present disclosure are as follows:
[0043] As can be seen from the above solution, the embodiment of the present invention provides a method for the architecture of a game-oriented multi-agent decision-making model. First, according to the hybrid system of knowledge rules and data learning, a multi-agent decision-making framework is constructed by combining the hierarchical binding idea with the inclusive structure to generate a multi-agent decision-making framework. Then, according to the operator type, the training direction of the reinforcement learning algorithm in the action is guided by the knowledge rule strategy to design the task decision. At the same time, based on the multi-agent proximal policy optimization algorithm, the micro-operations of the agent are constructed through data learning to design the action decision. Finally, by introducing the task decision and action decision into the multi-agent decision-making framework, a multi-agent model is constructed, and the model is verified through game confrontation to complete the architecture of the multi-agent decision-making software. Thus, the embodiment of the present disclosure realizes the construction of multi-agents driven by the hybrid of knowledge rules and data learning, and can ensure that in a complex wide-area environment, the agent can efficiently cope with adversarial games. And by using logical rules to solve the macro policy decision and iterative learning to improve the micro action decision, the decision-making ability and game level of the agent are comprehensively improved.
[0044] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.
[0045] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0046] Figure 1Flow chart of a game-oriented multi-agent decision-making model architecture method according to an embodiment of the present disclosure method;
[0047] Figure 2 Business framework flow chart of a game-oriented multi-agent decision-making model architecture method according to an embodiment of the present disclosure method;
[0048] Figure 3 Software-level decision logic flow chart of a game-oriented multi-agent decision-making model architecture method according to an embodiment of the present disclosure method;
[0049] Figure 4 Structural block diagram of a game-oriented multi-agent decision-making model architecture device according to an embodiment of the present disclosure method;
[0050] Figure 5 Block diagram of an electronic device according to an embodiment of the present disclosure method. Detailed implementation manners
[0051] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0052] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, materials, devices, steps, etc. may be adopted. In other cases, well-known structures, methods, devices, implementations, materials or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0053] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more software-hardened modules, or in different networks and / or processor devices and / or microcontroller devices.
[0054] In the embodiments of the present disclosure, first, a game-oriented multi-agent decision-making model architecture method is provided; referring to Figure 1 as shown in, the game-oriented multi-agent decision-making model architecture method may include the following steps:
[0055] Step S110, based on the hybrid system of knowledge rules and data learning, a multi-intelligence decision-making framework is constructed by combining the hierarchical conclusion idea with the inclusive structure to generate a multi-intelligence decision-making framework.
[0056] Step S120, according to the operator type, guide the training direction of the reinforcement learning algorithm in action through the knowledge rule strategy to design the task decision.
[0057] Step S130, based on the multi-agent proximal strategy optimization algorithm, construct the agent's micro-manipulation actions through data learning and design action decisions.
[0058] Step S140, verifying the multi-intelligent decision-making framework that introduces the task decision and action decision through game confrontation, and completing the architecture of the multi-agent decision-making model.
[0059] Below, as Figures 2 - 3 As shown, a game-oriented multi-agent decision-making model architecture method in an embodiment of the present disclosure will be further explained.
[0060] In step S110, a multi-intelligence decision-making framework can be generated by constructing a framework for multi-intelligence decision-making based on a hybrid system of knowledge rules and data learning by combining the hierarchical conclusion idea with the inclusive structure.
[0061] In some optional embodiments of this example, based on the hybrid system architecture design concept, combined with hierarchical decomposition and inclusive structure, a business framework of a multi-intelligent decision-making system is constructed. This multi-intelligent decision-making framework is developed around the task and mainly includes five major components: battlefield environment, perception system, decision-making agent system, weapon control system and operator. In this framework system, information interaction between the decision-making agent and the operator is realized through the human-machine user interface to ensure that the operator can grasp the decision-making process in real time.
[0062] Among them, the decision-making intelligent agent system is the core of the entire framework, which consists of three modules: situation judgment, behavior strategy decision, and action execution and control. The situation judgment module analyzes the battlefield situation through perception information, the behavior strategy decision module generates the best action strategy based on environmental information, and the action execution and control module converts the decision results into actual operation actions and feeds them back to the battlefield environment.
[0063] In some optional embodiments of this example, in order to cope with the complexity of the game scenario, a software system of a multi-intelligent decision-making framework for game is designed. The system uses a human-machine user interface and a hybrid decision-making mechanism to optimize the decision-making of different game scenarios from the above three aspects.
[0064] In the perception stage, the agent collects all the situation information such as the entire environment, our operators, and relevant party operators from the game platform. The operator makes a situation judgment on this situation information through the human-machine user interface, generating situation judgment situation information. This situation judgment situation information can provide sufficient auxiliary information for key decisions in the subsequent decision-making stage.
[0065] In the decision-making stage, based on the situation judgment situation information collected by the agent in the perception stage, the operator makes decision control using limited resource conditions, generating the decision result of the agent confrontation behavior strategy with the highest game confrontation winning rate. This strategy is an important manifestation of verifying the agent's cognitive level. This also makes the behavior strategy decision the core of the game agent.
[0066] In the control stage, the agent accepts the decision result of the behavior strategy generated in the decision-making stage, parses it into an action execution sequence executable by the decision environment, and returns it to the environment for operation to complete the action execution of the agent.
[0067] In this way, through the modeling of the business framework and the architecture design of the software system, the construction of the multi-agent decision-making framework can be completed.
[0068] In step S120, according to the operator type, the training direction of the reinforcement learning algorithm in the action guided by the knowledge rule strategy is designed for task decision-making.
[0069] In the game task decision-making of some optional embodiments of this example, first, a corresponding knowledge rule system is designed for different task requirements. The task decision-making strategy of the game agent will be adjusted according to whether the enemy operator is visible. If the enemy operator is not visible, the task focuses on search and detection, and a search and detection task strategy is constructed. If the enemy operator is visible, the agent will design a more complex rule assignment task strategy according to the battle situation and select the most suitable tactical strategy for the current situation.
[0070] After that, according to different types of agent platforms, the task decision-making rules are also different. For unmanned aerial vehicle clusters, unmanned surface vehicle clusters, and unmanned underwater vehicle clusters, etc., refined task decision-making rules are formulated respectively. Each platform adopts a different rule system for task planning and allocation according to its own characteristics. For example, the task decision-making rules of unmanned aerial vehicle clusters focus more on air superiority task decision-making; the task decision-making rules of unmanned surface vehicle clusters and unmanned underwater vehicle clusters focus more on task execution in the underwater environment.
[0071] Meanwhile, to improve the decision-making level of the agent, a rule-guided system based on the reinforcement learning algorithm is designed. By setting the rule parameters, the training direction of the reinforcement learning model can be effectively guided. This combination not only optimizes the decision-making process of the game agent but also improves the game level of the game strategy. Through rule-driven, the system can effectively prevent the agent from falling into a local optimal solution in a complex environment and further improve the decision-making quality.
[0072] In step S130, based on the multi-agent proximal policy optimization algorithm, the micro-operations of the agent are constructed through data learning, and the action decision is designed.
[0073] In some alternative embodiments of this example, due to the complexity and hierarchy of the game scenario, the micro-operations of the agent mainly used for lower-level action decisions are constructed based on the multi-agent proximal policy optimization algorithm (Multi-Agent Proximal Policy Optimization, abbreviated as MAPPO).
[0074] First, for different types of agent platforms, such as unmanned aerial vehicles, unmanned surface vessels, and unmanned underwater vehicles, different multi-agent state spaces are used for modeling. Unmanned aerial vehicles are mainly modeled based on their flight attitude, speed, and position, while unmanned surface vessels and unmanned underwater vehicles focus more on the motion state on the water surface or underwater. Through reasonable state space modeling, the behavior decisions of the agent in different scenarios can be ensured to be efficient and accurate.
[0075] Second, appropriate operation actions are defined according to different types of agents to model the multi-agent action space. The action space of unmanned aerial vehicles includes maneuvering actions, evading actions, speed adjustment actions, etc., and the action space for unmanned surface vessels and unmanned underwater vehicles is adjusted accordingly. By setting heterogeneous action spaces, sufficient flexibility and response capabilities can be guaranteed for each agent when performing tasks.
[0076] Third, to guide the training of the reinforcement learning model, two types of reward functions are designed - the battle damage reward function and the distance reward function. These reward functions are evaluated according to the performance of the agent in the game process and provide feedback for each agent. The battle damage reward function focuses on the attack effect of the agent on the enemy in the task, while the distance reward function is related to the change in the distance between the agent and the target. The design of the reward function not only enhances the adaptability of the agent but also improves the execution effect of the overall strategy.
[0077] In step S140, the multi-intelligent decision framework introducing the task decision and the action decision is verified through game confrontation, and the architecture of the multi-agent decision model is completed.
[0078] In some alternative embodiments of this example, first, it is necessary to introduce task decisions driven by knowledge rules and action decisions driven by data learning into the core of the multi-intelligent decision-making framework (i.e., the behavior strategy decision-making module) to construct a multi-agent decision-making model.
[0079] After that, a red-blue game model is used for game confrontation to verify the multi-agent decision-making model.
[0080] The first step is to construct different opponent strategies. In a specific example, 3 different types of expert opponent strategies are used to train the multi-agent decision-making model to ensure the effectiveness of training. The opponent strategies are respectively a herringbone attack, a three-team attack on the left, middle, and right, and a seven-team attack.
[0081] The second step is to verify the game ability of the agent by statistically calculating the winning rate of the multi-agent decision-making model based on the win-loss results. In the test results of a specific example, the agent maintains a 100% winning rate when facing a herringbone attack, and the winning rate still exceeds 80% when facing complex strategies such as a seven-team attack.
[0082] The third step is to collect the winning rate of the game confrontation and other feedback data for in-depth evaluation to generate the evaluation result of the multi-agent decision-making model. It should be noted that this evaluation result not only focuses on the winning rate but also includes multiple dimensions such as response time, resource consumption, and strategy execution efficiency.
[0083] The fourth step is to preset the benchmark value of the multi-agent decision-making ability. When the evaluation result does not meet this benchmark value, through iterative training, the strategy decision rule is improved to generate an iterative decision-making system, and the adaptability of the model in complex scenarios is enhanced according to the iterative decision-making system.
[0084] The fifth step is to further improve the performance of the multi-agent decision-making model in different battlefield environments by combining algorithm tuning and strategy optimization.
[0085] The sixth step is to stop the iterative training and algorithm optimization when the evaluation result meets the above benchmark value. Output the optimized multi-agent decision-making model. At this time, this model can effectively cope with diverse confrontation situations and has high strategic flexibility and adaptability.
[0086] It should be noted that although the steps of the method in this disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0087] In addition, in the embodiments of this example, a multi-agent decision-making model architecture device for games is also provided. Refer toFigure 3 As shown, the multi-agent decision-making model architecture device 300 for games may include: a decision framework construction module 310, a task decision design module 320, an action decision design module 330, and a verification and optimization module 340. Among them:
[0088] The decision framework construction module 310 is used to construct a multi-agent decision framework by combining the hierarchical conclusion idea and the inclusive structure according to the hybrid system of knowledge rules and data learning, and generate a multi-agent decision framework;
[0089] The task decision design module 320 is used to design task decisions by guiding the training direction of the reinforcement learning algorithm in actions through knowledge rule strategies;
[0090] The action decision design module 330 is used to design action decisions by constructing the micro-operations of the agent through data learning;
[0091] The verification and optimization module 340 is used to verify the multi-agent decision framework introducing the task decision and the action decision through game confrontation, and complete the architecture of the multi-agent decision model.
[0092] The multi-agent decision-making software architecture device for games in the embodiments of the present disclosure corresponds to the embodiments of the above-mentioned multi-agent decision-making software architecture method of the present disclosure, and the relevant content can be referred to each other, which will not be elaborated here. The corresponding beneficial technical effects of the multi-agent decision-making software architecture device for games in the embodiments of the present disclosure can be seen in the corresponding beneficial technical effects of the above-mentioned corresponding exemplary method part, which will not be elaborated here.
[0093] It should be noted that although several modules or units of the multi-agent decision-making software architecture device 300 for games are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0094] Next, refer to Figure 5 to describe the electronic device according to the embodiments of the present disclosure. The electronic device may be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device can communicate with the first device and the second device to receive the input signals collected from them.
[0095] Figure 5 The block diagram of the electronic device according to the embodiments of the present disclosure is shown.
[0096] As Figure 5As shown, the electronic device includes one or more processors and a memory.
[0097] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.
[0098] The memory can store one or more computer program products. The memory can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products can be stored on the computer-readable storage media, and the processor can run the computer program products to implement the methods of the various embodiments of the present disclosure described above and / or other desired functions.
[0099] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0100] In addition, the input device may further include, for example, a keyboard, a mouse, and so on.
[0101] The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0102] Of course, for simplicity, Figure 5 only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0103] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the above part of this specification.
[0104] The computer program product may be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0105] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the foregoing part of this specification.
[0106] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0107] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A game-oriented multi-agent decision model architecture method, characterized in that: include: Based on the hybrid system of knowledge rules and data learning, a multi-intelligence decision-making framework is constructed by combining the hierarchical conclusion idea with the inclusive structure to generate a multi-intelligence decision-making framework; According to the operator type, the training direction of the reinforcement learning algorithm in action is guided by knowledge rule strategy to design task decision; Based on the multi-agent proximal strategy optimization algorithm, the micro-manipulation actions of the agent are constructed through data learning and action decisions are designed; The multi-intelligent decision-making framework that introduces the task decision and action decision is verified through game confrontation, and the architecture of the multi-agent decision-making model is completed.
2. The method according to claim 1, characterized in that The multi-intelligent decision-making framework includes battlefield environment, perception system, decision-making intelligent agent system, weapon control system, and operator.
3. The method according to claim 2, characterized in that The framework of multi-intelligence decision-making is constructed by combining the hierarchical conclusion idea with the inclusive structure, which also includes: Based on the human-machine user interface, the operator perceives the battlefield environment, friendly operators, and enemy operators through the perception system to generate situation judgment information; Based on the human-machine user interface, the operator makes a decision by judging the situation information and generating a behavior strategy decision result; Based on the human-machine user interface, the operator generates an action execution sequence by analyzing and executing the behavior strategy decision results; Based on the situation judgment information, behavior strategy decision results, and action execution sequence, the multi-intelligent decision framework is constructed.
4. The method according to claim 3, characterized in that Guide the training direction of reinforcement learning algorithms in action through knowledge rule strategies, including: When the enemy operator is not visible in the situation judgment information, the training direction is set according to the knowledge rule to generate a search and detection task strategy; When the enemy operator is visible in the situation judgment information, the training direction is set according to the knowledge rule, and the rule allocation task strategy is generated; According to the operator platform type, by allocating the rules of the rule allocation task strategy, the UAV swarm task decision rules, the unmanned boat swarm task decision rules, and the unmanned submarine swarm task decision rules are generated.
5. The method according to claim 3, characterized in that: The micro-manipulation actions of the intelligent agent are constructed through data learning, including: Based on the agent type, the state space of the agent is constructed through data learning to generate a multi-agent state space; Based on the micro-manipulation interface type of the agent unit, the action space of the agent is constructed through data learning to generate a multi-agent action space; Based on the situation information, the dense report is designed through data learning to generate a reward function.
6. The method according to claim 1, characterized in that The multi-intelligent decision-making framework introducing the task decision and action decision is verified through game confrontation, including: The model is constructed by introducing the task decision and action decision into the decision agent system to generate a multi-agent decision model; Using the red-blue game model, we set up expert opponent strategies for game training and generate winning and losing results; By collecting statistics on the winning and losing results, the winning rate of the multi-agent decision-making model is generated, and the multi-agent decision-making model is evaluated and verified according to the winning rate of the multi-agent decision-making model.
7. The method according to claim 6, characterized in that The multi-intelligent decision-making framework introducing the task decision and action decision is verified through game confrontation, including: When the evaluation result of the multi-agent decision model does not meet the preset multi-agent decision-making capability, iteratively training the decision agent system to generate an iterative decision system, and completing the architecture of the multi-agent decision model based on the iterative decision system; When the evaluation result of the multi-agent decision-making model meets the preset multi-agent decision-making capability, the architecture of the multi-agent decision-making model is completed.
8. A multi-agent decision model architecture device for game-based games, characterized in that: include: The decision framework building module is used to build a framework for multi-intelligence decision-making based on a hybrid system of knowledge rules and data learning by combining the hierarchical conclusion idea with the inclusive structure to generate a multi-intelligence decision-making framework; The task decision design module is used to guide the training direction of the reinforcement learning algorithm in action through knowledge rule strategies and design task decisions; Action decision design module, which is used to construct micro-manipulation actions of intelligent agents through data learning and design action decisions; The verification and optimization module is used to verify the multi-intelligent decision-making framework that introduces the task decision and action decision through game confrontation, and complete the architecture of the multi-agent decision-making model.
9. An electronic device, characterized in that: include: a memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, it implements the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Complex event analysis processing method and device for multi-agent cooperative game
CN120509413A
Public facility reachability evaluation method and system based on geographical multi-agent
CN121094346A