Multi-agent game playing method and system based on large language model

By combining large language models and multi-agent systems to build agents in complex scenarios, the existing systems have solved the problems of low efficiency, insufficient accuracy and limited independent optimization capabilities in decision-making solutions, and efficient and accurate decision-making solutions generation and automated optimization are achieved.

CN120124748APending Publication Date: 2025-06-10INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510220412.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing multi-agent system has low efficiency in generation of decision-making solutions in complex scenarios, insufficient accuracy, limited autonomous optimization capabilities, and there are limitations on agent scalability and resource scheduling when handling a large number of parallel tasks.

Method used

By combining large language models and multi-agent systems, environmental agents, director agents and participant agents are built, and the natural language understanding and reasoning capabilities of the large language model are used to improve the system's analysis and processing capabilities of complex scenarios, and efficient solution generation and automation optimization are achieved.

Benefits of technology

It significantly improves the efficiency and accuracy of decision-making solutions in complex scenarios, enhances the system's independent optimization capabilities, enables it to continuously improve the solutions in a dynamic environment, and solves the problems of agent scalability and resource scheduling limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124748A_ABST
    Figure CN120124748A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent game method and system based on a large language model, and belongs to the field of artificial intelligence. According to the invention, a highly simulated game environment is formed by introducing multiple types of agents such as an environment agent, a director agent, a participant agent and a scheme recording agent. Each agent performs independent fine adjustment and optimization for a specific task, so that the agent has the capabilities of information analysis, strategy formulation and execution, real-time feedback and self-lifting, and the modularization design ensures the collaboration and flexibility of the system. The big language model provides comprehensive data processing and analysis support for the intelligent agent, and ensures that a decision scheme is kept scientific and reasonable in the formulation and optimization process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a multi-agent game method and system based on a large language model. Background Art

[0002] In the modern decision-making field, scenario generation is a challenging task in a complex environment, and its success or failure directly affects the decision-making result. Traditional scenario generation relies on the personal experience of decision-makers and expert discussions, combined with manual simulation, but it is inefficient, subjective, and difficult to cope with environmental variability and uncertainty.

[0003] With the continuous progress of artificial intelligence technology, big data, and cloud computing, decision support systems are developing towards intelligence and automation. As a distributed intelligent system, a multi-agent system can handle complex decision-making problems and achieve self-adaptation and optimization in a changing environment through the interaction and cooperation of multiple agents. However, the current multi-agent system still faces challenges in scenario generation.

[0004] ● Lack of powerful natural language processing capabilities: Existing systems are difficult to effectively process and understand complex natural language descriptions, limiting their application in scenarios that require in-depth analysis of text information.

[0005] ● Insufficient complex scenario understanding capabilities: Existing systems have difficulties in understanding various factors and their interrelationships in complex scenarios, resulting in generated scenarios that may be inaccurate or unrealistic.

[0006] ● Difficulty in quickly generating high-quality scenarios: In a complex and dynamic environment, existing systems often have difficulty generating scientific and reasonable decision-making scenarios in a short time.

[0007] ● Limited autonomous optimization capabilities: Existing systems usually rely on external feedback or a large amount of simulation data for optimization, lacking the ability of autonomous learning and adjustment, which limits their adaptability in a rapidly changing environment.

[0008] ● Limitations in agent scalability and resource scheduling: Traditional multi-agent systems are often limited by agent scalability and resource scheduling when dealing with a large number of parallel tasks.

[0009] The present invention effectively solves the defects of the above traditional systems by combining the advantages of large language models and multi-agents. The powerful natural language understanding and reasoning capabilities of large language models improve the system's analysis and processing level of complex scenarios. The high degree of cooperation of agents makes scenario generation faster and more accurate. The feedback mechanism independent of the external simulation platform further enhances the system's autonomous optimization capabilities, enabling it to continuously improve scenarios in a dynamic environment. Summary of the Invention

[0010] The present invention proposes a multi-agent game method and system based on large language models, which realizes efficient solution generation and automated optimization in dynamic environments by enhancing real-time data processing and intelligent feedback mechanisms, and is applicable to complex decision-making requirements in multiple fields.

[0011] To achieve the above object, the technical solution of the present invention includes the following content.

[0012] A multi-agent game method based on large language models, the method includes:

[0013] Constructing agents using large language models; wherein, the agents include: an environment agent, a director agent, and several participant agents;

[0014] The environment agent generates the tasks of each participant agent and the setting conditions of the director agent according to the environmental situation, strategic goals, and the advantages and disadvantages of each participant in the game scenario;

[0015] Each participant agent generates an action plan for the participant agent according to the environmental situation and the task;

[0016] The director agent determines whether to adjust the deduction process according to the deduction stage, setting conditions, and the action plans of each participant, and sends a deduction process adjustment instruction to each participant agent when the deduction process needs to be adjusted;

[0017] Each participant agent adjusts the environmental situation according to the deduction process adjustment instruction, and generates a new action plan for the participant agent based on the adjusted environmental situation and the task.

[0018] Further, the constructing agents using large language models includes:

[0019] Selecting a pre-trained large language model according to task requirements and computing resources, the large language model includes: GPT series, LLaMA series, or other open-source / commercial models;

[0020] Fine-tuning the pre-trained large speech model based on a dataset in a specific field; wherein, the dataset contains input-output pairs related to the roles and tasks of the agents;

[0021] Evaluating the metrics of the fine-tuned large speech model; wherein, the metrics include: BLEU, ROUGE, and / or perplexity.

[0022] Further, the game scenario includes: commercial game scenario, military game scenario, policy game scenario, or crisis handling game scenario.

[0023] Further, in the case where the game scenario is a military game scenario, the task of generating each participant agent includes:

[0024] Embed the environmental situation, strategic goals, and the advantages and disadvantages of a participant into the task objective generation prompt template, and obtain the task of this participant based on the large language model used by the environmental agent; wherein, the environmental situation includes: terrain features, weather conditions, the force deployments of both sides, and resource situations.

[0025] Further, in the case where the game scenario is a military game scenario, the setting conditions for generating the director agent include:

[0026] Embed the environmental situation and strategic goals into the setting condition generation prompt template, and obtain the setting conditions for the director agent based on the large language model used by the environmental agent; wherein, the setting conditions include: initial conditions, deduction rules, possible events, and evaluation indicators, and the initial conditions include: the forces, resources, and terrain of each participant.

[0027] Further, in the case where the game scenario is a military game scenario, each participant agent generates the action plan of this participant agent according to the environmental situation and the task, including:

[0028] Use the large language model used by the participant agent to parse the task to obtain the task objective, battlefield environmental conditions, and time limit;

[0029] Construct a query statement according to the task objective, the battlefield environmental conditions, and the time limit, and perform a retrieval in the knowledge base based on the query statement to obtain a retrieval result; wherein, the retrieval result includes: historical battle cases, tactical plans, and strategy fragments corresponding to the task;

[0030] Extract the action plan cases of each participant from the retrieval result;

[0031] Embed the task, the environmental situation, and the action plan cases of each participant into the action plan generation template, and generate an action plan based on the large speech model used by this participant agent.

[0032] Further, in the case where the game scenario is a military game scenario, the director agent determines whether to adjust the deduction process according to the deduction stage, the setting conditions, and the action plans of each participant, and in the case where the deduction process needs to be adjusted, sends a deduction process adjustment instruction to each participant agent, including:

[0033] Embed the deduction stage, set conditions, and action plans of each participant into the deduction process adjustment instruction generation prompt template, and obtain the deduction process adjustment instruction based on the large language model used by the director agent; wherein, the deduction process adjustment instruction includes: triggering a new environmental situation, adjusting deduction rules, and sending instructions to the participants to adjust the action plan and / or provide intelligence support.

[0034] Furthermore, the method further includes:

[0035] Construct a plan recording agent using the large language model;

[0036] Embed the deduction stage, the action plans of each participant agent, and the adjustment information of the environmental situation into the plan recording prompt template, and obtain a feedback report based on the large language model used by the plan recording agent; wherein, the feedback report includes: evaluation of the execution effect of the existing strategy, analysis of the rationality of the troop deployment, analysis of the changes in the battlefield situation, assessment of potential risks and threats, and optimization suggestions;

[0037] Provide the feedback report to each participant agent so that each participant agent can adjust the action plan in combination with the feedback report.

[0038] A multi-agent game system based on a large language model, the system includes: an environmental agent, a director agent, and several participant agents constructed using the large language model; wherein,

[0039] The environmental agent is used to generate the tasks of each participant agent and the set conditions of the director agent according to the environmental situation, strategic goals, and the advantages and disadvantages of each participant in the game scenario;

[0040] The participant agent is used to generate the action plan of the participant agent according to the environmental situation and the task;

[0041] The director agent is used to judge whether to adjust the deduction process according to the deduction stage, set conditions, and action plans of each participant, and in the case where the deduction process needs to be adjusted, send the deduction process adjustment instruction to each participant agent, so that each participant agent can adjust the environmental situation according to the deduction process adjustment instruction, and generate a new action plan of the participant agent based on the adjusted environmental situation and the task.

[0042] Furthermore, a plan recording agent constructed using the large language model, the plan recording agent is used for:

[0043] Embed the deduction stage, the action plans of each participating agent, and the adjustment information of the environmental situation into the plan record prompt template, and obtain a feedback report based on the large language model used by the plan record agent; wherein, the feedback report includes: evaluation of the execution effect of the existing strategy, analysis of the rationality of force deployment, analysis of changes in the battlefield situation, assessment of potential risks and threats, and optimization suggestions.

[0044] Provide the feedback report to each participating agent so that each participating agent can adjust the action plan in combination with the feedback report.

[0045] Compared with the prior art, the present invention has at least the following beneficial effects.

[0046] (1) Improve the efficiency of generating decision-making plans.

[0047] Compared with the traditional method, the present invention significantly improves the efficiency of generating decision-making plans in complex scenarios through the cooperation of the large language model and the multi-agent system.

[0048] (2) Enhance the accuracy of decision-making plans.

[0049] By introducing the large language model, the present invention endows the multi-agent system with stronger environmental understanding and prediction capabilities. The agent can handle variable factors and analyze complex scenario elements in real time, including external environmental changes, resource conditions and constraints. Compared with the prior art, the present invention reduces the dependence on a large amount of simulation data through the semantic understanding and context reasoning capabilities of the large language model, improves the accuracy and adaptability of the decision-making plan in a dynamic environment. The self-regulating mechanism of the present invention refers to the continuous evaluation and feedback of the plan execution effect by the large language model, so that the agent can autonomously adjust the strategy parameters to achieve the dynamic optimization of the plan.

[0050] (3) Realize the automatic optimization of decision-making plans.

[0051] The present invention utilizes the reasoning and learning capabilities of the large language model to realize the automatic optimization of decision-making plans. The agent autonomously adjusts the strategy according to the environmental changes and feedback information without manual intervention. Compared with the traditional optimization method that relies on repeated training, the present invention improves the optimization efficiency and adaptive ability. The present invention integrates the process of plan generation and optimization into a unified framework, and realizes the coordination of generation and optimization through the cooperation and information sharing among agents, improving the forward-looking and adaptive ability of the technical solution.

[0052] (4) Improve the ability to perceive complex scenarios.

[0053] With the semantic understanding and context reasoning capabilities of large language models, the intelligent body can accurately perceive and interpret real-time changes in various complex scenarios. The environmental agent updates the scenario information in real time, including variables and constraints, enabling decision-makers to have a clearer understanding of the current dynamics and possible impacts. Compared with the prior art, the present invention improves the accuracy and efficiency of scenario perception through the fusion analysis of multi-source heterogeneous data by a large language model. The present invention records the feedback of the scenario recording agent to provide decision support for users, further demonstrating its technical advantages and innovation in complex decision-making scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is an overall structure diagram of a multi-agent game system based on a large language model.

[0055] Figure 2 It is the specific process of generating and optimizing a multi-agent game solution based on a large language model.

[0056] Figure 3 It is a schematic diagram of a multi-agent game simulation example. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following is a further detailed description of the present invention in conjunction with the drawings and embodiments. The embodiments given are only for clarifying the present invention and not for limiting the scope of the present invention.

[0058] The core of the present invention lies in the organic combination of large language model (LLM) technology and multi-agent systems, giving full play to the powerful natural language understanding, reasoning capabilities of large language models and their application potential in decision-making generation. The large language models adopted in the present invention include but are not limited to currently leading artificial intelligence models such as the GPT series and the LLaMA series, which can extract information from diverse data and perform effective decision-making reasoning. By embedding the large language model into the multi-agent framework, the accuracy and efficiency of the solution generation process are significantly improved, realizing the automated deduction and optimization of complex decision-making processes.

[0059] As Figure 1 shown, the present invention forms a highly simulated game environment by introducing various types of agents such as environmental agents, director agents, participant agents, and scenario recording agents. Each agent is independently fine-tuned and optimized for specific tasks, enabling it to have information parsing, strategy formulation and execution, real-time feedback, and self-improving capabilities. This modular design ensures the coordination and flexibility of the system. The large language model provides comprehensive data processing and analysis support for the agents, ensuring the scientificity and rationality of the decision-making solutions during formulation and optimization.

[0060] The complete optimization process of the present invention covers all aspects from environment setting, agent interaction, solution generation to optimization feedback, aiming to achieve the automated, intelligent generation and continuous optimization of decision-making solutions in complex scenarios. The technological breakthrough of the present invention provides direct, accurate and efficient solutions for multiple game scenarios such as commercial game scenarios, military game scenarios, policy game scenarios and crisis handling game scenarios, significantly improving the decision-making level and response speed of decision-makers in complex environments.

[0061] Example 1: Military game scenario.

[0062] As Figure 2 and Figure 3 shown, the multi-agent game method based on large language models in the military game scenario of the present invention includes the following steps.

[0063] Step 1.1: Construction of a multi-agent system.

[0064] In the present invention, multiple agent modules are combined with large language models to form an efficient multi-agent system. Each agent corresponds to an instance of a large language model, and through fine-tuning for specific roles and tasks, the agents achieve independent and collaborative decision-making functions. Specific agent types include environmental agents, director agents, participant agents, and solution recording agents, which together form a highly simulated game environment, creating a solid foundation for solution generation and optimization.

[0065] Each agent is specifically fine-tuned by a large language model to conform to its unique operating environment and task content. The selection and fine-tuning strategies of the large language model are as follows:

[0066] ● Model selection: According to specific task requirements and computing resources, select a suitable pre-trained large language model, such as GPT series, LLaMA series or other open-source / commercial models. The selection criteria include the performance, parameter scale, inference speed and fine-tunability of the model.

[0067] ● Fine-tuning method: Adopt the supervised learning method and use a dataset in a specific domain to fine-tune the pre-trained model. The dataset contains input-output pairs related to the agent's role and task. During the fine-tuning process, full-parameter fine-tuning or parameter-efficient fine-tuning (such as LoRA, Adapter, etc.) methods can be adopted.

[0068] ● Evaluation metrics: Use metrics such as BLEU, ROUGE, perplexity, etc. to evaluate the performance of the fine-tuned model. At the same time, combined with manual evaluation, examine whether the text generated by the model conforms to the agent role setting and task requirements.

[0069] Among agents, collaboration is achieved through the enhanced understanding and generation capabilities of large language models, enabling agents with different roles to communicate and exchange information effectively in complex environments:

[0070] 1) Environmental agent: Defines the decision-making scenario and background, including various environmental factors, and guides other agents to identify key variables. Prompt example: "The current battlefield environment is X, the terrain features are Y, the weather conditions are Z, which may affect Strategy A." Environmental agent prompt template example:

[0071] Agent for task participants:

[0072] You are an environmental agent responsible for generating tasks for the [Red / Blue] participant according to the current military situation and strategic objectives.

[0073] Current military situation: [Describe terrain features, weather conditions, force deployments of both sides, resource situation, etc.]

[0074] Strategic objective: [Describe the overall strategic objective, such as occupying the target area, destroying the enemy's key facilities, etc.]

[0075] Advantages of the [Red / Blue] side: [Describe the advantages of this participant, such as technological advantages, force advantages, etc.]

[0076] Disadvantages of the [Red / Blue] side: [Describe the disadvantages of this participant, such as resource shortage, unfavorable terrain, etc.]

[0077] Please generate a specific task for the [Red / Blue] side, including task objectives, time limits, available resources, etc.

[0078] Agent for the director:

[0079] You are an environmental agent responsible for generating the setting conditions for the director's agent to conduct a military exercise according to the current military situation.

[0080] Current military situation: [Describe the battlefield environment, force deployments of both sides, resource situation, key terrain, etc.]

[0081] Strategic objective: [Describe the overall strategic objective, such as occupying the target area, destroying the enemy's key facilities, etc.]

[0082] Please generate the setting conditions for the director's agent to conduct a military exercise, including:

[0083] - Initial conditions: [Forces, resources, terrain, etc. of both sides]

[0084] - Rules of engagement: [Rules of action, rules for determining victory or defeat, etc.]

[0085] - Possible events: [Such as enemy reinforcements, weather changes, equipment failures, etc.]

[0086] - Evaluation metrics: [Metrics used to evaluate the quality of a plan, such as casualty rate, resource consumption, mission completion rate, etc.]

[0087] 2) Director agent: Responsible for the overall control of the game deduction, with the ability to adjust strategies and initiate new events. Hint example: "The deduction has entered the Nth stage, and the red side has achieved the M goal. Should a new battlefield event be triggered? The blue side's troop losses exceed 50%, and the strategy needs to be adjusted." Director agent hint template example:

[0088] You are a director agent, responsible for controlling the process of military strategy deduction.

[0089] Deduction setting conditions: [Setting conditions provided by the environment agent]

[0090] Current deduction stage: [Describe the current stage of the deduction]

[0091] Action plan of [Red / Blue] side: [Action plan generated by the participating agent]

[0092] Please judge whether it is necessary to adjust the deduction process according to the current situation, for example:

[0093] - Trigger new events: [Such as enemy reinforcements, weather changes, equipment failures, etc.]

[0094] - Adjust deduction rules: [Such as modifying action rules, win-loss judgment rules, etc.]

[0095] - Issue instructions to the participating parties: [Such as requiring adjustment of the action plan, providing intelligence support, etc.]

[0096] 3) Participating agent: Utilize the large language model to process historical data and the current situation to provide a basis for its action planning. Hint example: For example, the participating red and blue sides utilize the large language model to process historical battle cases and the current battlefield situation. Hint example: "Your goal is to occupy Hill X. Currently, your troop strength is Y. It is recommended to execute Action Z first." Participating agent hint template example:

[0097] You are a participating agent, representing [Red / Blue] side to participate in the military strategy deduction.

[0098] Task: [Task provided by the environment agent]

[0099] Current military situation: [Describe the battlefield environment, the troop deployments of both sides, resource situation, key terrain, etc.]

[0100] Our action plan: [Describe our current action plan]

[0101] Enemy action plan: [Describe the possible action plan of the enemy]

[0102] Generate or adjust your action plan according to the current situation, including:

[0103] - Objective: [Specific objective, such as capturing a certain high ground, destroying the enemy's radar station, etc.]

[0104] - Force deployment: [How to deploy forces, such as attack, defense, reserve forces, etc.]

[0105] - Resource allocation: [How to allocate resources, such as ammunition, fuel, supplies, etc.]

[0106] - Action steps: [Specific action steps, such as reconnaissance, attack, retreat, etc.]

[0107] - Risk assessment: [Assess the risks of the action plan, such as possible resistance encountered, possible losses caused, etc.]

[0108] These factors are embedded in the prompt template, which can generate different mission objectives for different parties. For each party, the input of the environmental agent includes the common input of the parties (such as the overall battlefield situation, strategic objectives, etc.) and the unique input of each party (such as force deployment, mission objectives, etc.) to ensure the generation of accurate mission objectives for each party.

[0109] 4) Plan recording agent: Track and record the key events that occur during the interaction, providing data support for further analysis and optimization. Prompt example: "Record the time T, the event is X, the impact is Y, and the response of the Red / Blue side is Z." Example of the prompt template of the plan recording agent:

[0110] You are a plan recording agent responsible for recording the key information during the military strategic deduction process.

[0111] Current deduction stage: [Describe the current deduction stage]

[0112] [Red / Blue]'s action plan: [The action plan generated by the party agent]

[0113] Environmental changes: [Describe the changes in the battlefield environment, such as weather changes, terrain changes, etc.]

[0114] Please record the following information:

[0115] - Time: [Record the time when the event occurred]

[0116] - Event: [Describe the event that occurred, such as firefight, reconnaissance, air raid, etc.]

[0117] - Party involved: [The party involved in the event]

[0118] - Action: [The action taken by the party]

[0119] - Outcome: [The outcome of the action, such as the achievement of the goal, the casualty situation, the resource consumption situation, etc.]

[0120] - Impact: [The impact of the event on the battle situation]

[0121] To clarify the above settings, consider a military strategic deduction scenario. Among them, the red side and the blue side represent two opposing military forces respectively. The environmental agent sets the battlefield conditions, such as terrain, weather, the initial troop deployments of both sides, etc.; the director agent controls the deduction process in real time, simulates possible battlefield events, such as enemy reinforcements, air strikes, supply line interruptions, etc.; the participating agent unfolds strategy planning according to the prompts, such as attacking, defending, retreating, adjusting troop deployments, etc.; and the scenario recording agent accurately records every key decision and market reaction for evaluation.

[0122] This example will be processed and applied in more detail in the subsequent steps. From task setting, strategy deduction to final evaluation, by analyzing the interactions between the red side and the blue side in the entire military game, it comprehensively demonstrates the efficiency and innovation of the multi-agent system of the present invention in actual operation.

[0123] Step 1.2: Task setting and assignment.

[0124] In the multi-agent system, the environmental agent undertakes the key responsibilities of initial battlefield environment analysis and task objective setting. With the help of the large language model, the environmental agent can deeply analyze the complex variables in the military game scenario, such as terrain features, weather conditions, the troop deployments of both sides, resource situations, etc., and accordingly set accurate task objectives. For example, the environmental agent may set the task that "the red side needs to occupy High Ground X within the next 24 hours, with the casualty rate controlled below Y%, while ensuring the smoothness of the main supply line". This intelligent task setting process significantly improves the task accuracy and the cooperation efficiency among agents.

[0125] When setting tasks, the environmental agent will generate two types of setting conditions: >

[0126] Task setting for participating agents: Clearly define the specific task objectives, time limits, available resources, etc. for each participating party (such as the red side and the blue side). This is to ensure that each participating party clearly understands its role and objectives in the deduction.

[0127] Deduction setting for the director agent: Provide the initial conditions, deduction rules, possible events, and evaluation indicators for the entire deduction. This is to enable the director agent to comprehensively control the deduction process and make adjustments as needed.

[0128] The environmental agent generates these set conditions by calling a pre-fine-tuned large language model instance and inputting a prompt containing the above key information (the specific prompt template has been detailed in Step 1).

[0129] After receiving the task, the participating agents (such as the Red Side and the Blue Side) use the Retrieval-Augmented Generation (RAG) technique combined with the powerful semantic understanding ability of the large language model to analyze the task requirements. This process includes accurately identifying key elements such as the task objective, battlefield environmental conditions, time limit, etc., and clarifying the types and scope of information to be retrieved, such as classic tactics, force configuration plans, logistics support strategies, etc.

[0130] 1) Task analysis and requirement clarification

[0131] After receiving the task assigned by the environmental agent, the participating agent first uses the large language model to analyze the task, accurately identify the task objective, battlefield environmental conditions, time limit and clarify the retrieval requirements. The retrieval requirements will be based on existing data and historical records, such as historical battle cases, tactical manuals, intelligence analysis reports, etc.

[0132] 2) Information retrieval

[0133] The participating agent constructs highly semantically relevant query statements to retrieve information in a military knowledge base (which can be a vector database or other forms of databases). Using vector representation techniques (such as Word2Vector, BERT, etc.), the documents in the database are transformed into semantic vectors, thus achieving fast and effective semantic similarity retrieval. Through appropriate retrieval strategies and parameter settings, the sorting of relevant results is optimized. The most suitable historical battle cases, tactical plans and strategy fragments for the current task requirements are screened out.

[0134] 3) Generate a preliminary plan

[0135] The participating agent extracts useful information from the retrieval results, such as successful tactical templates, force configuration plans, etc., and integrates and analyzes them in combination with the current battlefield situation. With the assistance of the large language model, a preliminary action plan is generated, including key information such as clear objectives, action steps, force deployment, resource allocation, etc. The large language model can provide subsequent action suggestions for the participating agent, such as: "According to the current battlefield situation and the possible actions of the enemy, it is recommended to give priority to strengthening the reconnaissance of Hill X and prepare two sets of offensive plans, Plan A and Plan B."

[0136] In a military strategic deduction scenario, the environmental agent sets the initial conditions based on the battlefield situation (such as terrain, weather, the force deployments of both sides, etc.). For example, the Red Army's mission is to capture Hill X, so the RAG method is used to retrieve historical mountain offensive battle cases and relevant tactical materials. Through the analysis and integration of this information by the large language model, an initial action plan is preliminarily formed, such as selecting a suitable attack route, deploying fire support, formulating a feint plan, etc. The plan recording agent saves the detailed records of each step of the decision-making process, providing data support for subsequent optimization and evaluation.

[0137] This method that combines the large language model and RAG technology realizes the efficient utilization of historical experience and existing knowledge, provides more comprehensive and reliable information support for the participating agents, and thus forms a more reasonable and feasible initial action plan.

[0138] Step 1.3: Strategy simulation execution, interaction, and result evaluation.

[0139] In the simulation of military strategic deduction, the director agent controls and guides the strategy deduction through the powerful reasoning ability and deep semantic understanding of the large language model. With the help of the large language model, the director agent can not only execute the preset strategy procedures but also dynamically adjust the deduction path, simulating the complexity of battlefield changes and emergencies in the real world. For example, when it is identified that enemy reinforcements have arrived, the director agent may trigger new tactical challenges or adjust the force deployment strategy to test the reactions of the participants.

[0140] The participating agents (such as the Red Army and the Blue Army) must quickly respond to the constantly changing battlefield environment during the deduction process. Utilizing the efficient analysis ability of the large language model, the participating agents can evaluate the impact of their decisions on the enemy in real time and accordingly adjust their action plans. For example, after receiving a signal of an enemy air raid warning or a threatened supply line, the Red Army agent may redeploy its air defense forces or adjust the supply route. In this way, the quick response ability supported by the large language model enables more rapid and precise strategic adjustments.

[0141] 1) Initialization state

[0142] Based on the key battlefield information determined in the military strategic deduction model, set the agent roles and their initial tasks.

[0143] 2) Director agent controls the deduction

[0144] The director agent analyzes the battlefield change trend in real time through the large language model.

[0145] Input information: The set conditions provided by the environmental agent, the current deduction stage, the action plans of each participating party, etc.

[0146] Prompt: (The prompt template has been described in detail in Step 1)

[0147] ● Action strategy adjustment: It refers to the modification of the participating agent's own action plan according to the changes in the battlefield situation. For example, the red side discovers that the enemy's defense is weak and decides to adjust the attack route.

[0148] ● Action strategy notification: It means that the participating agent informs the director agent of its own action plan (including the adjusted plan). This is to enable the director agent to understand the action plans of each participating party for global coordination and control.

[0149] ● Generation method: The action strategy adjustment is generated by the participating agent using the large language model, combined with the current battlefield situation and its own task objectives. The action strategy notification is that the participating agent sends the adjusted action plan to the director agent in a specific format (such as JSON).

[0150] 3) Participating agent responds to instructions

[0151] The participating agent uses the task parsing ability extended by the large language model to respond to complex instructions from the director agent.

[0152] Battlefield condition changes: The environmental agent simulates battlefield condition changes by updating battlefield situation information (such as weather changes, enemy actions, etc.). The participating agent perceives the changes by receiving new information from the environmental agent.

[0153] Core strategy adjustment: Based on the large language model. The participating agent adjusts the troop deployment, action steps, etc. according to the new battlefield situation and task requirements.

[0154] Negotiate new resource allocation:

[0155] ● Original resource allocation: In the action plan generated in Step 2.

[0156] ● Negotiation process: The participating agent sends a request to the director agent, stating the types and quantities of resources that need to be adjusted.

[0157] The director agent decides whether to approve the request based on the current battlefield situation and resource conditions.

[0158] 4) The plan recording agent records in real time

[0159] The plan recording agent, combined with the understanding framework constructed by the large language model, records and stores important strategy changes and battlefield responses in each deduction stage by refining and tracking the global battlefield dynamics, forming a complete historical file of decision-making plans for post-event analysis.

[0160] Input information: current deduction stage, action plans of each participant, environmental changes, etc. (The hint template has been described in detail in Step 1).

[0161] 5) Interaction loop

[0162] The interaction among the director agent, participant agents, and scenario recording agent forms a loop.

[0163] The director agent controls the deduction process, sends instructions to the participant agents, and triggers scenario events (such as enemy reinforcements, weather changes, etc.).

[0164] The participant agents respond to the instructions, execute decision-making actions, and send requests or suggestions to the director agent (requests can include: requests for reinforcements, requests for supplies, requests for adjusting mission objectives, requests for intelligence support, etc.).

[0165] The scenario recording agent records the scenario changes and historical records during the entire deduction process.

[0166] Consider a specific scenario of a military confrontation game. The Red Army and the Blue Army represent two opposing military forces respectively. With the change of the battlefield environment, the battlefield conditions have increased the difficulty of troop mobilization and strategic objectives. To cope with this, the director agent triggers a series of additional battlefield emergencies, including unpredictable political changes and sudden changes in battlefield requirements.

[0167] Red Army strategy adjustment: After the new situation makes the defense cost of a certain key area increase, the Red Army agent analyzes the battlefield prospect through the large language model and decides to reduce the dependence on this area. At the same time, the Red Army agent mobilizes troops into a new combat area, reducing the risks that may arise due to local battlefield restrictions.

[0168] Blue Army countermeasures: The Blue Army agent relies on the large language model to re-evaluate historical battle cases, identifies combat directions with high benefits and low risks, and then adjusts its combat strategy. To confront sudden battlefield challenges, the Blue Army agent requests the director agent to provide battlefield intelligence to optimize the long-term troop deployment strategy.

[0169] This real-time interaction and strategy optimization method based on agents significantly improves the flexibility and adaptability of the entire military game system, fully demonstrating the unique advantages of the large language model in the actual decision-making process.

[0170] 6) Deduction end and result output

[0171] When the preset end conditions are met (such as occupying the target area or achieving specific strategies), the data related to the deduction process is archived, and these historical data are used for refined evaluation of the strategy, forming a set of reliable and applicable intelligent decision-making guidance programs.

[0172] Internal Review and Index Evaluation: First, use the intelligent evaluation tool built by the large language model to conduct a preliminary evaluation of the decision-making plan from multiple dimensions. The key evaluation indicators include: the accuracy of the plan, effectiveness (meeting the actual conditions), troop losses, resource utilization rate, and the rationality of the time arrangement.

[0173] Expert Review and Simulation Verification: Based on the internal evaluation, relevant experts in the field can be invited to review the generated plan. In addition, use the simulation system to simulate the actual environment of the plan and test the possible results of the plan in the envisioned scenario.

[0174] 7) Generate a Deduction Report

[0175] Generate a detailed deduction report according to the data collected during the deduction process and the finally determined decision-making plan. The content of the report should include: deduction objectives, process overview, key events, final decision-making plan, and evaluation results. An automated report generation tool can be used to ensure the efficiency of report production and the accuracy of information, and to ensure the traceability and readability of key data.

[0176] Through this series of steps, ensure that the decision-making plan not only has theoretical feasibility but also has been tested through practical simulations, improving the overall decision-making level and execution success rate.

[0177] Step 1.4: Plan Recording and Feedback Optimization.

[0178] During the deduction process, the plan recording agent plays a crucial role. Its main responsibility is to collect and record key data in the decision-making process in real time. These data include but are not limited to troop deployment, decision-making instructions, battlefield environment changes, action results, and resource consumption. To support efficient data management, the plan recording agent stores all the collected data in a structured, scalable, and efficiently queryable database (which can be a relational database, graph database, or time-series database, determined according to the data type and query requirements). This design not only adapts to the continuous growth of the data volume but also helps to ensure the timeliness and accuracy of information retrieval.

[0179] 1) Real-time Recording and Storage

[0180] Data Collection: The plan recording agent, through the improved information understanding ability of the large language model, listens in real time for strategy update signals, environment change instructions, and resource usage status (such as troop mobilization, etc.) from other agents. For example, when the battlefield environment affects the red side's adjustment of the defense deployment, the corresponding adjustment will be recorded by the plan recording agent in a timely manner.

[0181] Database Design: When interacting with the database, use a relational data model to establish relationship mapping tables between agents (such as force - strategy mapping, strategy - environment mapping), and optimize query efficiency through indexing. Special attention should be paid to the plasticity of data linkage in database design to enable multi - dimensional cross - verification and correlation reasoning in subsequent data processing and analysis.

[0182] 2) Feedback Mechanism and Optimization Adjustment

[0183] The scenario - recording agent not only records data but also continuously monitors and analyzes the implementation effect of strategies through a feedback mechanism. This process includes the following steps:

[0184] Data Analysis and Feedback Generation: Based on the large - language model, conduct multi - level analysis on the stored data to form a feedback report. The content of the report includes:

[0185] ● Evaluation of the execution effect of existing strategies (such as whether the expected goals are achieved, casualty situations, resource consumption situations, etc.).

[0186] ● Analysis of the rationality of force deployment

[0187] ● Analysis of changes in the battlefield situation

[0188] ● Assessment of potential risks and threats

[0189] ● Optimization suggestions (such as adjusting force deployment, changing tactical strategies, etc.)

[0190] Feedback Link: (Provide a timeline and data - flow example of the feedback loop)

[0191] ● T0 - Trigger the feedback mechanism: When new battlefield data indicates a change, the scenario - recording agent triggers the feedback mechanism.

[0192] ● T1 - Generate a feedback report: The scenario - recording agent uses the large - language model to generate a detailed feedback report.

[0193] ● T2 - Analyze the impact of feedback: The large - language model (which can be a participating - party agent or a specialized evaluation agent) evaluates the feedback report to identify strategy defects, risks of force losses, and potential threats.

[0194] ● T3 - Agent negotiation for optimization: Referring to the feedback report, the participating - party agent negotiates strategies with the director agent

[0195] to propose an optimization plan. The elements included in the report generation are:

[0196] ■ Deduction stage

[0197] ■ Actions and results of all parties

[0198] ■ Resource consumption

[0199] ■Key events (such as firefights, reconnaissance, airstrikes, etc.)

[0200] ■Preliminary assessment (such as strengths, weaknesses, risks, optimization suggestions, etc.)

[0201] ●T4—Execution plan optimization: The participating agents implement the optimization plan according to the negotiation structure and adjust the troop deployment and tactical strategies.

[0202] ●T5—Record optimization results: The solution recording agent records the optimized solution and execution results.

[0203] In a military battle simulation, if the Red Agent finds that its offensive strategy does not achieve the expected results due to the opponent's stubborn resistance and terrain restrictions, the plan recording agent will automatically generate a feedback report. Under this feedback mechanism, through consultation with the director agent, the Red Agent decides to optimize its force deployment and concentrate superior forces to break through the opponent's weak defense line to achieve the occupation of the target area. For the Blue Agent, the recording agent will associate its strategy execution with the dynamic changes of the battlefield to ensure real-time adjustment of the defense strategy.

[0204] Through such a systematic real-time feedback and optimization process, the solution recording agent ensures that the strategy of the entire game system is highly adaptable and responsive, and on this basis continuously improves the effectiveness and competitiveness of the entire system.

[0205] Example 2: Business game scenario.

[0206] Step 2.1: Construction of multi-agent system.

[0207] In the present invention, multiple agent modules are combined with a large language model to form an efficient multi-agent system. Each agent corresponds to a large language model instance. Through fine-tuning for specific roles and tasks, the agent realizes independent and collaborative decision-making functions. Specific agent types include environmental agents, director agents, participant agents, and solution recording agents, which together form a highly simulated economic game environment, creating a solid foundation for solution generation and optimization.

[0208] Each agent is fine-tuned by a dedicated large language model to suit its unique operating environment and task content. For example, the environmental agent uses the fine-tuned large language model to parse complex variables in the economic environment and set key parameters in the scenario, such as resource distribution, market trends, and legal constraints. The director agent uses the fine-tuned large language model to formulate evaluation criteria for each stage, monitor the game process in real time, and adjust the deduction path.

[0209] Among agents, collaboration is achieved through the enhanced understanding and generation capabilities of large language models, enabling agents with different roles to communicate effectively and exchange information in complex environments:

[0210] Environmental agent: Defines the decision-making scenario and economic background, including resources and market conditions, and guides other agents to identify key variables. Prompt example: "The current economic state is X, the influencing factors of resource distribution are Y, and the main limitation is Z, which may affect Strategy A."

[0211] Director agent: Responsible for the overall control of game deduction, with the ability to adjust strategies and initiate new events. Prompt example: "The game has entered the Nth stage. The red side has achieved the M goal. Should a new market dynamic be initiated? The blue side has used more than 50% of its resources and needs a strategy adjustment."

[0212] Participant agent: For example, the participating companies are divided into the red side and the blue side. They use large language models to process historical data and the current market situation to provide a basis for their action plans. Prompt example: "Your goal is to increase the market share of X. Currently, the resources are Y. It is recommended to execute Action Z first."

[0213] Scenario recording agent: Tracks and records the key events that occur during the interaction, providing data support for further analysis and optimization. Prompt example: "Record time T, the event is X, the impact is Y, and the response of the red / blue side is Z."

[0214] To clarify the above settings, consider an economic confrontation game scenario. Among them, the red side and the blue side represent two competing companies respectively. The environmental agent sets market conditions such as economic growth rate and tax policies. The director agent controls the process of the economic game in real time, simulating possible market fluctuations. The participant agents carry out strategy planning according to the prompts, such as adjusting the product line or exploring new market areas. The scenario recording agent accurately records every key decision and market reaction for evaluation.

[0215] This example will be processed and applied in more detail in the subsequent steps. From task setting, strategy deduction to final evaluation, by analyzing the interaction between the red side and the blue side in the whole economic game, it comprehensively demonstrates the efficiency and innovation of the multi-agent system of the present invention in actual operation.

[0216] Step 2.2: Task setting and assignment.

[0217] In a multi-agent system, the environmental agent is responsible for background information analysis and task objective setting. With the help of large language models, the environmental agent can deeply analyze complex variables in economic game scenarios, such as economic growth rate, market conditions, and tax policies, etc., and set precise task objectives accordingly, such as "The red side needs to occupy X% of the market share within the next 3 months". This intelligent setting process significantly improves task accuracy and the collaboration efficiency among agents.

[0218] After receiving the task, participating agents (such as the red side and the blue side) use the Retrieval-Augmented Generation (RAG) method combined with the powerful semantic understanding ability of large language models to analyze task requirements. This process includes the identification of key elements such as objectives, environmental conditions, time limits, etc., and clarifies the types and scopes of retrieved information, such as strategy types, resource allocation, market utilization, etc.

[0219] In this step, the application process of the Retrieval-Augmented Generation (RAG) method is as follows:

[0220] 1) Task analysis and requirement clarification

[0221] After receiving the task, the participating agents first use large language models for task analysis, accurately identify task objectives, environmental conditions, time limits, and clarify retrieval requirements. Based on existing data and historical records, consumer market behavior and enterprise strategies are used as the main reference elements.

[0222] 2) Information retrieval

[0223] By constructing highly semantically relevant query statements, the participating agents conduct information retrieval in the vector database. Using advanced vector representation techniques, the documents in the database are transformed into semantic vectors, thus enabling fast and effective semantic similarity retrieval. Through appropriate retrieval strategies and parameter settings, the accurate ranking of relevant results is optimized, such as adjusting the indexing algorithm to improve retrieval efficiency. The historical decision-making schemes and strategy fragments that best meet the task requirements are screened out.

[0224] 3) Generate a preliminary plan

[0225] From information extraction and integration: The participating agents extract useful information from the retrieval results, such as successful strategy templates, resource allocation plans, etc., and conduct integrated analysis in combination with the current market situation. With the assistance of large language models, a preliminary action plan is generated, which includes key information such as clear goal setting, action steps, and resource allocation. During the generation process, the large language model will generate subsequent prompts for the participating agents, such as: "Your goal is to increase X% of the market share, the existing resources are Y, and it is recommended that the primary action be Z."

[0226] During the economic game process, the environmental agent sets background conditions based on various indicators of the market economy, such as the current market growth forecast and tax adjustments. The Red Party's task is to increase market share, so it uses the RAG method to retrieve the company's historical successful cases and market response data. The initial formation of the plan is promoted through the large language model, such as adjusting the price strategy or entering a new market. The plan recording agent saves every detailed decision-making process for future optimization and evaluation.

[0227] This method maximizes the efficient utilization of historical experience and existing knowledge, providing stronger and more reliable information support for the participating agents, thus forming a more reasonable and feasible initial plan.

[0228] Step 2.3: Execution and interaction of the strategy simulation process.

[0229] In the simulation of the economic confrontation game, the director agent controls and guides the strategy deduction through the powerful reasoning ability and deep semantic understanding of the large language model. With the help of the large language model, the director agent can not only execute the preset strategy procedures, but also dynamically adjust the deduction path to simulate the complexity of market fluctuations and policy changes in the real world. For example, when identifying a slowdown in market growth, the director agent may trigger new economic challenges or adjust the resource allocation strategy to test the participants' reactions.

[0230] The participating agents (such as the Red Party and the Blue Party) must quickly respond to the changing economic environment during the deduction process. With the help of the efficient analysis ability of the large language model, the participating agents can evaluate the impact of their decisions on competitors in real time and adjust the action plan accordingly. For example, after receiving signals of market turmoil or resource depletion, the Red Party agent may reallocate resources or cut non-critical product lines. In this way, the quick response ability supported by the large language model makes strategic adjustments faster and more precise.

[0231] 1) Initial state

[0232] In the React application, create a global state manager (such as Redux or Context) to stabilize various shared states during the deduction process. During initialization, set the agent roles and their initial tasks according to the key economic indicators determined in the economic game model.

[0233] 2) Director agent controls the deduction

[0234] The director agent analyzes the economic change trends in the context of big data in real time through the large language model, and monitors the changes in the deduction steps in the global state. The large language model can help the director agent identify and predict potential changes in the market and automatically generate deduction instructions, improving the dynamic adaptability of the deduction process.

[0235] 3) The participating agent responds to the instruction

[0236] The participating agent utilizes the task parsing ability extended by the large language model to respond to complex instructions from the director agent. During the deduction activity, if the market conditions change drastically, the participating agent flexibly adjusts its core strategy and negotiates with the director agent for new resource allocation when necessary, forming a benign cycle adaptation mechanism.

[0237] 4) The scenario recording agent records the scenario

[0238] The scenario recording agent, in combination with the understanding framework constructed by the large language model, records and stores important strategy changes and market responses in each deduction stage by refining and tracking the global economic dynamics, forming a complete historical file of decision-making scenarios for post-event analysis.

[0239] 5) Interaction loop

[0240] The interaction among the director agent, the participating agent, and the scenario recording agent forms a loop. The director agent controls the deduction process, sends instructions to the participating agent, and triggers scenario events. The participating agent responds to the instructions, executes decision-making actions, and sends requests or suggestions to the director agent. The scenario recording agent records the scenario changes and historical records throughout the deduction process.

[0241] Consider a specific scenario of an economic confrontation game, where the Red side and the Blue side represent two highly competitive companies respectively. With the change of the economic environment, the market conditions have increased the difficulty of resource acquisition and investment return. To cope with this, the director agent triggers a series of additional market perturbation events, including unpredictable policy changes and sudden market demand changes.

[0242] Red side strategy adjustment: After the new policy increases the acquisition cost of a certain type of resource, the Red side agent analyzes the market prospect through the large language model and decides to reduce the dependence on this resource. At the same time, the Red side agent mobilizes resources to enter a new market area, reducing the risk that may be caused by local market restrictions.

[0243] Blue side countermeasures: The Blue side agent relies on the large language model to re-evaluate historical data, identifies product lines with high benefits and low risks, and then adjusts its product strategy. To confront sudden market challenges, the Blue side agent requests the director agent to provide market outlook data to optimize the long-term resource allocation strategy.

[0244] This real-time interaction and strategy optimization method based on agents significantly improves the flexibility and adaptability of the entire economic game system, fully demonstrating the unique advantages of the large language model in the actual decision-making process.

[0245] 6) Deduction end and result output

[0246] When the preset end conditions are met (such as market share targets or achieving specific strategies), the data related to the deduction process is archived, and these historical data are used for refined evaluation of the strategy to form a set of reliable and applicable intelligent decision-making guidance programs.

[0247] Step 2.4: Program Recording and Feedback Optimization.

[0248] During the deduction process, the program recording agent plays a crucial role. Its main responsibility is to collect key data in the decision-making process in real time by monitoring the changes in the global state manager. These data include but are not limited to resource allocation, decision instructions, environmental state changes, and response results. To support efficient data management, the program recording agent stores all the collected data in a structured, scalable, and efficiently queryable database. This design not only adapts to the continuous growth of data volume but also helps ensure the timeliness and accuracy of information retrieval.

[0249] 1) Real-time Recording and Storage

[0250] Data Acquisition: The program recording agent, through the enhanced information understanding ability of the large language model, listens in real time for strategy update signals, environmental change instructions, and resource usage status from other agents. For example, when new market trends affect the resource management of the red side, its corresponding adjustments will be recorded in a timely manner.

[0251] Database Design: When interacting with the database, a relational data model is used to establish relationship mapping tables between agents (such as resource - strategy mapping, strategy - environment mapping), and query efficiency is improved through index optimization. Special attention should be paid to the plasticity of data linkage in database design to facilitate multi-dimensional cross-validation and correlation reasoning in subsequent data processing and analysis.

[0252] 2) Feedback Mechanism and Optimization Adjustment

[0253] The program recording agent not only records data but also continuously monitors and analyzes the implementation effect of the strategy through the feedback mechanism. This process includes the following steps:

[0254] Data Analysis and Feedback Generation: Based on the large language model, multi-level analysis is performed on the stored data to generate a feedback report. The content of the report includes: whether the implementation effect of the existing strategy meets expectations, and the rationality and timeliness analysis of resource allocation.

[0255] Feedback Link: (Provide an example of the timeline and data flow of the feedback loop)

[0256] ● T0 - Trigger the feedback mechanism: When new market data indicates a change, the program recording agent generates a preliminary feedback.

[0257] ● T1 - Analyze feedback impact: Evaluate the initial feedback impact of the large language model, identify strategic deficiencies and risks of resource imbalance.

[0258] ● T2 - Agent negotiation optimization: Refer to the report, and the participating agent and the director agent conduct strategic negotiations to propose optimization plans.

[0259] ● T3 - Implementation plan optimization: Implement the optimized plan verified through analysis, and adjust resource allocation and strategic instructions.

[0260] Feasibility of optimization and adjustment: The optimization is carried out based on considering the actual needs of users and the constraints of the existing resource environment, making the adjusted plan operable and effective.

[0261] In the economic confrontation game, assume that the red agent discovers that its market strategy fails to achieve the expected effect due to the influence of uncontrollable environmental factors. The plan recording agent will automatically generate a feedback report. Under this feedback mechanism, through negotiation with the director agent, the red agent decides to optimize its resource allocation strategy, concentrating funds on growth jump points to maximize the increase in market share. Similarly for the blue side, the recording agent will bind its strategy execution to market dynamics to ensure real-time resource adjustment.

[0262] Through such a systematic real-time feedback and optimization process, the plan recording agent ensures that the strategies of the entire game system are highly adaptable and responsive, and continuously improves the effectiveness and competitiveness of the entire system on this basis.

[0263] Step 2.5: Result output and evaluation.

[0264] After the deduction ends, the system automatically generates the final decision plan, covering the complete action plan, resource allocation strategy, and time arrangement. This stage is a crucial step in comprehensively evaluating the generated plan, verifying its scientificity, feasibility, and potential impact.

[0265] 1) Plan evaluation and verification

[0266] Internal review and indicator evaluation First, use the intelligent evaluation tool constructed by the large language model to conduct a preliminary evaluation of the decision plan from multiple dimensions. The key evaluation indicators include: the accuracy of the plan, effectiveness (meeting real-world conditions), resource utilization rate, and rationality of time arrangement.

[0267] Based on the internal evaluation, expert review and simulation verification can invite experts in relevant fields to review the generated solutions to further confirm the usefulness and innovativeness of the solutions with the help of professional experience. In addition, the simulation system is used to simulate the actual environment of the solutions to test the possible results of the solutions in the envisioned scenarios. During the simulation process, the agents adjust their strategies by combining historical data and real-time variables to provide scientific execution path suggestions for decision-making. For example, by simulating the impact of different policy perturbations on the red side's market strategy, its adaptability is verified.

[0268] 2) Generate a deduction report

[0269] Based on the data collected during the deduction process and the finally determined decision-making solution, a detailed deduction report is generated. The report content should include: deduction objectives, process overview, key events, final decision-making solution, and evaluation results. Automated report generation tools can be used to ensure the efficiency and information accuracy of report production, and the traceability and readability of key data need to be ensured.

[0270] Through this series of steps, the system ensures that the decision-making solution not only has theoretical feasibility but also has been tested through practical simulations, improving the overall decision-making level and execution success rate. In this economic game example, the strategy optimization and implementation effects of the red side and the blue side are clear and intuitive, providing a detailed and reliable basis for decision-making.

[0271] In summary, the present invention constructs an efficient intelligent decision-making support framework by integrating advanced large language model technology and multi-agent systems. In various complex scenarios, through the interaction and cooperation between agents, the automated generation and dynamic optimization of decision-making solutions are realized, significantly improving the feasibility, scientificity, and adaptability of the solutions in practical applications.

[0272] The present invention fully exploits the capabilities of large language model technology in natural language processing, reasoning, and understanding of online questions, effectively improving the efficiency and quality of decision-making solution generation and optimization. This method not only provides a faster and more accurate solution for modern decision-making but also enhances the adaptability of solution implementation.

[0273] The present invention flexibly combines the advantages of large language models and multi-agent systems to achieve the goal of intelligent generation and optimization of decision-making solutions in complex scenarios. This method is applicable to various application scenarios such as military strategy deduction, business strategy formulation, policy formulation, and crisis management, greatly improving the scientificity and efficiency of decision-making and showing broad application potential and market value.

[0274] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it. It cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. A multi-agent game method based on a large language model, characterized in that: The method comprises: Using a large language model to construct an intelligent agent; wherein the intelligent agent includes: an environment intelligent agent, a director intelligent agent and a plurality of participant intelligent agents; The environment agent generates the tasks of each participant agent and the setting conditions of the director agent according to the environment situation, strategic goals and the strengths and weaknesses of each participant in the game scenario; Each participant agent generates an action plan for the participant agent according to the environmental situation and the task; The director agent determines whether to adjust the deduction process according to the deduction stage, set conditions and the action plans of each participant, and sends the deduction process adjustment instructions to the agents of each participant if the deduction process needs to be adjusted; Each participating intelligent agent adjusts the environmental situation according to the deduction process adjustment instruction, and generates a new action plan for the participating intelligent agent based on the adjusted environmental situation and the task.

2. The method according to claim 1, characterized in that The method of using a large language model to construct an intelligent agent includes: According to task requirements and computing resources, select a pre-trained large language model, which includes: GPT series, LLaMA series or other open source / commercial models; Fine-tuning the pre-trained large speech model based on a domain-specific dataset; wherein the dataset contains input-output pairs related to the agent role and task; The fine-tuned large speech model is evaluated by indicators, wherein the indicators include: BLEU, ROUGE and / or perplexity.

3. The method according to claim 1, characterized in that The game scenarios include: commercial game scenarios, military game scenarios, policy game scenarios or crisis management game scenarios.

4. The method according to claim 3, characterized in that In the case where the game scenario is a military game scenario, the task of generating each participant agent includes: The environmental situation, strategic goals, and the strengths and weaknesses of a participant are embedded into a task goal generation prompt template, and the task of the participant is obtained based on the large language model used by the environmental agent; wherein the environmental situation includes: terrain features, weather conditions, troop deployment of both sides, and resource conditions.

5. The method according to claim 3, characterized in that: In the case where the game scenario is a military game scenario, the setting conditions for generating the director agent include: The environmental situation and strategic goals are embedded in the setting condition generation prompt template, and the setting conditions of the director agent are obtained based on the large language model used by the environmental agent; wherein the setting conditions include: initial conditions, deduction rules, possible events and evaluation indicators, and the initial conditions include: the military strength, resources and terrain of each participant.

6. The method according to claim 3, characterized in that In the case where the game scenario is a military game scenario, each participant agent generates an action plan for the participant agent according to the environmental situation and the task, including: Utilize the large language model used by the participating agents to parse the task and obtain the task objectives, battlefield environmental conditions and time limits; Constructing a query statement according to the mission objective, the battlefield environment conditions and the time limit, and searching the knowledge base based on the query statement to obtain a search result; wherein the search result includes: historical battle examples, tactical plans and strategy fragments corresponding to the mission; Extracting action plan cases of each participant from the search results; The task, the environmental situation and the action plan cases of each participant are embedded into an action plan generation template, and an action plan is generated based on the large speech model used by the participant's intelligent agent.

7. The method according to claim 3, characterized in that In the case where the game scenario is a military game scenario, the director agent determines whether to adjust the game process according to the game stage, set conditions and the action plan of each participant, and sends the game process adjustment instruction to each participant agent if the game process needs to be adjusted, including: The deduction stage, set conditions and action plans of each participant are embedded into a prompt template for generating deduction process adjustment instructions, and the deduction process adjustment instructions are obtained based on the large language model used by the director agent; wherein the deduction process adjustment instructions include: triggering a new environmental situation, adjusting deduction rules, and issuing instructions to the participants to adjust action plans and / or provide intelligence support.

8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Using a large language model to build a program recording agent; The action plans of the deduction phase, the action plans of the participating agents, and the adjustment information of the environmental situation are embedded in the plan recording prompt template, and a feedback report is obtained based on the large language model used by the plan recording agent; wherein the feedback report includes: evaluation of the execution effect of the existing strategy, rationality analysis of force deployment, analysis of battlefield situation changes, potential risk and threat assessment and optimization suggestions; The feedback report is provided to each participating agent so that each participating agent can adjust the action plan based on the feedback report.

9. A multi-agent game system based on a large language model, characterized in that: The system includes: an environment agent, a director agent and several participant agents constructed using a large language model; wherein, The environment agent is used to generate tasks for each participant agent and setting conditions for the director agent according to the environment situation, strategic goals and advantages and disadvantages of each participant in the game scenario; The participant agent is used to generate an action plan for the participant agent according to the environmental situation and the task; The director intelligent agent is used to determine whether to adjust the deduction process according to the deduction stage, set conditions and the action plan of each participant, and if the deduction process needs to be adjusted, send a deduction process adjustment instruction to each participant intelligent agent, so that each participant intelligent agent adjusts the environmental situation according to the deduction process adjustment instruction, and generates a new action plan for the participant intelligent agent based on the adjusted environmental situation and the task.

10. The system according to claim 9, characterized in that The system further includes: a solution recording agent constructed using a large language model, wherein the solution recording agent is used to: The action plans of the deduction phase, the action plans of the participating agents, and the adjustment information of the environmental situation are embedded in the plan recording prompt template, and a feedback report is obtained based on the large language model used by the plan recording agent; wherein the feedback report includes: evaluation of the execution effect of the existing strategy, rationality analysis of force deployment, analysis of battlefield situation changes, potential risk and threat assessment and optimization suggestions; The feedback report is provided to each participating agent so that each participating agent can adjust the action plan based on the feedback report.

Citation Information

Cited By

  • An air game self-evolution decision tree opponent generation method based on a large model

    CN122527639A