Layered decision-making method and system for fusion of large language model and reinforcement learning

By constructing a hierarchical decision-making framework and a multi-agent sequential collaboration mechanism, the collaborative decision-making between large language models and reinforcement learning is optimized, solving the problems of insufficient accuracy and low collaboration efficiency in existing systems for complex tasks, and achieving autonomous evolution and efficient collaboration.

CN121997969APending Publication Date: 2026-05-08NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANKAI UNIV
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing collaborative decision-making systems based on large language models and reinforcement learning suffer from problems such as insufficient accuracy, slow response, reliance on manual design, and high computational requirements in complex tasks. They are unable to achieve efficient online learning and multi-agent collaboration and lack transparent and efficient collaboration mechanisms.

Method used

A hierarchical decision-making framework is constructed, and the underlying execution module is trained using the proximal gradient pruning algorithm. A memory feedback optimization module and a multi-agent sequential collaborative decision-making mechanism are designed. The prompts of the large language model are optimized through environmental feedback, and the collaborative relationship between agents is explicitly modeled.

Benefits of technology

It achieves the organic integration of large language models and reinforcement learning, improves the system's decision-making performance and collaborative efficiency, endows the system with autonomous evolution capabilities, and enhances the decision-making transparency and collaborative efficiency of multi-agent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997969A_ABST
    Figure CN121997969A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-agent autonomous decision-making, in particular to a hierarchical decision-making method and system for fusion of a large language model and reinforcement learning, and the method comprises the following steps: constructing a hierarchical decision-making framework comprising an upper-layer agent and a bottom-layer execution module; a bottom layer execution module for designing a hierarchical decision framework based on reinforcement learning training; designing an upper-layer intelligent agent of the hierarchical decision framework based on a large language model; a prompt instruction optimization iteration mechanism is designed, environment feedback is used as a signal, and continuous evolution of a prompt instruction is achieved through self-reflection of a large language model; a multi-agent sequential collaborative decision-making mechanism based on a thinking chain is introduced, and explicit reasoning and modeling of a multi-agent collaborative relation are achieved. According to the method and the system provided by the invention, the decision-making performance of the system is ensured, and meanwhile, the decision-making capability and the cooperation efficiency of the system in a multi-intelligent-body complex scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent autonomous decision-making technology, and in particular to a hierarchical decision-making method and system that integrates large language models and reinforcement learning. Background Technology

[0002] Collaborative decision-making using large language models and reinforcement learning is a key direction for enhancing the cognitive and executive capabilities of intelligent agents in complex tasks. Large language models, pre-trained with massive amounts of knowledge, possess powerful semantic understanding and task planning capabilities, and their decision-making processes exhibit high interpretability. However, they suffer from insufficient accuracy and slow response in real-time control tasks. Reinforcement learning, on the other hand, optimizes decision-making strategies through autonomous interaction with the environment, excelling in low-level precise control. However, it faces challenges such as a large policy search space, reliance on extensive interaction data, and poor interpretability of its decision logic. Integrating large language models and reinforcement learning effectively combines the former's high-level reasoning capabilities with the latter's low-level execution advantages, enhancing both decision-making intelligence and system interpretability. Therefore, it holds broad application prospects in open and complex decision-making tasks.

[0003] Previous research on collaborative decision-making between large language models and reinforcement learning has mostly adopted a static combination strategy. The prompts of large language models often rely on manual design and are fixed, unable to evolve autonomously based on environmental feedback. Alternatively, reinforcement learning can be used to fine-tune the large language model to make its output more consistent with the task reward signal, but this requires too much computing power and is difficult to apply in real-world scenarios. Moreover, in complex collaborative tasks, the decision-making is mostly implicit, lacking a transparent and efficient collaborative mechanism. This relatively loose and static coupling method, while combining the reasoning advantages of large language models with the execution capabilities of reinforcement learning to some extent, fails to form an organically collaborative decision-making closed loop. This limitation at the architectural level makes it difficult for existing systems to achieve efficient online learning and multi-agent collaboration while maintaining interpretability, severely restricting the practical application effect and reliability in complex decision-making tasks. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a hierarchical decision-making method and system that integrates large language models and reinforcement learning, which significantly improves the decision-making ability and collaborative efficiency of the system in complex multi-agent scenarios while ensuring the system's decision-making performance.

[0005] A hierarchical decision-making method integrating a large language model and reinforcement learning includes the following steps: S1: Construct a hierarchical decision-making framework that includes upper-level intelligent agents and lower-level execution modules; S2: The underlying execution module of a hierarchical decision-making framework based on reinforcement learning training design; S3: The upper-level intelligent agent of a hierarchical decision-making framework designed based on a large language model; S4: The upper-layer intelligent agent outputs macro-planning to the lower-layer execution module. The lower-layer execution module generates specific execution actions based on the macro-planning, applies the specific execution actions to the environment to obtain environmental feedback, and transmits the environmental feedback to the memory feedback optimization module of the large language model. The memory feedback optimization module iteratively optimizes the system prompt instructions of the large language model based on the environmental feedback, generates the optimal system prompt instructions, and obtains the upper-layer intelligent agent under the optimal system prompt instructions. S5: The multi-agent sequential collaborative decision-making module of the large language model defines the decision-making order of the upper-level agents under the optimal system prompt instructions, and generates the decision actions of the upper-level agents under the optimal system prompt instructions according to the decision-making order of the upper-level agents under the optimal system prompt instructions. Then, based on the decision actions of the upper-level agents under the optimal system prompt instructions, explicit reasoning is performed to obtain the actions and analysis of the upper-level agents under the optimal system prompt instructions. The actions and analysis of the upper-level agents under the optimal system prompt instructions are used as the planning information of the upper-level agents and input to the lower-level execution module. The lower-level execution module generates the corresponding actions for execution.

[0006] Furthermore, in step S1, the upper-layer intelligent agent is responsible for task planning of the hierarchical decision-making framework, the lower-layer execution module is responsible for executing specific actions, and the upper-layer intelligent agent and the lower-layer execution module exchange information and transmit instructions through a structured interface.

[0007] In the optimized version, step S2 employs a proximal gradient pruning algorithm for reinforcement learning, and designs the underlying execution module of the hierarchical decision framework.

[0008] Furthermore, in step S2, the proximal gradient pruning algorithm is used for reinforcement learning. When designing the bottom-level execution module of the hierarchical decision framework, the loss function is calculated according to equation (1): (1); in: Value network loss function, This indicates the operation of taking the average. Indicates the system status. Indicates system state The value function at that point, Represents the objective function. Indicates the network loss of the policy. Indicates a probability factor. Represents the dominance function. This represents the truncation function. This represents the gradient clipping factor.

[0009] Furthermore, in step S4, the memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback, which includes two parts: decision storage and reflection iteration. The decision storage part calculates the complete decision trajectory of the upper-level agent according to equation (2), and the reflection iteration part generates prompts through multiple rounds of reflection iteration, and evaluates the prompts according to equation (3) to generate the optimal prompts. The evaluation formula is as follows: (2); (3); in: This represents the complete decision-making trajectory of the upper-level intelligent agent. express The game dynamics after the textualization of moments express The upper-level decisions generated in real time express Always give back, express The game dynamics after the textualization of moments Indicates the length of the trajectory. This indicates the optimal prompt instruction. This indicates a prompt to find the maximum value of the evaluated function. Represents the evaluation function, Indicates the first The prompts generated by the round of iterations, This represents the total number of iterations.

[0010] Furthermore, in step S5, the agent sequential collaborative decision-making module performs explicit reasoning based on equation (4) to obtain the actions and analysis of the upper-level agent under the optimal system prompt instruction: (4); in: Indicates the first Actions generated by a higher-level intelligent agent Indicates the first Analysis of the generation of upper-level intelligent agents This represents the upper-level game strategy based on a large language model. Indicates the system status. This represents the sequence of actions and analysis of the preceding upper-level intelligent agent.

[0011] In the optimized step S2, when designing the underlying execution module of the hierarchical decision-making framework based on reinforcement learning, the number of training rounds is 1000.

[0012] In the optimized step S4, the memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback for 100 rounds.

[0013] A hierarchical decision-making system integrating a large language model and reinforcement learning is used to execute a hierarchical decision-making method integrating a large language model and reinforcement learning as described above, comprising a hierarchical decision-making framework, a bottom-level execution module design unit, a large language model and an upper-level intelligent agent design unit. The hierarchical decision-making framework includes an upper-layer intelligent agent and a lower-layer execution module. The upper-layer intelligent agent and the lower-layer execution module are connected through a structured interface. The upper-layer intelligent agent is responsible for the task planning of the hierarchical decision-making framework, and the lower-layer execution module is used to execute specific actions. The underlying execution module design unit designs the underlying execution module based on reinforcement learning training. The upper-layer intelligent agent design unit designs upper-layer intelligent agents based on a large language model and generates task planning for a hierarchical decision framework. The large language model includes a memory feedback optimization module and a multi-agent sequential collaborative decision-making module. The memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback to generate the optimal system prompts. The multi-agent sequential collaborative decision-making module specifies the decision order of the upper-level agents under the optimal system prompts and generates the decision actions of the upper-level agents under the optimal system prompts based on the decision order of the upper-level agents under the optimal system prompts. Then, explicit reasoning is performed based on the decision actions of the upper-level agents under the optimal system prompts to obtain the actions and analysis of the upper-level agents under the optimal system prompts.

[0014] Beneficial effects of the invention: This invention provides a hierarchical decision-making method and system that integrates a large language model and reinforcement learning. By introducing a hierarchical decision-making framework, it achieves effective integration of the large language model and reinforcement learning, and proposes a reflective iterative optimization and multi-agent sequential collaborative decision-making mechanism to enable continuous system evolution and improve its performance in multi-agent collaborative tasks. By constructing a hierarchical architecture of "large language model-reinforcement learning," the high-level reasoning ability of the large language model is organically combined with the precise control ability of reinforcement learning, effectively improving the system's low-level micro-management decision-making performance and high-level cognitive reasoning ability. By designing a prompting iterative optimization mechanism based on environmental feedback, the autonomous evolution of prompting instructions is achieved, significantly reducing the dependence on manual prompting and endowing the system with the ability to continuously and autonomously evolve. By introducing a multi-agent sequential collaborative mechanism based on thought chains, implicit collaborative relationships are transformed into explicit and traceable reasoning chains, effectively improving the decision-making transparency and collaborative efficiency of the multi-agent system. This invention demonstrates excellent adaptability and generalization ability in various complex decision-making scenarios, providing an interpretable, evolvable, and highly collaborative solution for intelligent decision-making systems. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the process of this invention. Detailed Implementation

[0016] A hierarchical decision-making method integrating a large language model and reinforcement learning includes the following steps, the flowchart of which is shown below. Figure 1 As shown: S1: Construct a hierarchical decision-making framework that includes upper-level intelligent agents and lower-level execution modules; the construction of the hierarchical decision-making framework provides a foundation for the integration of reinforcement learning and large language models.

[0017] Specifically, the upper-layer intelligent agent is responsible for task planning within the hierarchical decision-making framework, while the lower-layer execution module is responsible for executing specific actions. The upper-layer intelligent agent and the lower-layer execution module interact and exchange instructions through a structured interface.

[0018] The underlying execution module, as the basic action execution module, can include four typical actions: attack, defense, detection, and evasion, each suitable for different stages and situations. The upper-layer agent, as the policy scheduling center, is responsible for dynamically assigning specific task objectives to each agent based on the global state. By dynamically allocating different task objectives, it achieves optimized control of the macro-level situation, effectively improving the system's adaptability and decision-making level in changing environments.

[0019] S2: The underlying execution module of the hierarchical decision-making framework based on reinforcement learning training design can generate high-performance underlying micro-operation execution strategies, effectively improving the underlying micro-operation decision-making performance of the system.

[0020] Specifically, a near-end gradient pruning algorithm can be used for reinforcement learning to design the underlying execution module of a hierarchical decision-making framework. The actions of the underlying execution module can be trained to cope with different multi-agent collaborative scenarios.

[0021] The four typical actions of the underlying execution module can use the same state space and action space. The state space consists of three parts: the agent's own body information, teammate information, and detectable opponent information. The action space includes maneuvering actions and equipment actions. Maneuvering actions are the control quantities of throttle and rudder angle, while equipment actions control whether to launch.

[0022] During training, different reward functions can be used for the four typical actions. Specifically, offensive actions are rewarded for hitting the opponent; defensive actions are rewarded for moving away from the opponent; probing actions are rewarded for obtaining the opponent's position information; and evasive actions are rewarded for breaking free from the opponent's lock-on.

[0023] Given that different typical actions require different training data, each underlying action corresponds to a set of value networks and policy networks. Specifically, the loss functions of the value networks and policy networks can be calculated according to equation (1): (1); in: Value network loss function, This indicates the operation of taking the average. Indicates the system status. Indicates system state The value function at that point, Represents the objective function. Indicates the network loss of the policy. Indicates a probability factor. Represents the dominance function. This represents the truncation function. This represents the gradient clipping factor.

[0024] Specifically, when designing the underlying execution module of a hierarchical decision-making framework based on reinforcement learning training, the number of training rounds can preferably be 1000 rounds.

[0025] S3: An upper-level intelligent agent based on a hierarchical decision-making framework designed using a large language model, capable of generating high-level decision-making strategies with efficient reasoning and interpretability.

[0026] S4: The upper-layer intelligent agent outputs macro-planning to the lower-layer execution module. The lower-layer execution module generates specific execution actions based on the macro-planning, applies the specific execution actions to the environment to obtain environmental feedback, and transmits the environmental feedback to the memory feedback optimization module of the large language model. The memory feedback optimization module iteratively optimizes the system prompt instructions of the large language model based on the environmental feedback, generates the optimal system prompt instructions, and obtains the upper-layer intelligent agent under the optimal system prompt instructions. The hierarchical decision-making framework's task planning comprises four parts: situational semantic transformation, prompts and instructions, decision parsing and generation, and memory feedback optimization. Situational semantic transformation primarily converts state numerical information (such as current pose information) into structured natural language descriptions that include spatial relationships, contextual information, and task objectives. Prompts and instructions are the input prompts required for the large language model's decision-making process, including system prompts and multi-agent decision prompts. System prompts predefine the decision-making roles, core task objectives, and referable strategies of the large language model, and strictly define the format of its output text to constrain the large language model's thinking framework and output paradigm. Multi-agent decision prompts specify the order and reasoning logic of multiple agents generating text decisions to define the large language model's... The model defines the collaborative relationships among multiple agents (e.g., defining agents to simultaneously output text decisions or output text sequentially); decision parsing and generation transforms the natural language decision text generated by the large language model into structured control instructions that the system can execute, which are then handed over to the trained underlying execution module to execute specific micro-operation instructions; memory feedback optimization continuously records and analyzes the decision trajectory, including the situational semantic decisions generated by each decision, the input prompt instructions, the semantic decisions generated by the large language model, and the final win / loss reward signal obtained after the decision is executed. Based on this memory data, and through the prompt instruction iterative optimization mechanism designed below, the system prompt instructions can be continuously optimized to achieve autonomous optimization and iterative improvement of multiple agents, ensuring the efficiency and rationality of strategy allocation and target decision-making.

[0027] The upper-level decision-making strategy is constructed based on a large language model, which endows the agent with efficient upper-level cognitive reasoning ability. Secondly, an iterative optimization mechanism for prompting instructions based on reflection of the large language model is designed. The environmental feedback is used as an optimization signal to drive the continuous autonomous evolution of prompting instructions, realizing the autonomous evolution of prompts, significantly reducing the system's dependence on manual prompting, and endowing the system with the ability to continuously and autonomously evolve.

[0028] Specifically, the memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback, which includes two parts: decision storage and reflection iteration. The decision storage part calculates the complete decision trajectory of the upper-level agent according to equation (2), and the reflection iteration part generates prompts through multiple rounds of reflection iteration, and evaluates the prompts according to equation (3) to generate the optimal prompts. The evaluation formula is as follows: (2); (3); in: This represents the complete decision-making trajectory of the upper-level intelligent agent. express The game dynamics after the textualization of moments express The upper-level decisions generated in real time express Always give back, express The game dynamics after the textualization of moments Indicates the length of the trajectory. This indicates the optimal prompt instruction. This indicates a prompt to find the maximum value of the evaluated function. Represents the evaluation function, Indicates the first The prompts generated by the round of iterations, This represents the total number of iterations.

[0029] Specifically, the memory feedback optimization module can preferably perform iterative optimization of the system prompts of the large language model based on environmental feedback for 100 rounds, with 50 decision data sets collected in each round. In the multi-agent collaborative decision-making mechanism, the data is executed sequentially in ascending order of the agent number.

[0030] S5: The multi-agent sequential collaborative decision-making module of the large language model defines the decision-making order of the upper-level agents under the optimal system prompt instructions, and generates the decision actions of the upper-level agents under the optimal system prompt instructions according to the decision-making order of the upper-level agents under the optimal system prompt instructions. Then, based on the decision actions of the upper-level agents under the optimal system prompt instructions, explicit reasoning is performed to obtain the actions and analysis of the upper-level agents under the optimal system prompt instructions. The actions and analysis of the upper-level agents under the optimal system prompt instructions are used as the planning information of the upper-level agents and input to the lower-level execution module. The lower-level execution module generates the corresponding actions for execution.

[0031] This step creates a traceable chain of thought, transforming implicit collaborative relationships into explicit, traceable reasoning chains, effectively improving the decision-making transparency and collaborative efficiency of multi-agent systems.

[0032] Specifically, the agent sequential collaborative decision-making module can perform explicit reasoning based on equation (4) to obtain the actions and analysis of the upper-level agent under the optimal system prompt instruction: (4); in: Indicates the first Actions generated by a higher-level intelligent agent Indicates the first Analysis of the generation of upper-level intelligent agents This represents the upper-level game strategy based on a large language model. Indicates the system status. This represents the sequence of actions and analysis of the preceding upper-level intelligent agent.

[0033] This invention provides a hierarchical decision-making method that integrates a large language model and reinforcement learning. First, it constructs a hierarchical decision-making architecture, providing a foundation for the integration of reinforcement learning and the large language model. Then, it generates low-level execution modules based on reinforcement learning, effectively improving the system's low-level micro-management decision-making performance. Next, it constructs high-level decision-making strategies based on the large language model, endowing the agent with efficient high-level cognitive reasoning capabilities. Second, it designs an iterative optimization mechanism for prompts based on reflection within the large language model, using environmental feedback as an optimization signal to drive the continuous autonomous evolution of prompts, achieving autonomous progression of prompts. Finally, it introduces a multi-agent sequential collaborative decision-making mechanism based on thought chains, explicitly modeling the collaborative relationships between agents and improving the collaborative efficiency among multiple agents. This invention endows the system with the ability for continuous evolution and multi-agent collaboration, improving the decision-making generalization and efficiency under complex multi-agent tasks, and providing an efficient and comprehensive solution for agent learning in various fields that need to cope with complex environments.

[0034] A hierarchical decision-making system integrating a large language model and reinforcement learning is used to execute a hierarchical decision-making method integrating a large language model and reinforcement learning as described above, comprising a hierarchical decision-making framework, a bottom-level execution module design unit, a large language model and an upper-level intelligent agent design unit. The hierarchical decision-making framework includes an upper-layer intelligent agent and a lower-layer execution module. The upper-layer intelligent agent and the lower-layer execution module are connected through a structured interface. The upper-layer intelligent agent is responsible for the task planning of the hierarchical decision-making framework, and the lower-layer execution module is used to execute specific actions. The underlying execution module design unit designs the underlying execution module based on reinforcement learning training. The upper-layer intelligent agent design unit designs upper-layer intelligent agents based on a large language model and generates task planning for a hierarchical decision framework. The large language model includes a memory feedback optimization module and a multi-agent sequential collaborative decision-making module. The memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback to generate the optimal system prompts. The multi-agent sequential collaborative decision-making module specifies the decision order of the upper-level agents under the optimal system prompts and generates the decision actions of the upper-level agents under the optimal system prompts based on the decision order of the upper-level agents under the optimal system prompts. Then, explicit reasoning is performed based on the decision actions of the upper-level agents under the optimal system prompts to obtain the actions and analysis of the upper-level agents under the optimal system prompts.

[0035] In summary, the present invention provides a hierarchical decision-making method and system that integrates a large language model and reinforcement learning. By embedding the large language model and reinforcement learning into a unified hierarchical decision-making framework, and designing a prompting iterative optimization mechanism based on self-reflection and a multi-agent sequential collaborative decision-making mechanism, the decision-making logic and collaboration mode are dynamically optimized. While ensuring the system's decision-making performance, the decision-making ability and collaboration efficiency in complex multi-agent scenarios are significantly improved.

[0036] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A hierarchical decision-making method integrating a large language model and reinforcement learning, characterized in that: Includes the following steps: S1: Construct a hierarchical decision-making framework that includes upper-level intelligent agents and lower-level execution modules; S2: The underlying execution module of a hierarchical decision-making framework based on reinforcement learning training design; S3: The upper-level intelligent agent of a hierarchical decision-making framework designed based on a large language model; S4: The upper-layer intelligent agent outputs macro-planning to the lower-layer execution module. The lower-layer execution module generates specific execution actions based on the macro-planning, applies the specific execution actions to the environment to obtain environmental feedback, and transmits the environmental feedback to the memory feedback optimization module of the large language model. The memory feedback optimization module iteratively optimizes the system prompt instructions of the large language model based on the environmental feedback, generates the optimal system prompt instructions, and obtains the upper-layer intelligent agent under the optimal system prompt instructions. S5: The multi-agent sequential collaborative decision-making module of the large language model defines the decision-making order of the upper-level agents under the optimal system prompt instructions, and generates the decision actions of the upper-level agents under the optimal system prompt instructions according to the decision-making order of the upper-level agents under the optimal system prompt instructions. Then, based on the decision actions of the upper-level agents under the optimal system prompt instructions, explicit reasoning is performed to obtain the actions and analysis of the upper-level agents under the optimal system prompt instructions. The actions and analysis of the upper-level agents under the optimal system prompt instructions are used as the planning information of the upper-level agents and input to the lower-level execution module. The lower-level execution module generates the corresponding actions for execution.

2. The hierarchical decision-making method integrating a large language model and reinforcement learning as described in claim 1, characterized in that: The upper-layer intelligent agent described in step S1 is responsible for task planning of the hierarchical decision-making framework, and the lower-layer execution module is responsible for executing specific actions. The upper-layer intelligent agent and the lower-layer execution module exchange information and transmit instructions through a structured interface.

3. The hierarchical decision-making method integrating a large language model and reinforcement learning as described in claim 1, characterized in that: In step S2, the proximal gradient pruning algorithm is used for reinforcement learning to design the underlying execution module of the hierarchical decision framework.

4. The hierarchical decision-making method integrating a large language model and reinforcement learning as described in claim 3, characterized in that: Step S2 employs the proximal gradient pruning algorithm for reinforcement learning. When designing the underlying execution module of the hierarchical decision framework, the loss function is calculated according to equation (1): (1); in: Represents the network loss function. This indicates the operation of taking the average. Indicates the system status. Indicates system state The value function at that point, Represents the objective value function, Indicates the network loss of the policy. Represents a probability factor. Represents the dominance function. This represents the truncation function. This represents the gradient clipping factor.

5. The hierarchical decision-making method integrating a large language model and reinforcement learning according to claim 4, characterized in that: In step S4, the memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback. This includes two parts: decision storage and reflection iteration. The decision storage part calculates the complete decision trajectory of the upper-level agent according to equation (2). The reflection iteration part generates prompts through multiple rounds of reflection iteration and evaluates the prompts according to equation (3) to generate the optimal prompts. The evaluation formula is as follows: (2); (3); in: This represents the complete decision-making trajectory of the upper-level intelligent agent. express The game dynamics after the textualization of moments express The upper-level decisions generated in real time express Always give back, express The game dynamics after the textualization of moments Indicates the length of the trajectory. This indicates the optimal prompt instruction. This indicates a prompt to find the maximum value of the evaluated function. Represents the evaluation function, Indicates the first The prompts generated by the round of iterations, This represents the total number of iterations.

6. The hierarchical decision-making method integrating a large language model and reinforcement learning according to claim 1, characterized in that: In step S5, the agent sequential collaborative decision-making module performs explicit reasoning based on equation (4) to obtain the actions and analysis of the upper-level agent under the optimal system prompt instruction: (4); in: Indicates the first Actions generated by a higher-level intelligent agent Indicates the first Analysis of the generation of upper-level intelligent agents This represents the upper-level game strategy based on a large language model. Indicates the system status. This represents the sequence of actions and analysis of the preceding upper-level intelligent agent.

7. The hierarchical decision-making method integrating a large language model and reinforcement learning according to claim 1, characterized in that: In step S2, when designing the underlying execution module of the hierarchical decision-making framework based on reinforcement learning training, the number of training rounds is 1000.

8. The hierarchical decision-making method integrating a large language model and reinforcement learning according to claim 1, characterized in that: In step S4, the memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback for 100 rounds.

9. A hierarchical decision-making system integrating a large language model and reinforcement learning, used to execute a hierarchical decision-making method integrating a large language model and reinforcement learning as described in any one of claims 1 to 8, characterized in that: This includes a hierarchical decision-making framework, a lower-level execution module design unit, a large language model, and an upper-level intelligent agent design unit; The hierarchical decision-making framework includes an upper-layer intelligent agent and a lower-layer execution module. The upper-layer intelligent agent and the lower-layer execution module are connected through a structured interface. The upper-layer intelligent agent is responsible for the task planning of the hierarchical decision-making framework, and the lower-layer execution module is used to execute specific actions. The underlying execution module design unit designs the underlying execution module based on reinforcement learning training. The upper-layer intelligent agent design unit designs upper-layer intelligent agents based on a large language model and generates task planning for a hierarchical decision framework. The large language model includes a memory feedback optimization module and a multi-agent sequential collaborative decision-making module. The memory feedback optimization module iteratively optimizes the system prompts of the large language model based on environmental feedback to generate the optimal system prompts. The multi-agent sequential collaborative decision-making module specifies the decision order of the upper-level agents under the optimal system prompts and generates the decision actions of the upper-level agents under the optimal system prompts based on the decision order of the upper-level agents under the optimal system prompts. Then, explicit reasoning is performed based on the decision actions of the upper-level agents under the optimal system prompts to obtain the actions and analysis of the upper-level agents under the optimal system prompts.