Policy decision-making method and device based on world model
By combining the multi-agent tree search algorithm with the MCTS world model planning improved prediction model network, the problem of low sample efficiency in multi-agent reinforcement learning is solved, and efficient training in complex environments is achieved.
Patent Information
- Application Number
- CN202510910449.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multi-agent reinforcement learning algorithms suffer from low sample efficiency in large-scale decision-making tasks and are difficult to train effectively in complex multi-agent environments, especially due to training difficulties caused by environmental non-stationarity, high dimensionality of joint action space, and complexity of policy search space.
We adopt a world model-based policy decision-making method, combining multi-agent tree search algorithm and MCTS world model planning. We generate latent vector representations through an improved prediction model network, select decision actions through multi-agent tree search algorithm, and update network parameters by combining training data pool.
Under the premise of the same policy performance, it significantly improves the sample efficiency of multi-agent reinforcement learning, especially the search efficiency in large action spaces.
Smart Images

Figure CN120975172A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-agent reinforcement learning technology, and in particular to a policy decision-making method and apparatus based on a world model. Background Technology
[0002] In recent years, Multi-Agent Reinforcement Learning (MARL) has achieved remarkable success, with its algorithms making significant breakthroughs in solving large-scale decision-making tasks. It has been applied to real-time strategy games (Arulkumaran et al., 2019; Ye et al., 2020), card games (Bard et al., 2020), sports games (Kurach et al., 2020), autonomous driving (Zhou et al., 2020), and multi-robot navigation (Long et al., 2018).
[0003] However, most existing MARL algorithms (primarily MAPPO and QMIX) are model-free, limiting sample efficiency and hindering their application in more challenging scenarios. For example, MAPPO often requires more than 5,000,000 interactive training steps to converge in a simple 3v3 scenario in the SMAC environment.
[0004] A key issue contributing to the low sample efficiency of multi-agent algorithms lies in the non-stationarity of the multi-agent environment setup. Agents continuously update their policies based on observations and rewards, leading to constant changes in the environment faced by each agent (Nguyen et al., 2020). Furthermore, the dimensionality of the joint action space grows exponentially with the number of agents, resulting in a massive policy search space (Hernandez-Leal et al., 2020). These challenges, along with issues related to partial observability, coordination, and credit assignment, necessitate that MARL requires a large number of samples for effective training (Gronauer & Diepold, 2022).
[0005] In contrast, model-based reinforcement learning (MBRL) has demonstrated value in sample efficiency in single-agent RL scenarios, both in practice (Wang et al., 2019) and in theory (Sun et al., 2019). Unlike model-free methods, MBRL methods typically focus on learning parameterized models to characterize the transition or reward functions of the real environment (Sutton & Barto, 2018; Corneil et al., 2018; Ha & Schmidhuber, 2018). Based on how the learned model is used, MBRL methods can be broadly divided into two categories: model-based planning (Hewing et al., 2020; Nagabandi et al., 2018; Wang & Ba, 2019; Schrittwieser et al., 2020; Hansen et al., 2022) and model-based data augmentation (Kurutach et al., 2018; Janner et al., 2019; Hafner et al., 2019; 2020; 2023). Due to the forward-looking nature of the planning method and the theoretically guaranteed convergence, the MBRL combined with planning usually exhibits significantly higher sample efficiency and faster convergence speed (Zhang et al., 2020).
[0006] The well-known planning-based MBRL method is MuZero, which performs MCTS through a value equivalence learning model. MuZero has demonstrated superhuman performance on limited data in many tasks, such as Atari video games and board games such as Go, chess and shogi (Schrittwieser et al., 2020).
[0007] However, extending single-agent MBRL methods to multi-agent environments is extremely challenging. On the one hand, existing single-agent algorithm models do not consider multi-agent-specific biases (such as near-independence of agents), and directly adopting the flattened model of single-agent MBRL is difficult to support efficient learning in real multi-agent environments. On the other hand, the state-action space of a multi-agent environment is much more complex than that of a single agent, leading to an exponential increase in search complexity, forcing us to explore dedicated search algorithms for complex action spaces. Furthermore, the model form and its generalization ability jointly constrain the form and efficiency of the search algorithm, and vice versa, making model design and search algorithm design highly correlated.
[0008] Therefore, in scenarios such as policy interaction, it is not possible to directly improve sample training efficiency by extending the single-agent MBRL method to a multi-agent environment. Summary of the Invention
[0009] The technical problem to be solved by this invention is how to improve the training efficiency of multi-agent reinforcement learning in the field of strategy interaction, and to provide a strategy decision-making method and device based on a world model.
[0010] According to embodiments of the present invention, a strategy decision-making method based on a world model includes: S10, acquire situational information during strategic interactions; S20, a prediction model network based on the Muzero model network is used to generate a latent vector representation based on the situation information; S30, the decision action is selected and determined based on the latent vector representation using a multi-agent tree search algorithm, and then output to the agents in the policy interaction.
[0011] The world-model-based policy decision-making method of this invention, according to embodiments of the present invention, is the first to employ a multi-agent model algorithm that combines multi-agent reinforcement learning with MCTS world model planning, thereby improving search efficiency in large action spaces compared to the most advanced model-free multi-agent methods currently available. Experiments demonstrate that, compared to existing multi-agent reinforcement learning algorithms, it exhibits superior sample efficiency while maintaining the same policy performance.
[0012] According to some embodiments of the present invention, the method further includes: S40, save strategy interaction game data to the training data pool; S50, the network parameters of the prediction model network are updated based on the data in the training data pool using a policy optimization algorithm.
[0013] In some embodiments of the present invention, the prediction model network in step S20 includes: The representation function is used to map the agent's current observation history into the individual's latent state vector; A dynamic function is used to derive the potential state vector of an individual at the next moment based on the individual's potential state vector and future actions; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
[0014] According to some embodiments of the present invention, the prediction model network in step S20 includes: The representation function is used to map the agent's current observation history to an individual potential state vector; Communication functions are used to generate collaborative features for each agent through an attention mechanism; Dynamic functions are used to derive the next local latent state vector based on individual state vectors, future actions, and communication features; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
[0015] In some embodiments of the present invention, the information stored for each edge in the multi-intelligent tree search algorithm in step S30 is as follows: N(s, a), P (s, a), Q(s, a), R(s, a), S(s, a), Where s represents the state information of the policy interaction environment at the current moment, a represents the action taken by each agent, N represents the number of times the state-action pair is accessed, P represents the policy probability distribution, Q represents the state value, R represents the reward value, and S represents the predicted state at the next moment.
[0016] According to an embodiment of the present invention, a strategy decision-making device based on a world model includes: The information acquisition module is used to acquire situational information during strategy interactions; The prediction model network, which is an improved prediction model network based on the Muzero model network, is used to generate latent vector representations based on the situation information. The search module is used to select and determine decision actions based on the latent vector representation using a multi-agent tree search algorithm, and output the results to the agents in the policy interaction.
[0017] The world-model-based policy decision-making device according to embodiments of the present invention is the first to employ a multi-agent model algorithm that combines multi-agent reinforcement learning with MCTS world model planning, thereby improving search efficiency in large action spaces compared to the most advanced model-free multi-agent methods currently available. Experiments demonstrate that, compared to existing multi-agent reinforcement learning algorithms, it exhibits superior sample efficiency while maintaining the same policy performance.
[0018] According to some embodiments of the present invention, the apparatus further includes: The training data pool is used to store strategy interaction game data; The parameter update module is used to update the network parameters of the prediction model network based on the data in the training data pool using a policy optimization algorithm.
[0019] In some embodiments of the present invention, the prediction model network includes: The representation function is used to map the agent's current observation history into the individual's latent state vector; A dynamic function is used to derive the potential state vector of an individual at the next moment based on the individual's potential state vector and future actions; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
[0020] According to some embodiments of the present invention, the prediction model network includes: The representation function is used to map the agent's current observation history to an individual potential state vector; Communication functions are used to generate collaborative features for each agent through an attention mechanism; Dynamic functions are used to derive the next local latent state vector based on individual state vectors, future actions, and communication features; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
[0021] In some embodiments of the present invention, the information stored for each edge in the multi-intelligent tree search algorithm in the search module is as follows: N(s, a), P (s, a), Q(s, a), R(s, a), S(s, a), Where s represents the state information of the policy interaction environment at the current moment, a represents the action taken by each agent, N represents the number of times the state-action pair is accessed, P represents the policy probability distribution, Q represents the state value, R represents the reward value, and S represents the predicted state at the next moment. Attached Figure Description
[0022] Figure 1 This is a flowchart of a strategy decision-making method based on a world model according to an embodiment of the present invention; Figure 2 This is a flowchart of a multi-agent reinforcement learning method that incorporates a world model according to an embodiment of the present invention. Detailed Implementation
[0023] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments.
[0024] The steps described in the specification and the flowcharts in the accompanying drawings of this invention are not necessarily to be strictly followed according to the step numbers; the execution order of the steps can be changed. Furthermore, certain steps can be omitted, multiple steps can be combined into one step, and / or one step can be broken down into multiple steps.
[0025] This invention aims to combine world model planning with multi-agent reinforcement learning algorithms, improving the sample efficiency of multi-agent reinforcement learning algorithms in the policy interaction domain by employing a world model-based approach. However, integrating planning and search methods into multi-agent systems faces significant challenges, and no previous work has combined the two. Therefore, this invention is pioneering in this field. Policy activity can be understood as a game scenario involving multiple agents and policy interactions, such as chess.
[0026] The overall framework of this invention is as follows Figure 1 As shown. The strategy interaction strategy decision-making method based on a world model according to an embodiment of the present invention includes: S10, acquire situational information during strategic interactions; S20 uses a prediction model network improved based on the Muzero model network to generate latent vector representations based on situational information; It should be noted that the existing Muzero algorithm network structure cannot be directly applied to multi-agent inference environments involving policy interactions. Therefore, this invention, referencing the Muzero algorithm network structure, proposes a new network structure for specific application scenarios of policy interactions.
[0027] The original MuZero algorithm model includes: Representation function: Transforms observations (e.g., images) into hidden states; Dynamics function: updates the hidden state iteratively, taking the previous hidden state and the assumed next action as input, and outputting the new hidden state; Prediction function: predicts the policy (e.g., the action to be taken) and the state value (a good or bad assessment of the state) based on the hidden state.
[0028] In some embodiments of the present invention, in step S20, the improved prediction model network includes: The representation function is used to map the agent's current observation history into the individual's latent state vector; A dynamic function is used to derive the potential state vector of an individual at the next moment based on the individual's potential state vector and future actions; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
[0029] In other embodiments of the present invention, the improved prediction model network in step S20 includes: The representation function is used to map the agent's current observation history to an individual potential state vector; Communication functions are used to generate collaborative features for each agent through an attention mechanism; Dynamic functions are used to derive the next local latent state vector based on individual state vectors, future actions, and communication features; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
[0030] S30 selects and determines decision actions based on latent vector representations using a multi-agent tree search algorithm and outputs the results to the agents in the policy interaction.
[0031] In some embodiments of the present invention, the information stored for each edge in the multi-intelligent tree search algorithm in step S30 is as follows: N(s, a), P (s, a), Q(s, a), R(s, a), S(s, a), Where s represents the state information of the policy interaction environment at the current moment, a represents the action taken by each agent, N represents the number of times the state-action pair is accessed, P represents the policy probability distribution, Q represents the state value, R represents the reward value, and S represents the predicted state at the next moment.
[0032] According to some embodiments of the present invention, the method further includes: S40, save strategy interaction game data to the training data pool; S50 updates the network parameters of the prediction model network based on the data in the training data pool using a policy optimization algorithm.
[0033] According to an embodiment of the present invention, a strategy interaction and strategy decision-making device based on a world model includes: an information acquisition module, a prediction model network, and a search module.
[0034] The information acquisition module is used to acquire situational information in the strategy interaction; the prediction model network adopts a prediction model network improved based on the Muzero model network, which is used to generate latent vector representations based on situational information; the search module is used to select and determine decision actions based on latent vector representations through a multi-agent tree search algorithm, and output them to the agents in the strategy interaction.
[0035] According to some embodiments of the present invention, the apparatus further includes a training data pool and a parameter update module.
[0036] The training data pool is used to store strategy interaction game data; the parameter update module is used to update the network parameters of the prediction model network based on the data in the training data pool through the strategy optimization algorithm.
[0037] The present invention has the following beneficial effects: This invention proposes a policy solution method that combines a world model in the field of policy interaction. It is the first multi-agent model algorithm that combines multi-agent reinforcement learning algorithm with MCTS world model planning and is implemented under the CTDE framework.
[0038] Compared to state-of-the-art model-free multi-agent methods, this invention proposes a novel network structure and develops two new technologies: a multi-agent tree search algorithm and a policy optimization algorithm, to improve search efficiency in large action spaces. Experiments demonstrate that, compared to existing multi-agent reinforcement learning algorithms, it achieves superior sample efficiency while maintaining the same policy performance.
[0039] The present invention will now be described in detail with reference to the accompanying drawings and two specific embodiments. It is to be understood that the following description is merely exemplary and should not be construed as a specific limitation of the present invention.
[0040] Example 1: like Figure 1 As shown, firstly, this embodiment proposes a novel prediction model network based on the Muzero algorithm network structure, which includes five key modules: Representation function: maps the agent's current observation history to the individual's latent state vector.
[0041] Dynamic function: Based on the individual's potential state vector and future actions, derive the individual's potential state vector at the next moment.
[0042] Reward prediction function: Predicts team reward from the global latent state vector and the joint action vector of all individuals.
[0043] Value prediction function: Given a global latent state vector, predict the value of the global state.
[0044] Policy prediction function: Generates the corresponding policy based on the latent state vector of each individual.
[0045] Among these, the representation function, dynamic function, and policy prediction function all operate using local information and support distributed execution; other functions handle value information and team collaboration, requiring centralized training for effective learning. It should be noted that the "local information" mentioned above can be understood as the state and actions of a single agent in policy interaction. "Centralized training" can be understood as the state and action information of all agents in policy interaction, as well as historical game information related to policy interaction.
[0046] Next, based on the aforementioned neural network structure, this embodiment proposes a tree search algorithm based on the perspective of a single agent. Here, in this embodiment, 's' represents the state information of the policy interaction environment at the current moment, 'a' represents the action taken by each agent, and the information stored in each edge is: N(s, a), P(s, a), Q(s, a), R(s, a), and S(s, a) represent: the number of times the state-action pair is visited (N), the policy (action) probability distribution (P), the state value (Q), the reward value (R), and the next state predicted by the state transition function (S). Details of the tree search algorithm are as follows: 1. Selection. Starting from the root node, recursively select the "most valuable" child node until a leaf node (i.e., a node without child nodes) is found. The selection strategy is based on a balance between the node's value, prior probability, and number of visits. Specifically, the action with the highest score is selected using the UCB (Upper Confidence Bound) formula below, where c is a parameter that can be adjusted.
[0047] ; 2. Expansion. The search continues using the selection method described above until a leaf node is encountered. The leaf node is added to the search tree, and the available actions for the current state are determined by the environment's mechanisms. These available actions are then used to expand the child nodes, while simultaneously initializing their information.
[0048] 3. Evaluation. Estimate the action probabilities and state values of the available actions using the policy network and value network, and store the action probabilities on the corresponding edges. Estimate the probability distribution and state value of the set of available actions corresponding to the leaf nodes using the policy network and value network, and store each action probability value P(s, a) on the corresponding edge.
[0049] 4. Backtracking. Near the end of the k-th simulation, update V and N on the search path. The evaluation value is backpropagated back to all nodes in the search path. For each visited node, its visit count is incremented by 1, and its value is updated to the average value of the k simulations already performed.
[0050] The above process is repeated multiple times (k times). After the tree search is completed, the root node returns the set of visit counts for state-action pairs, N(s, a). Then, an action is selected based on the visit counts of the root node's child nodes. Specifically, it selects the action with the highest visit count or samples from the distribution derived from N(s, a) to obtain the final action to be executed.
[0051] Finally, this embodiment proposes a policy optimization algorithm. After a round of policy interaction ends, the sample data is stored in the experience replay pool, and samples are taken from it to train the network results proposed in the first step of this embodiment. The value network is optimized using mean squared error loss, and the policy network is optimized using cross-entropy loss. In this embodiment, the reward prediction network is trained by minimizing the error between the predicted reward and the actual observed reward, and the dynamic network is trained by minimizing the error between the predicted next time step state and the actual observed next time step state. The representation network updates its parameters synchronously during the above training process. As these networks are continuously improved, the training samples are also continuously improved, thereby gradually enhancing the capabilities of the trained agent.
[0052] In the strategic interaction scenario, this embodiment selects a complete, interconnected water network of paddy fields as the scenario, with the opponent being 109. A 50-core CPU and an A100 80G graphics card are used as the training environment. Each evaluation experiment involves 100 rounds against the 109 opponent to eliminate randomness. Experimental results show that, with only 40,000 training steps, the traditional multi-agent reinforcement learning algorithms mapp and qmix have a win rate of around 54%, while the algorithm in this embodiment achieves a win rate of 77%, improving performance by approximately 50%.
[0053] In summary, this embodiment aims to improve the sampling efficiency of MARL by adopting a model-based approach. There is currently no work that combines the two, so this embodiment is a pioneering work in this field.
[0054] This embodiment first proposes a novel network structure comprising five key modules, combining multi-agent reinforcement learning with a world model. Next, it proposes a tree search algorithm within the multi-agent reinforcement learning field. This algorithm, in conjunction with the neural network mentioned in the first point, performs policy search and improvement. Finally, this embodiment proposes a policy optimization algorithm to enhance the training efficiency of the neural network.
[0055] Example 2: like Figure 2 As shown, in this embodiment, a novel network structure is first proposed, which includes six key modules: Representation function hθ : The current observation history of agent i o ≤ ti Mapped to individual potential states st ,0 i .
[0056] Communication functions eθ Attention mechanisms are used to assist each agent. i Generate collaborative features et , ki .
[0057] Dynamic functions gθ Based on individual status st , ki Future Actions at + ki and communication characteristics et , ki Derivation of the next local potential state st , k +1 i .
[0058] Reward prediction function: from global hidden state st , k =( st , k 1,…, st , kN ) and joint actions at + k =( at + k 1,…, at + kN Predict team rewards rt , k .
[0059] Value prediction function Vθ Predict each global hidden state st , k value vt , k .
[0060] Policy prediction function P θ: Based on individual state st , ki Generate corresponding strategies pt , ki .
[0061] In this context, it represents a function.hθ Dynamic functions gθ and policy prediction function Pθ All functions use local information to run and support distributed execution; other functions handle valuable information and team collaboration, and require centralized training for effective learning.
[0062] Next, given that the algorithm learns a deterministic world model, this embodiment designs an optimistic search algorithm to better utilize the model's characteristics. In previous research, the MCTS selection phase used two metrics (value score and exploration reward) to evaluate interest in specific actions. In deterministic models, mean estimation appears overly conservative. Analogous to the multi-armed slot machine problem, if pulling a particular arm always produces a deterministic result, there is no need to repeatedly sample the same arm and calculate the average as in the UCB algorithm. Similarly, in the tree-based version of the UCT algorithm, when the environment is deterministic, a more optimistic estimate can replace the average of all simulated values in the subtree. Therefore, based on the deterministic model, this embodiment focuses on managing model generalization error rather than environmental randomness error, and designs a value score calculation method that emphasizes optimism: For each node s, define: ; Pick quantiles value.
[0063] Calculate the weighted average: ; Calculate the optimistic advantage: ; To utilize the value information calculated by the optimistic search algorithm, this embodiment proposes an advantage-weighted strategy optimization algorithm that incorporates the optimistic advantage into the behavioral cloning loss: ; Where θ represents the parameters of the learning model. Network prediction strategies to be improved To support action subsets T ( s The search strategy for ) Let α be the optimistic advantage derived from OS(M), and α > 0 be a hyperparameter controlling the degree of optimism. Theoretically, advantage-weighted strategy optimization can be viewed as... and The cross-entropy loss between them, where It is a nonparametric solution to the following constrained optimization problem: ; In summary, this embodiment proposes a novel network architecture comprising six key modules, combining multi-agent reinforcement learning with a world model. Two training optimization algorithms—an optimistic search algorithm and a dominance-weighted policy optimization algorithm—are also proposed to improve the training efficiency of the neural network.
[0064] Through the description of specific embodiments, a more in-depth and specific understanding should be gained of the technical means and effects adopted by the present invention to achieve the intended purpose. However, the accompanying drawings are only provided for reference and illustration and are not intended to limit the present invention.
Claims
1. A strategy decision-making method based on a world model, characterized in that, include: S10, acquire situational information during strategic interactions; S20, a prediction model network based on the Muzero model network is used to generate a latent vector representation based on the situation information; S30, the decision action is selected and determined based on the latent vector representation using a multi-agent tree search algorithm, and then output to the agents in the policy interaction.
2. The strategy decision-making method based on a world model according to claim 1, characterized in that, The method further includes: S40, save strategy interaction game data to the training data pool; S50, the network parameters of the prediction model network are updated based on the data in the training data pool using a policy optimization algorithm.
3. The strategy decision-making method based on a world model according to claim 1, characterized in that, The prediction model network in step S20 includes: The representation function is used to map the agent's current observation history into the individual's latent state vector; A dynamic function is used to derive the potential state vector of an individual at the next moment based on the individual's potential state vector and future actions; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
4. The strategy decision-making method based on a world model according to claim 1, characterized in that, The prediction model network in step S20 includes: The representation function is used to map the agent's current observation history to an individual potential state vector; Communication functions are used to generate collaborative features for each agent through an attention mechanism; Dynamic functions are used to derive the next local latent state vector based on individual state vectors, future actions, and communication features; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
5. The strategy decision-making method based on a world model according to claim 1, characterized in that, The information stored for each edge in the multi-intelligent tree search algorithm in step S30 is as follows: N(s, a), P (s, a), Q(s, a), R(s, a), S(s, a), Where s represents the state information of the policy interaction environment at the current moment, a represents the action taken by each agent, N represents the number of times the state-action pair is accessed, P represents the policy probability distribution, Q represents the state value, R represents the reward value, and S represents the predicted state at the next moment.
6. A strategy decision-making device based on a world model, characterized in that, include: The information acquisition module is used to acquire situational information during strategy interactions; The prediction model network, which is an improved prediction model network based on the Muzero model network, is used to generate latent vector representations based on the situation information. The search module is used to select and determine decision actions based on the latent vector representation using a multi-agent tree search algorithm, and output the results to the agents in the policy interaction.
7. The strategy decision-making device based on a world model according to claim 6, characterized in that, The device further includes: The training data pool is used to store strategy interaction game data; The parameter update module is used to update the network parameters of the prediction model network based on the data in the training data pool using a policy optimization algorithm.
8. The strategy decision-making device based on a world model according to claim 6, characterized in that, The prediction model network includes: The representation function is used to map the agent's current observation history into the individual's latent state vector; A dynamic function is used to derive the potential state vector of an individual at the next moment based on the individual's potential state vector and future actions; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
9. The strategy decision-making device based on a world model according to claim 6, characterized in that, The prediction model network includes: The representation function is used to map the agent's current observation history to an individual potential state vector; Communication functions are used to generate collaborative features for each agent through an attention mechanism; Dynamic functions are used to derive the next local latent state vector based on individual state vectors, future actions, and communication features; A reward prediction function is used to predict team rewards from the global latent state vector and the joint action vector of all individuals. The value prediction function is used to predict the value of the global state from the input global potential state vector. The policy prediction function is used to generate a corresponding policy based on the potential state vector of each individual.
10. The strategy decision-making device based on a world model according to claim 6, characterized in that, The information stored for each edge in the multi-intelligent tree search algorithm in the search module is as follows: N(s, a), P (s, a), Q(s, a), R(s, a), S(s, a), Where s represents the state information of the policy interaction environment at the current moment, a represents the action taken by each agent, N represents the number of times the state-action pair is accessed, P represents the policy probability distribution, Q represents the state value, R represents the reward value, and S represents the predicted state at the next moment.