This invention provides a heterogeneous wargame
simulation decision-making method, apparatus, device, and storage medium based on LLM-MARL collaborative driving. By extracting features and quantifying military terminology from the
battlefield situation, and generating structured situation text according to a protocol template, the LLM strategic planner outputs a tactical intent vector, achieving high-level
semantic mapping to a
tensor space. Using the tactical intent vector and
local agent observations as query and key-value sources respectively, an intent-guided feature is generated via a cross-
attention network, driving the policy network to output actions, thus guiding the underlying actions based on strategic intent. The actions and observation trajectories are mapped to a shared
semantic space, and joint training combining consistency loss and environmental rewards constrains the convergence of tactical trajectories towards strategic intent. Furthermore, reflective feedback is generated based on execution deviation, driving the LLM to update the tactical intent vector, forming a two-way collaborative optimization
closed loop.