Multi-domain combat dynamic simulation system based on deep reinforcement learning
By constructing a three-layer nested dynamic architecture and multi-domain combat data processing technology, the complexity of multi-domain combat systems has been solved, enabling efficient simulation and training of multi-domain combat strategies and improving the accuracy of combat decision-making and training effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2026-04-03
AI Technical Summary
Existing multi-domain combat dynamic simulation systems based on deep reinforcement learning are ill-suited to the complexities of multi-domain operations. This results in complex system operation logic, limited agent behavior, and insufficient data processing. These systems are unable to effectively handle complex data and system states, and cannot identify self-organizing emergent phenomena, thus affecting the accuracy of combat strategies and training effectiveness.
It adopts a three-layer nested dynamic architecture of environmental perception, adversarial game, and emergent mining, including a dynamic adaptive adversarial network, a chaotic data weaver, and an emergent capture and evolution guide. Combined with a two-way feedback mechanism, a cross-activation mechanism, and adaptive adjustment nodes, it realizes real-time interaction and cross-level collaboration of multi-domain information. Through a hybrid mechanism of inverse reinforcement learning and generative adversarial network and topology data analysis technology, it monitors and processes multi-domain combat data in real time.
It significantly improves the accuracy of multi-domain linkage simulation, shortens decision response time, enhances the decision-making flexibility of intelligent agents, optimizes dynamic data processing capabilities, manages chaotic edge states, improves the decision-making ability and training effectiveness of combat personnel, and promotes the upgrading of military training models.
Smart Images

Figure CN120930487B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of military automation technology, specifically to a multi-domain combat dynamic simulation system based on deep reinforcement learning. Background Technology
[0002] Multi-domain operational dynamic simulation systems based on deep reinforcement learning have become a key technology for addressing modern complex warfare. Multi-domain operations integrate land, sea, air, space, and cyber domains, and the battlefield environment changes rapidly. Traditional rule-based or simple data statistics-based simulation methods are insufficient to simulate complex adversarial scenarios under multi-domain coordination, and cannot meet the needs of modern military operations for pre-research of operational plans and tactical verification. Deep reinforcement learning, with its ability to learn through interaction between agents and the environment, offers a new direction for multi-domain operational simulations. However, in practical applications, existing technologies struggle to adapt to the complexity of multi-domain operations, preventing the full realization of the system's potential value.
[0003] Existing technologies primarily focus on optimizing conventional strategies and simulating known combat modes using deep reinforcement learning in multi-domain combat simulations. At the agent behavior control level, pre-defined rules are commonly used to establish fixed decision-making logic for the agent. For example, in air combat simulations, rules for maneuvering and attacking enemy aircraft are predefined for the fighter agent. This approach severely limits the agent's flexibility, making it difficult to adapt to unexpected and unconventional situations in multi-domain combat. Regarding data processing, the focus is mainly on structuring historical combat data and constructing specific datasets for training, such as organizing ship navigation trajectories and weapon launch data in naval battles into fixed formats. However, multi-domain combat data is characterized by real-time, dynamic, and diverse features. The continuous influx of new data leads to frequent changes in data distribution and characteristics. Existing structured processing methods are simply unable to effectively cope with the dynamic complexity of the data, and deep reinforcement learning models struggle to capture the complex potential relationships between data points.
[0004] Because of the significant differences in combat units, rules, and data types across various domains in multi-domain operations, integrating them into a deep reinforcement learning simulation system directly results in an extremely complex system structure and operational logic. Taking land-air joint operations as an example, the advance pace of ground troops and the response speed of air support differ in time and space, and their combat rules also differ. This multi-dimensional interplay of differences makes the interactions between agents within the system intricately complex. As the number of agents increases, the amount of data expands, and the interactions become increasingly complex, the system gradually approaches a state of near-chaos. In this state, while unexpected self-organizing emergent phenomena such as collaborative combat patterns may occur, the limitations on agent behavior and insufficient data processing in the early stages further exacerbate the system's approach to the edge of chaos. Furthermore, current technologies lack both the means to monitor the system entering a state of near-chaos and the ability to identify and analyze these self-organizing emergent phenomena. For example, in large-scale multi-domain joint simulations, the failure to detect the system entering a state of near-chaos has resulted in the undetected discovery of unexpectedly efficient collaborative strategies among agents, strategies that could potentially turn the tide of battle in real combat.
[0005] If the aforementioned problems arising from the complexity of multi-domain operations are not resolved, multi-domain operational dynamic simulation systems based on deep reinforcement learning will struggle to overcome existing developmental bottlenecks. The inability to effectively handle complex data and system states will lead to increased deviations between simulation results and real combat scenarios, making it difficult to generate operational strategies with practical guidance. The inability to uncover self-organizing emergent phenomena in chaotic edge states means missing opportunities to discover new operational theories and tactics, hindering the innovative development of military operational theory. In terms of tactical training, because the simulation system cannot simulate the complexity and uncertainty of the real battlefield, combat personnel will struggle to receive adequate and effective training, failing to improve their ability to respond to emergencies and complex situations, thus impacting the overall combat level and actual combat capability of the troops.
[0006] In view of this, a multi-domain combat dynamic simulation system based on deep reinforcement learning is provided to overcome the above problems. Summary of the Invention
[0007] The purpose of this invention is to provide a multi-domain combat dynamic simulation system based on deep reinforcement learning to solve the problems mentioned in the background art.
[0008] To address the aforementioned technical problems, this invention provides a multi-domain combat dynamic simulation system based on deep reinforcement learning, comprising a three-layer nested dynamic architecture of environmental perception, adversarial game, and emergent mining. This three-layer nested dynamic architecture deploys a bidirectional feedback mechanism, a cross-activation mechanism, and adaptive adjustment nodes. The three-layer nested dynamic architecture includes a dynamic adaptive adversarial network, a chaotic data weaver, and an emergent capture and evolutionary guide; wherein:
[0009] The dynamic adaptive adversarial network consists of multiple heterogeneous adversarial subnets, each subnet corresponding to a combat domain. The subnets interact with each other through a dual-mode adversarial-cooperative interaction, and the subnets adopt a hybrid mechanism of inverse reinforcement learning and generative adversarial network.
[0010] Chaotic data weavers treat multi-source heterogeneous data as a dynamic data cloud and use topological data analysis techniques to capture the topological structure and dynamic change trends of data in high-dimensional space.
[0011] Emergent Capture and Evolution Guide is based on the principle of breaking the information cocoon. It monitors the interaction information and data flow of agents in the system in real time. When it detects that the system is close to the edge of chaos, it intervenes through a three-stage operation of information perturbation, pattern amplification and value assessment.
[0012] Furthermore, the two-way feedback mechanism deploys intelligent information routers between levels. The intelligent information routers have built-in time-series correlation analysis modules and priority judgment algorithms to achieve real-time information interaction and efficient processing.
[0013] Furthermore, the cross-activation mechanism adopts an event-rule-resource driven model. When the emergence mining layer identifies a potential combat mode, it generates an activation command that includes event triggering conditions, rule constraints, and resource requirements, thereby activating the simulation module of the adversarial game layer and the dynamic monitoring task scheduler of the environment perception layer.
[0014] Furthermore, the adaptive adjustment node integrates a multi-indicator fusion evaluation engine and a control strategy library. The multi-indicator fusion evaluation engine uses an algorithm that combines principal component analysis and fuzzy logic to reduce the monitoring indicators to three comprehensive dimensions: system load, decision complexity, and data correlation, which are used to judge the system status and control resources.
[0015] Furthermore, an environmental mutation detector is embedded in the heterogeneous adversarial subnet. Based on the sliding window algorithm, the rate of change of environmental state characteristics is calculated in real time. When the rate of change exceeds a preset threshold, the subnet is triggered to switch from cooperative mode to adversarial mode.
[0016] Furthermore, in the hybrid mechanism of inverse reinforcement learning and generative adversarial network, the generator introduces a policy fragment memory pool, and the discriminator adopts a multi-scale evaluation method to comprehensively score the policy from the time, space and resource scales.
[0017] Furthermore, the chaotic data weaver employs an incremental topology analysis method to construct a data feature hash table. When new data flows in, it only calculates the connection relationship between the new data point and the existing topology to update the topology graph.
[0018] Furthermore, the chaotic data weaver introduces a knowledge graph-guided weaving weight allocation mechanism, which assigns semantic weights to topological connections based on the strength of associations between data entities in the knowledge graph, thus forming a data weave.
[0019] Furthermore, the information perturbation wave generation module of the emergence capture and evolution guide adopts an adversarial example-based perturbation generator, which generates adversarial examples against the agent's decision-making model as information perturbation waves through adversarial training; the pattern amplification and value assessment module adopts a social network propagation simulation mechanism, combined with combat simulation stress testing to evaluate the strategy value.
[0020] Compared with the prior art, the beneficial effects of the present invention are:
[0021] 1. Enhanced multi-domain linkage simulation capabilities:
[0022] A three-layer nested dynamic architecture of "environmental perception - adversarial game - emergent mining" is constructed, breaking through the traditional hierarchical fragmentation and one-way transmission mode, forming the core logic of "dynamic collaboration - intelligent regulation - evolution drive", realizing real-time interaction and cross-level collaboration of multi-domain information, significantly reducing the simulation error of multi-domain linkage, and significantly improving the consistency between strategy and actual combat scenario.
[0023] The system deploys a two-way feedback mechanism, sets up intelligent information routers between levels, and incorporates a time-series correlation analysis module and priority judgment algorithm to achieve real-time information interaction and efficient processing, significantly shortening decision response time. For example, when simulating multiple waves of enemy attacks, the system can plan interception strategies in advance and dynamically optimize them, generating composite strategies that traditional systems cannot achieve.
[0024] 2. Enhanced decision-making flexibility of intelligent agents:
[0025] The Dynamic Adaptive Adversarial Network (DAAN) consists of multiple heterogeneous adversarial subnetworks that interact with each other through a dual-mode "adversarial-cooperative" interaction. Internally, it employs a hybrid mechanism of "inverse reinforcement learning-generative adversarial network (RL-GAN)". When the environmental mutation detector detects that the rate of change of environmental state features exceeds a threshold, the subnetwork automatically switches from cooperative mode to adversarial mode. Through competition, new strategies are generated and optimized collaboratively, breaking the limitations of preset rules on the agent and significantly improving the decision-making accuracy of the agent in unknown and sudden scenarios. It also efficiently and autonomously generates cross-domain composite strategies.
[0026] The generator introduces a policy fragment memory pool, and the discriminator adopts a multi-scale evaluation method to comprehensively score the policy from the time, space and resource scales, guiding the generation of better policies. This enables the agent to quickly adjust its decision-making mode according to changes in the battlefield and generate a coherent chain of response strategies in complex scenarios such as electronic warfare.
[0027] 3. Optimized dynamic data processing capabilities:
[0028] Chaotic Data Weaver (CDW) treats multi-source heterogeneous data as a dynamic "data cloud," utilizing Topological Data Analysis (TDA) technology to capture the topological structure and dynamic trends of the data in high-dimensional space. Employing an incremental topological analysis method, it constructs a data feature hash table. When new data flows in, it only calculates the connection relationships between the new data points and the existing topological structure to update the topological graph, avoiding the repetitive work of traditional full-scale computation, significantly improving data processing efficiency and reducing time complexity.
[0029] By introducing a knowledge graph-guided weaving weight allocation mechanism, semantic weights are assigned to topological connections based on the strength of associations between data entities in the knowledge graph, forming a more semantically deep "data weave" that deeply mines potential associations between data, significantly increasing the amount of effective information mined.
[0030] 4. Breakthrough in chaotic edge state management capabilities:
[0031] Emergent Capture and Evolutionary Guide (ECEG) is based on the principle of "breaking the information cocoon." It monitors the system status in real time and intervenes through a three-stage process of "information perturbation, pattern amplification, and value assessment" when it detects that the system is approaching the edge of chaos. The information perturbation wave generation module uses a perturbation generator based on adversarial examples to generate adversarial examples that are injected into the system, breaking the inherent decision-making patterns of the agent and increasing the probability of policy emergence.
[0032] The pattern amplification and value assessment module adopts a social network propagation simulation mechanism, combined with combat simulation stress testing, to optimize the capture accuracy of emerging patterns. This enables the system to successfully capture a variety of valuable self-organized emerging phenomena during simulations, fostering the prototype of a new combat theory and filling the gap in traditional technology for monitoring and utilizing chaotic edge states.
[0033] 5. Multi-module collaborative optimization enhances practical value:
[0034] DAAN generates diverse combat strategies, CDW provides precise intelligence support, and ECEG simulates chaotic scenarios. The collaborative work of multiple modules enables the simulation system to realistically reproduce the complexity of the battlefield, significantly improving the decision-making ability and emergency response speed of combat personnel in complex scenarios, promoting the upgrading of military training models, and providing a more realistic and effective simulation environment for the pre-research of combat plans and tactical verification, making the pre-research plans more in line with actual combat needs. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the multi-domain combat dynamic simulation system based on deep reinforcement learning, as described in this invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Please see Figure 1 The present invention provides a technical solution:
[0038] See Figure 1 As shown, an embodiment of a multi-domain combat dynamic simulation system based on deep reinforcement learning is presented:
[0039] I. System Overall Architecture:
[0040] A three-layer nested dynamic architecture of "environmental perception-adversarial game-emergent mining" is constructed. This architecture breaks through the hierarchical fragmentation and unidirectional transmission mode of traditional inference systems. With "dynamic collaboration-intelligent regulation-evolutionary drive" as the core logic, it integrates the principles of biological neural networks and ecosystem self-regulation to form a highly adaptable complex system.
[0041] Traditional simulation systems break down multi-domain operations into independent modules, essentially using a mechanical approach to handle complex systems, leading to information lag and rigid decision-making. This architecture's breakthrough is reflected in three levels:
[0042] Two-way feedback mechanism: Drawing inspiration from the two-way signal transmission characteristics of neurons in biological neural networks, this architecture breaks through the unidirectional "data input-decision output" link of traditional systems. For example, in real neural networks, after receiving a signal, a neuron not only transmits it to downstream neurons but also regulates the activity of upstream neurons through feedback connections. This architecture applies this principle to combat simulation, enabling information acquired by the environmental perception layer (such as satellite reconnaissance of enemy troop movements) to trigger real-time adjustments to combat strategies at the adversarial game layer. Simultaneously, changes in the battlefield situation after strategy execution (such as enemy countermeasures) are fed back to the perception layer, forming a dynamic closed loop of "perception-decision-feedback," thus solving the problems of information lag and inability to respond to sudden changes in traditional systems.
[0043] Cross-activation mechanism: Inspired by the co-evolution of species in ecosystems, a cross-level cross-activation mechanism is designed. In ecosystems, the interaction between predators and prey drives both to continuously evolve new survival strategies. Similarly, when the emergence mining layer discovers new combat modes (such as the coordinated attack of drone swarms and electronic warfare units), the cross-activation mechanism will activate the adversarial game layer to conduct strategy simulation verification, and guide the environmental awareness layer to specifically monitor relevant combat elements (such as enemy air defense system responses and electromagnetic spectrum changes), breaking the limitations of independent operation and inability to co-evolve at each level in traditional systems.
[0044] Adaptive Adjustment Nodes: Unlike the human body's automatic regulation system, which relies solely on a single physiological indicator (such as body temperature or blood pressure), the adaptive adjustment nodes in this architecture employ a composite mechanism of "multi-dimensional state assessment - dynamic threshold calibration - hierarchical control strategy." By monitoring 12 core indicators, including data flow, information entropy, interaction conflict index, and resource utilization, a dynamic assessment model is constructed. This model automatically adjusts the control strategy based on the current operational phase (offense, defense, stalemate), achieving intelligent allocation of system resources and preventing traditional systems from falling into chaos due to information overload or resource mismatch.
[0045] Specifically:
[0046] A two-way feedback mechanism is implemented: "Intelligent information routers" are deployed between layers, each with a built-in time-series correlation analysis module and priority judgment algorithm. For example, when the environmental perception layer receives a warning of an enemy missile launch, the time-series correlation analysis module quickly matches historical patterns of similar events, and the priority judgment algorithm determines the information transmission path and speed based on the threat level (e.g., missile range, target type). High-threat information is transmitted directly to the core decision-making unit of the adversarial game layer with millisecond-level latency, while simultaneously pushing related battlefield environmental data (e.g., weather conditions, terrain features) to assist in generating interception strategies. The effectiveness data after strategy execution (e.g., interception success rate, enemy secondary attack trends) is then fed back to the perception layer via a reverse link to optimize subsequent monitoring parameters.
[0047] Cross-activation mechanism: Employing an "event-rule-resource" driven model. When the emergence mining layer identifies a potential combat mode (such as the timing coordination of multi-domain electronic jamming and fire strikes), it generates an activation command that includes event triggering conditions (such as enemy radar activation, communication frequency band switching), rule constraints (such as the synergistic relationship between jamming intensity and fire coverage), and resource requirements (such as the number of UAVs and electromagnetic spectrum resources). This command activates the simulation module of the adversarial game layer on the one hand, using virtual simulation to verify the feasibility of the mode; on the other hand, it triggers the "dynamic monitoring task scheduler" of the environmental perception layer, adjusting the sensor acquisition frequency and coverage, such as increasing the sampling rate of enemy electronic equipment signals, providing more accurate data support for mode optimization.
[0048] Adaptive Adjustment Nodes: Each node integrates a "multi-indicator fusion evaluation engine" and a "control strategy library." The evaluation engine employs an algorithm combining principal component analysis (PCA) and fuzzy logic to reduce the 12 monitoring indicators to three comprehensive dimensions: system load, decision complexity, and data correlation. Fuzzy logic is used to determine the system state (e.g., normal, warning, crisis). When the system enters a warning state (e.g., data flow exceeds a threshold and information entropy surges), the node selects a "data diversion-priority calculation-interaction circuit breaker" combination strategy from the control strategy library: non-critical data is temporarily stored in an edge cache, core data such as combat instructions and target identification are prioritized for transmission, and low-value interaction channels between agents are temporarily cut off, allowing system resources to focus on critical decisions and avoiding decision paralysis due to information congestion.
[0049] Traditional systems suffer from significant delays in information transmission and decision-making processes in multi-domain contingency scenarios, making it difficult to cope with dynamically changing battlefield situations. This architecture, through bidirectional feedback and intelligent routing mechanisms, enables real-time information interaction and efficient processing, significantly shortening decision response time. When simulating multi-wave enemy attacks, the system can pre-plan interception strategies and dynamically optimize them based on changes in enemy tactics, generating composite strategies that traditional systems cannot achieve, significantly improving combat response capabilities.
[0050] To more clearly demonstrate the effect of the two-way feedback mechanism on improving decision-making response speed, a mathematical model is constructed for quantitative analysis:
[0051] Let the decision response time of the traditional system be T. c Includes information transmission time t c1 Decision calculation time t c2 ,Right now:
[0052] T c =t c1 +t c2 ;
[0053] The decision response time of this architecture is T. n Information transmission time t n1 Decision calculation time t n2 Due to the two-way feedback and intelligent routing mechanism,
[0054] t n1 =α×t c1 ;
[0055] (0 < α < 1, representing the coefficient for improving information transmission efficiency), then:
[0056] t n2 =β×t c2 ;
[0057] (0 < β < 1, representing the coefficient for improving decision computation efficiency), then:
[0058] T n =t n1 +t n2 =α×t c1 +β×t c2 .
[0059] Improvement rate in decision response time:
[0060]
[0061] in:
[0062] T c Total decision-making response time of traditional systems;
[0063] t c1 Information transmission time in traditional systems;
[0064] t c2 Traditional system decision-making computation time;
[0065] T n Total decision response time for this architecture;
[0066] t n1 Information transmission time in this architecture;
[0067] t n2 The decision-making time for this architecture;
[0068] α: Information transmission efficiency improvement coefficient, reflecting the degree to which two-way feedback and intelligent routing accelerate information transmission;
[0069] β: The ratio of improvement in decision computation efficiency, reflecting the improvement in decision computation efficiency brought about by information optimization;
[0070] η: The percentage improvement in decision response time, dimensionless.
[0071] Traditional simulation systems operate independently at each level, resulting in low coordination efficiency. This architecture's cross-activation mechanism enables deep collaboration across levels. When the emergence mining layer discovers a new combat mode, it can quickly activate the adversarial game layer for simulation verification and coordinate with the environmental awareness layer to adjust monitoring priorities. This mechanism effectively improves the system's coordination efficiency, fosters new combat theories, and significantly enhances the combat effectiveness of troops in complex scenarios.
[0072] The effect of cross-activation mechanism on improving the efficiency of cross-level collaboration in the system can be demonstrated by establishing a collaborative efficiency evaluation formula:
[0073] Let the coordination efficiency of the traditional system be E. cThe amount of collaborative tasks completed is Q. c The time consumed is T. c3 ,but:
[0074]
[0075] The collaborative efficiency of this architecture is E. n The amount of collaborative tasks completed is Q. n (Because collaboration is more efficient,)
[0076] Q n =γ×Q c ,
[0077] γ > 1 (representing the collaborative task workload increase factor), with a time consumption of T. n3 (T n3 =δ×T c3 (where 0 < δ < 1), representing the coordination time reduction coefficient, then:
[0078]
[0079] Percentage increase in collaborative efficiency:
[0080]
[0081] in:
[0082] E c Traditional system collaboration efficiency, unit: task volume / second;
[0083] Q c The number of collaborative tasks completed by a traditional system, in units of: tasks.
[0084] T c3 Time taken by a traditional system to complete a collaborative task, in seconds;
[0085] E n The collaborative efficiency of this architecture is expressed in terms of task volume per second.
[0086] Q n The number of collaborative tasks completed by this architecture is expressed in units of: tasks.
[0087] T n3 The time taken for this architecture to complete the collaborative task, in seconds;
[0088] γ: Coefficient of improvement in collaborative task workload, reflecting the degree of enhancement in collaborative task processing capacity brought about by the cross-activation mechanism;
[0089] δ: Coordination time reduction coefficient, reflecting the optimization effect of the cross-activation mechanism on the execution time of collaborative tasks;
[0090] θ: The percentage increase in collaborative efficiency, dimensionless.
[0091] In highly complex multi-domain adversarial scenarios, traditional systems often suffer from decision-making conflicts among agents due to data overload, leading to system instability. This architecture, through adaptive node adjustment, dynamically regulates system resources, significantly reducing the decision-making conflict rate and effectively improving system stability. Furthermore, by leveraging the uncertainty of chaotic edges, the system can spontaneously form novel collaborative mechanisms, exhibiting self-organizing capabilities not found in traditional systems.
[0092] To verify the effectiveness of adaptive adjustment nodes in reducing the system's decision conflict rate, a formula for calculating the decision conflict rate is constructed to demonstrate this:
[0093] Let the decision-making conflict rate of the traditional system be R. c The number of conflicting decisions is N c The total number of decisions is M c ,but:
[0094]
[0095] The decision conflict rate of this architecture is R. n The number of conflict decisions is (This represents the coefficient for reducing the number of conflicting decisions), where the total number of decisions is M. n (Assuming the decision-making scenarios are the same)
[0096] M n =M c ),but:
[0097]
[0098] The percentage of time that decision-making conflict rate decreased:
[0099]
[0100] in:
[0101] R c : Decision conflict rate in traditional systems, dimensionless;
[0102] N c The number of conflicting decisions in traditional systems;
[0103] M c Total number of decisions in a traditional system;
[0104] R n : Decision conflict rate in this architecture;
[0105] N n Number of conflicting decisions in this architecture;
[0106] M n Total number of decisions in this architecture;
[0107] The conflict decision reduction coefficient reflects the effect of adaptive adjustment nodes on suppressing decision conflicts;
[0108] λ : The percentage decrease in decision-making conflict rate is dimensionless.
[0109] It should be noted here that:
[0110] The two-way feedback mechanism provides more accurate environmental information and policy feedback to the Dynamic Adaptive Adversarial Network (DAAN), accelerating its policy learning process and enabling it to quickly adjust its decision-making mode according to battlefield changes. In complex scenarios such as electronic warfare, DAAN can more agilely generate coherent response policy chains, improving the flexibility and effectiveness of agent decision-making.
[0111] The cross-activation mechanism guides the Chaotic Data Weaver (CDW) to conduct in-depth analysis around high-value data. In actual simulations, guided by specific combat modes, CDW can uncover deep-seated correlations between data, obtain key information that traditional analysis methods cannot discover, and provide more valuable intelligence support for combat decision-making.
[0112] Adaptive adjustment of nodes optimizes resource allocation, enhancing the performance of the Emergent Capture and Evolutionary Guide (ECEG). ECEG can more precisely control perturbation parameters, ensuring stable system operation while improving the accuracy of capturing self-organized emergent phenomena, thereby accelerating the exploration and formation of new combat theories.
[0113] II. Core Modules:
[0114] 1. Dynamic Adaptive Adversarial Network (DAAN):
[0115] In the field of intelligent agent behavior control, traditional preset rules and fixed decision-making logic are abandoned. DAAN consists of multiple heterogeneous adversarial subnets, each corresponding to a combat domain. The subnets interact through a dual-mode "adversarial-cooperative" interaction: in normal scenarios, the subnets cooperate with each other, share information, and make collaborative decisions; when encountering unconventional and sudden situations, the subnets automatically switch to adversarial mode, generate new strategic ideas through mutual competition, and then optimize and integrate the new strategies through cooperation.
[0116] To further enhance the decision-making capabilities and adaptability of DAAN, a hybrid mechanism of "Inverse Reinforcement Learning-Generative Adversarial Network (RL-GAN)" is adopted within the subnet. Inverse Reinforcement Learning (IRL) is a technique that infers expected behavioral patterns from expert demonstrations or environmental feedback. Unlike traditional forward reinforcement learning, which starts from a pre-defined reward function, IRL is more suitable for complex multi-domain combat scenarios where rewards are difficult to define precisely. Generative Adversarial Networks (GANs) consist of a generator and a discriminator, which continuously optimize the quality of the generated results through adversarial competition. In DAAN, combining inverse reinforcement learning with generative adversarial networks forms a unique RL-GAN hybrid mechanism, enabling the agent to autonomously learn and generate effective coping strategies in complex and ever-changing environments.
[0117] Specifically:
[0118] Subnet Adversarial-Cooperative Switching Mechanism: An "environmental mutation detector" is embedded in each heterogeneous adversarial subnet. This detector, based on a sliding window algorithm, calculates the rate of change of current environmental state characteristics in real time. When the rate of change exceeds a preset threshold (e.g., the characteristic change amplitude is greater than twice the average change amplitude within three consecutive time steps), the subnet is triggered to switch from cooperative mode to adversarial mode. For example, in land-air combat, if ground radar detects the sudden appearance of an enemy drone swarm, and the environment mutation detector detects a sharp change in the number, speed, and other characteristics of aerial targets, the land combat subnet and the air combat subnet automatically enter adversarial mode. This switching mechanism breaks the fixed pattern of traditional agent behavior control, enabling agents to autonomously adjust their decision-making strategies according to dynamic environmental changes.
[0119] Hybrid Mechanism Optimization Strategy: In the RL-GAN hybrid mechanism, a "policy fragment memory pool" is introduced for the generator. When generating a new policy, the generator prioritizes selecting historically valid policy fragments from the memory pool for combination. If the combined policy cannot meet the environmental requirements, a completely new policy fragment is generated through adversarial training. The discriminator adopts a "multi-scale evaluation method," which not only evaluates the matching degree between the policy and environmental requirements but also comprehensively scores it from three dimensions: time scale (the timeliness of the policy), spatial scale (the applicability of the policy in different combat areas), and resource scale (the combat resource consumption required to execute the policy), guiding the generator to generate a better policy. Taking the joint land-air encounter with new enemy electronic jamming as an example, the generator of the land combat subnet selects the "terrain concealment communication" policy fragment from the memory pool and combines it with the "frequency band hopping jamming avoidance" fragment generated by the air combat subnet. After multi-scale evaluation by the discriminator, it is optimized into a new policy of "low-altitude terrain-following flight + multi-frequency band rapid switching communication."
[0120] It should be noted here that:
[0121] Traditional rule-based agents, when faced with unexpected and unconventional situations in multi-domain operations, are like mechanical devices bound by fixed programs, struggling to adapt flexibly. DAAN completely changes this situation through a dual-mode interaction of "adversarial-cooperative" interaction in heterogeneous subnets and a hybrid RL-GAN mechanism. When encountering unprecedented electronic jamming tactics from the enemy, traditional agents can only respond in a limited way according to preset rules. However, DAAN's subnets can quickly switch from cooperative mode to adversarial mode. Each subnet autonomously explores new strategies based on its own operational domain characteristics. For example, the land combat subnet proposes concealed communication schemes using terrain characteristics, while the air combat subnet generates frequency band evasion strategies based on flight advantages. Subsequently, these strategies are integrated and optimized through cooperative mode to form a composite strategy of "low-altitude terrain-concealed communication + multi-band dynamic switching." This new decision-making mode frees agents from fixed rules, enabling them to proactively explore and generate diverse and effective response strategies in complex and ever-changing battlefield environments. This greatly improves the decision-making flexibility and effectiveness of agents in unknown scenarios, successfully solving the problem of rigid decision-making and difficulty in adapting to complex and unexpected situations in traditional agents in multi-domain operations.
[0122] Traditional structured processing methods for multi-domain operational data are like trying to mold a rapidly changing fluid with a fixed mold—unable to adapt to the real-time, dynamic, and diverse nature of the data. DAAN's RL-GAN hybrid mechanism and technologies such as the "strategy fragment memory pool" bring a completely new approach to data processing. In actual multi-domain operational simulations, new data is constantly and rapidly entering the market, and data distribution and characteristics change frequently, making it difficult for traditional methods to capture the complex potential relationships between data. DAAN's generator, when generating strategies, combines historical strategy fragments from the memory pool with real-time data characteristics to uncover the key information hidden behind the data. For example, in a multi-domain operational simulation involving land, sea, and air, when the navigation data of naval vessels, the flight trajectory data of fighter jets, and the deployment data of land forces all change simultaneously, DAAN can discover the enemy's operational intentions and behavioral patterns under specific time and space conditions from these dynamically changing data, and generate targeted strategies. This strategy generation method based on dynamic data not only effectively handles the complexity of multi-domain combat data, but also deeply mines the value of data, providing richer, more accurate, and real-time data support for combat decision-making, and solving the problem that traditional data processing methods cannot cope with the dynamic complexity of multi-domain combat data.
[0123] In multi-domain combat simulation systems, as the number of agents increases, data volume expands, and interactions between domains become increasingly complex, the system is prone to approaching a chaotic edge state. Traditional technologies are unable to effectively monitor this state or identify potentially valuable self-organizing emergent phenomena. DAAN's heterogeneous subnet adversarial-cooperative mode and adaptive policy generation mechanism provide an effective solution to this problem. When the system approaches the chaotic edge, DAAN's subnets can stimulate innovative policies through adversarial modes. These policies propagate and diffuse within the system, potentially triggering new cooperative combat modes and other self-organizing emergent phenomena. Simultaneously, the policies generated by DAAN can guide other modules to adjust their operating methods. For example, they provide a clearer data mining direction for the Chaotic Data Weaver (CDW), enabling it to focus on key policy-related data and extract information that helps stabilize the system state; and they provide rich "materials" for the Emergent Capture and Evolutionary Guide (ECEG), helping ECEG to more accurately monitor the system state and identify and analyze self-organizing emergent phenomena. For example, in large-scale multi-domain joint simulations, DAAN, through its own strategy adjustments, has prompted a new cross-domain collaborative communication strategy among agents in the system. This strategy not only stabilizes the system operation but also provides new tactical ideas for subsequent operations, successfully solving the problem that traditional systems cannot effectively monitor and utilize self-organizing emergent phenomena in chaotic edge states.
[0124] It should also be noted here that:
[0125] DAAN's powerful decision-making capabilities and flexible strategy generation mechanism provide a more realistic and effective simulation environment for operational plan pre-research and tactical verification. Previously, limited by the constraints of traditional simulation systems, operational plan pre-research was often based on limited preset scenarios and fixed rules, making it difficult to fully consider the complexity and uncertainty of multi-domain operations, resulting in significant deviations between pre-researched plans and actual operational needs. However, simulation systems based on DAAN can simulate various complex and ever-changing battlefield situations, generating a rich variety of operational strategies. Military researchers can utilize these strategies to conduct more in-depth and comprehensive verification and optimization of different operational plans, identifying problems and potential risks in advance, thereby developing operational plans that are more aligned with actual operational needs and are more feasible and effective. For example, in the pre-research of new multi-domain collaborative operational plans, the various strategies generated by DAAN can help researchers evaluate the advantages and disadvantages of plans from different perspectives, optimize the coordination methods and resource allocation between different operational domains, significantly improve the quality and efficiency of operational plan pre-research, and promote the development of military operational plan research in a more scientific and precise direction.
[0126] Traditional multi-domain combat simulation systems, unable to realistically simulate the complexity and uncertainty of the battlefield, hinder combat personnel from gaining sufficient and effective training, and fail to improve their ability to respond to emergencies and complex situations. The introduction of DAAN enables simulation systems to more realistically recreate multi-domain combat scenarios, generating various complex and ever-changing combat situations and contingencies. In such a training environment, combat personnel need to constantly face unknown challenges, applying their acquired knowledge and skills to analyze, make decisions, and respond, thereby effectively improving their judgment, decision-making, collaborative combat, and emergency response capabilities in complex battlefield environments. For example, in DAAN-based multi-domain combat simulation training, combat personnel may encounter various complex emergencies such as sudden multi-domain joint attacks by the enemy, communication disruptions, and resource shortages. By dealing with these situations, combat personnel can accumulate rich combat experience, improve their psychological resilience and adaptability. This combat-oriented training method can significantly improve the overall combat level and combat capability of combat personnel, making the troops more combat-effective in actual combat and effectively compensating for the shortcomings of traditional training methods that are disconnected from actual combat.
[0127] 2. Chaotic Data Weaver (CDW):
[0128] In response to the dynamic complexity of multi-domain operational data, CDW no longer performs structured processing on the data. It treats multi-source heterogeneous data as a dynamic "data cloud" and uses Topological Data Analysis (TDA) technology to capture the topological structure and dynamic changing trends of the data in high-dimensional space.
[0129] Specifically:
[0130] Dynamic updates to the data topology: Employing an "incremental topology analysis" method, when new data flows in, the entire data topology is not recalculated; instead, only the connections between the new data points and the existing topology are calculated. Specifically, a "data feature hash table" is constructed to quickly locate similar feature points between new and historical data, and the topology is updated based on these similarities. For example, in naval warfare data processing, newly acquired ship navigation data is quickly found using the hash table to locate historical data points with similar trajectories, and only the associated topological edges are updated, significantly improving data processing efficiency. This method avoids the repetitive calculation of large amounts of historical data in traditional data processing, effectively enhancing the system's ability to process dynamic data.
[0131] To quantify the efficiency improvement of incremental topology analysis methods in processing dynamic data, a data update time complexity model is constructed as follows:
[0132] The time complexity of traditional full topology analysis is:
[0133] T full =O(n) 2 );
[0134] The time complexity of incremental topology analysis is:
[0135] T incremental =O(k);
[0136] The efficiency improvement rate is:
[0137]
[0138] in:
[0139] T full The time complexity of traditional full-scale topology analysis is positively correlated with the square of the total number of data points.
[0140] T incremental The time complexity of incremental topology analysis depends only on the number of neighbors of the new data point.
[0141] n: Total number of data points, usually in the millions or more;
[0142] k: The number of neighboring nodes of the new data point, typically ranging from tens to hundreds (k < k). <n);
[0143] η cdw : The percentage increase in data processing efficiency, dimensionless.
[0144] Semantic Weaving Optimization: A "knowledge graph-guided weaving weight allocation" mechanism is introduced. A multi-domain operational knowledge graph, containing operational concepts, entity relationships, and other information, is pre-constructed. During the "weaving" operation on the data topology graph, semantic weights are assigned to different topological connections based on the strength of associations between data entities in the knowledge graph. For example, when associating ship navigation trajectories with meteorological data, if the knowledge graph shows a correlation between certain meteorological conditions and specific navigation evasion strategies, this association is assigned a higher semantic weight, making the system more inclined to mine such valuable association patterns and form a more semantically deep "data weave." By combining knowledge graphs with topological data analysis, deep mining from the data structure to the semantic level is achieved.
[0145] By constructing an information gain model, we demonstrate the effectiveness of knowledge graph-guided semantic weaving for deep association mining:
[0146] Let the information gain of traditional topology analysis be:
[0147]
[0148] The information gain after introducing knowledge graphs is:
[0149]
[0150] The percentage increase in information mining depth is:
[0151]
[0152] in:
[0153] G traditional Information gain in traditional topology analysis measures the improvement in information purity after data partitioning.
[0154] H(D): Information entropy of data set D, reflecting data uncertainty;
[0155] D i : The data is divided into subsets, where \(|D_i|\) is the size of the subset;
[0156] G KG-guided Information gain guided by knowledge graphs, including semantic association weights;
[0157] w j : Semantic weights of entity associations in a knowledge graph, with values ranging from [0,1];
[0158] I(E j ;C): Entity E j Mutual information with operational concept C measures the strength of semantic association;
[0159] θ cdw The proportion of information mining depth improvement, dimensionless.
[0160] 3. Emergent Capture and Evolutionary Guide (ECEG):
[0161] To address the chaotic edge state of a system, ECEG is designed based on the principle of "breaking the information cocoon". It monitors the interaction information and data flow of agents in the system in real time, and intervenes through a three-stage operation of "information perturbation - pattern amplification - value assessment" when it detects that the system is approaching the chaotic edge.
[0162] Specifically:
[0163] Information Disturbance Wave Generation: A "disturbance generator based on adversarial examples" is designed. This generator produces adversarial examples targeting agent decision-making models as information disturbance waves through adversarial training. Specifically, a subset of agent decision-making models are selected as targets. Then, through iterative optimization, adversarial examples that can mislead these models into making unconventional decisions are generated. For example, against the target recognition model of an air combat agent, false radar echo data resembling real targets but causing misjudgments can be generated as disturbance waves. Injecting these waves into the system disrupts the agent's original interaction and decision-making patterns. This method of using adversarial examples to break the agent's inherent decision-making patterns makes it possible for new strategies to emerge in the system.
[0164] By constructing a perturbation-emergence probability model, the promoting effect of adversarial examples on the emergence of novel strategies is quantified:
[0165] The emergent probability of strategies in a traditional system is:
[0166]
[0167] The emergence probability after introducing adversarial example perturbation is:
[0168] P ECEG =P traditional ·(1+λ·σ);
[0169] The increase in the emergence probability is:
[0170]
[0171] in:
[0172] P traditional The emergent probability of strategies in traditional systems is dimensionless.
[0173] N emergence The number of new strategies emerging in traditional systems;
[0174] N total The total number of strategies generated in a traditional system;
[0175] P ECEG : The probability of strategy emergence after ECEG intervention;
[0176] λ: The perturbation strength coefficient of the adversarial sample, which is controlled by the perturbation generator;
[0177] σ: The proximity of the system to the edge of chaos; the closer it is to 1, the more likely it is to emerge.
[0178] ρ eceg The probability increase of strategy emergence is dimensionless.
[0179] Pattern Amplification and Value Assessment: In the pattern amplification phase, a "social network propagation simulation" mechanism is employed. Newly emerging behavioral patterns are treated as information within a social network. By simulating the propagation path and speed of information within the agent network, the propagation of pattern-related information is artificially enhanced. Specifically, based on factors such as the historical interaction frequency and trust level between agents, information propagation weights are calculated, prioritizing the propagation of pattern information to agents with frequent interactions and high trust levels. In the value assessment phase, a "combat simulation stress test" stage is introduced. Selected potential strategy patterns are placed in simulated extreme combat scenarios for testing. By evaluating indicators such as the survivability of the strategies under stress scenarios and the degree of improvement in combat effectiveness, their potential value is more accurately determined.
[0180] By constructing a propagation-evaluation model, we demonstrate that social network propagation simulation optimizes the capture of emerging patterns.
[0181] Traditional capture accuracy is:
[0182]
[0183] The capture accuracy after introducing mode magnification is:
[0184] A ECEG =A traditional ·(1+μ·v);
[0185] The percentage increase in capture accuracy is:
[0186]
[0187] in:
[0188] A traditional The emergence capture accuracy of traditional methods is dimensionless.
[0189] T correct The number of valuable emerging patterns correctly captured;
[0190] T detected The number of all emergent patterns detected by the system;
[0191] A ECEG Emergent capture accuracy after ECEG intervention;
[0192] μ: Propagation weight coefficient based on agent interaction frequency, with a value of [0,1];
[0193] v: A connectivity metric for intelligent agent networks, reflecting the efficiency of information propagation;
[0194] ω eceg Emergent capture accuracy improvement ratio, dimensionless.
[0195] Summarize:
[0196] The three-layer nested dynamic architecture and bidirectional feedback mechanism solve the problem of insufficient simulation of complex multi-domain linkage in traditional inference systems. By constructing an "environmental perception-adversarial game-emergent mining" architecture and a millisecond-level information closed loop of intelligent information router, it realizes real-time interaction of multi-domain information and cross-level collaboration, significantly reduces the simulation error of multi-domain linkage, and significantly improves the consistency between strategy and actual combat scenario.
[0197] DAAN’s heterogeneous subnet adversarial-cooperative mode and RL-GAN hybrid mechanism solve the problems of rigid decision-making and insufficient flexibility of agents. By dynamically switching subnet modes through an environmental mutation detector, and combining inverse reinforcement learning and generative adversarial networks, it significantly improves the decision-making accuracy of agents in unknown and sudden scenarios, efficiently and autonomously generates cross-domain composite strategies, and breaks the limitations of preset rules on agents.
[0198] CDW's incremental topology analysis and knowledge graph-guided semantic weaving solve the problem of weak dynamic data processing capabilities. By reducing the time complexity of updating the data topology graph and introducing knowledge graph weight allocation, it significantly improves the scale of data processing and dynamic response speed of the system, deeply mines the potential relationships between data, and significantly increases the amount of effective information mined.
[0199] ECEG's adversarial sample perturbation generation and social network propagation simulation solves the problem of lack of system chaotic edge state management. By injecting adversarial samples to increase the probability of policy emergence and simulate information propagation paths, the system successfully captures a variety of valuable self-organizing emergence phenomena in the simulation, giving rise to the prototype of a new combat theory and filling the gap in traditional technology for monitoring and utilizing chaotic edge states.
[0200] Multi-module collaborative optimization (DAAN / CDW / ECEG) solves the problems of large deviations between simulation results and actual combat and limited training effects. By generating diverse combat strategies through DAAN, providing accurate intelligence through CDW, and simulating chaotic scenarios through ECEG, it significantly improves the decision-making ability and emergency response speed of combat personnel in complex scenarios, enabling the simulation system to realistically reproduce the complexity of the battlefield and promote the upgrading of military training models.
Claims
1. A multi-domain combat dynamic simulation system based on deep reinforcement learning, characterized in that, This includes a three-layer nested dynamic architecture encompassing environmental perception, adversarial game theory, and emergent discovery. This architecture deploys bidirectional feedback mechanisms, cross-activation mechanisms, and adaptive adjustment nodes. The three-layer nested dynamic architecture also includes a dynamic adaptive adversarial network, a chaotic data weaver, and an emergent capture and evolutionary guide. Among these components: The dynamic adaptive adversarial network consists of multiple heterogeneous adversarial subnets, each subnet corresponding to a combat domain. The subnets interact with each other through a dual-mode adversarial-cooperative interaction, and the subnets adopt a hybrid mechanism of inverse reinforcement learning and generative adversarial network. Chaotic data weavers treat multi-source heterogeneous data as a dynamic data cloud and use topological data analysis techniques to capture the topological structure and dynamic change trends of data in high-dimensional space. Emergent Capture and Evolution Guide is based on the principle of breaking the information cocoon. It monitors the interaction information and data flow of agents in the system in real time. When it detects that the system is close to the edge of chaos, it intervenes through a three-stage operation of information perturbation, pattern amplification and value assessment. The two-way feedback mechanism deploys intelligent information routers between levels. The intelligent information routers have built-in time-series correlation analysis modules and priority judgment algorithms to realize real-time information interaction and efficient processing. The cross-activation mechanism adopts an event-rule-resource driven model. When the emergence mining layer identifies a potential combat mode, it generates an activation command that includes event triggering conditions, rule constraints, and resource requirements, activating the simulation module of the adversarial game layer and the dynamic monitoring task scheduler of the environmental perception layer. The adaptive adjustment node integrates a multi-indicator fusion evaluation engine and a control strategy library. The multi-indicator fusion evaluation engine uses an algorithm that combines principal component analysis and fuzzy logic to reduce the monitoring indicators to three comprehensive dimensions: system load, decision complexity, and data correlation, which are used to judge the system status and control resources.
2. The multi-domain combat dynamic simulation system based on deep reinforcement learning as described in claim 1, characterized in that: An environmental mutation detector is embedded in the heterogeneous adversarial subnet. Based on the sliding window algorithm, the rate of change of environmental state characteristics is calculated in real time. When the rate of change exceeds a preset threshold, the subnet is triggered to switch from cooperative mode to adversarial mode.
3. The multi-domain combat dynamic simulation system based on deep reinforcement learning as described in claim 2, characterized in that: In the hybrid mechanism of inverse reinforcement learning and generative adversarial network, the generator introduces a policy fragment memory pool, and the discriminator adopts a multi-scale evaluation method to comprehensively score the policy from the time, space and resource scales.
4. The multi-domain combat dynamic simulation system based on deep reinforcement learning as described in claim 1, characterized in that: The chaotic data weaver uses an incremental topology analysis method to construct a data feature hash table. When new data flows in, it only calculates the connection relationship between the new data point and the existing topology to update the topology graph.
5. The multi-domain combat dynamic simulation system based on deep reinforcement learning as described in claim 4, characterized in that: The chaotic data weaver introduces a knowledge graph-guided weaving weight allocation mechanism, which assigns semantic weights to topological connections based on the strength of associations between data entities in the knowledge graph, thus forming a data weave.
6. The multi-domain combat dynamic simulation system based on deep reinforcement learning as described in claim 1, characterized in that: The information perturbation wave generation module of the Emergent Capture and Evolution Guide adopts an adversarial example-based perturbation generator, which generates adversarial examples against the agent's decision-making model as information perturbation waves through adversarial training; the pattern amplification and value assessment module adopts a social network propagation simulation mechanism, combined with combat simulation stress test to evaluate the strategy value.
Citation Information
Patent Citations
Unmanned cluster system evolution and feedback evolution method driven by bionic behavior normal form
CN117454926A
Intelligent inspection path planning method and system based on reinforcement learning
CN119990496A