Confrontation game cooperative control method and system for heterogeneous unmanned ship cluster
Through the hierarchical game strategy engine and hybrid equilibrium solution architecture, the problems of poor strategy coordination, low game solution efficiency and insufficient communication real-time in dynamic confrontation scenarios are solved, efficient and robust collaborative control is achieved, and task execution efficiency and real-time performance are improved.
Patent Information
- Application Number
- CN202510479119.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-08
AI Technical Summary
The existing unmanned boat clusters have problems such as insufficient strategic heterogeneity, low game equilibrium solution efficiency, lack of dynamic role adaptation mechanism, and limited communication and computing resources in terms of collaborative control, dynamic confrontation game and heterogeneous group task allocation, resulting in low task execution efficiency and limited coordination capabilities.
The hierarchical game strategy engine, role dynamic adaptation module, hybrid equilibrium solution architecture and communication optimization module are adopted, combined with dynamic training modules, differentiated strategy coordination of heterogeneous unmanned boats is realized, task execution efficiency is improved, game equilibrium solution time is shortened, dynamic role priority adjustment is supported, and distributed decision consistency is ensured under communication restriction conditions.
It improves task execution efficiency, shortens the game equilibrium solution time to 1.5 seconds, meets high real-time requirements, and takes only 0.3 seconds to ensure the consistency of distributed decisions with a communication delay of ≤200ms, providing an efficient and robust intelligent maritime operation collaborative control solution.
Smart Images

Figure CN120447543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-agent game control and distributed collaborative decision-making technology, and specifically to a confrontation game collaborative control method and system for heterogeneous unmanned boat clusters. Background Art
[0002] With the rapid development of unmanned boat swarm technology, its applications in military and civilian fields such as maritime attack and defense, reconnaissance and patrol, and electronic jamming are becoming increasingly widespread. However, existing technologies still face significant bottlenecks in multi-agent collaborative control, dynamic confrontation games, and heterogeneous group task allocation, specifically the following key issues:
[0003] 1. Insufficient strategy heterogeneity: Existing UAV swarm collaborative control systems are mostly based on homogeneous model designs and employ a unified strategy generation mechanism. This makes it difficult to effectively coordinate the differentiated behaviors of heterogeneous UAVs, such as reconnaissance, interception, and jamming vessels. For example, interception vessels require rapid response to path planning, while jamming vessels require dynamic frequency band strategy adjustments. Traditional homogeneous control models are unable to optimize for the characteristics of different types of UAVs, resulting in low mission execution efficiency and limited collaborative capabilities.
[0004] 2. Low efficiency in solving game equilibrium: In dynamic confrontation scenarios, traditional game theory methods, such as virtual games and repeated elimination of dominant strategies, need to search for equilibrium solutions in a high-dimensional strategy space, and the computational complexity increases exponentially. For example, in a collaborative mission involving 30 heterogeneous unmanned boats, the virtual game solver takes more than 12 seconds to solve the equilibrium, which cannot meet the requirements of high-real-time tasks such as port blockades, which require a response in seconds. In addition, existing methods are mostly limited to a single equilibrium type, such as only Nash equilibrium or Stackelberg equilibrium, and are difficult to adapt to mixed game modes in complex confrontation scenarios;
[0005] 3. Lack of dynamic role adaptation mechanism: Existing task allocation methods generally use static priorities or fixed role divisions, lacking the ability to dynamically adjust based on real-time situational awareness. For example, when enemy strategies suddenly change or the battlefield environment deteriorates, traditional methods cannot quickly reallocate the task priorities of reconnaissance and jamming vessels, resulting in low resource utilization.
[0006] 4. Communication and computing resource limitations: Distributed UAV swarms must achieve collaborative decision-making under conditions of limited communication bandwidth and latency sensitivity. Existing consensus protocols, such as Paxos and PBFT, suffer from high computational overhead and synchronization latency, typically exceeding 500ms, making them unable to meet the real-time requirements of maritime communication standards. Furthermore, traditional reinforcement learning methods are susceptible to exploration noise during heterogeneous policy optimization, resulting in slow policy convergence and poor stability.
[0007] In response to the above-mentioned technical defects, the present invention proposes a confrontation game collaborative control method and system for heterogeneous unmanned boat clusters. Through the comprehensive design of modules such as a layered game strategy engine, a dynamic role adaptation module, a hybrid equilibrium solution framework, and a communication optimization module, the following breakthroughs are achieved: solving the problem of poor coordination of heterogeneous strategies and improving task execution efficiency; reducing the game equilibrium solution time from 12 seconds to 1.5 seconds to meet high real-time requirements; supporting dynamic role priority adjustment, and task allocation takes only 0.3 seconds; ensuring distributed decision consistency under communication-restricted conditions (delay ≤ 200ms). Summary of the Invention
[0008] The purpose of the present invention is to provide a confrontation game collaborative control method and system for heterogeneous unmanned boat clusters, filling the technical gap in efficient collaborative control of heterogeneous unmanned boat clusters in dynamic confrontation scenarios, providing core support advantages for the actual deployment of intelligent maritime combat systems, and solving the core problems of poor strategy coordination, low game solving efficiency, rigid task allocation and insufficient real-time communication of heterogeneous unmanned boats in dynamic confrontation scenarios.
[0009] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a confrontation game collaborative control method and system for heterogeneous unmanned boat clusters, the system comprising a hierarchical game strategy engine, a role dynamic adaptation module, a distributed equilibrium solver, a communication optimization module, a confrontation strategy selection module and unmanned boats, wherein the hierarchical game strategy engine is used to coordinate the differentiated strategies of heterogeneous unmanned boats, and the hierarchical game strategy engine is composed of an upper-layer potential game task allocation module and a lower-layer reinforcement learning behavior optimization module, the upper-layer potential game task allocation module and the lower-layer reinforcement learning behavior optimization module are communicated with each other; the role dynamic adaptation module dynamically adjusts the unmanned boat task allocation module based on the real-time confrontation situation The system prioritizes tasks and realizes social maximization of task allocation based on the auction model of VCG mechanism, and the role dynamic adaptation module is connected to the hierarchical game strategy engine and the distributed equilibrium solver respectively; the distributed equilibrium solver integrates the hybrid solution architecture of Nash equilibrium and Stackelberg equilibrium to support fast game decision-making; the communication optimization module is connected to the distributed equilibrium solver, the adversarial strategy selection module and the unmanned boat respectively, and the communication optimization module realizes neighborhood strategy synchronization through the Raft consensus protocol to ensure that the maximum delay of distributed decision-making under communication constraints is ≤200ms; the adversarial strategy selection module is used to generate forged strategy signals with a preset probability.
[0010] Preferably, the system further comprises a dynamic training module, which is based on a training model of countermeasure evolution and is connected to the unmanned boat.
[0011] Preferably, the specific process of the dynamic training module is as follows:
[0012] ① Offline strategy pre-training unit: simulates 1,000+ adversarial scenarios in the digital twin platform;
[0013] ② Online strategy fine-tuning unit: collects enemy strategy data in real time and updates network weights through meta-learning. The specific formula is as follows:
[0014]
[0015] ③ Historical strategy library backtracking unit: When a sudden change in enemy strategy is detected, it matches similar patterns from the strategy library and loads the corresponding parameters;
[0016] Among them, θ new is the updated neural network weight parameter; θ old is the neural network weight parameter before updating; η is the learning rate; is the loss function Gradient with respect to parameter θ; D online An online dataset consisting of enemy strategy data collected in real time.
[0017] Preferably, in the upper potential game task allocation module of the hierarchical game strategy engine, for the i-th type of unmanned boat (type T i ), its task utility function is defined as:
[0018]
[0019] Among them, R task is the task benefit, C collab is the collaborative penalty term, and α, β, and γ are type weight coefficients (e.g., α = 0.7 for interceptor boats and β = 0.6 for reconnaissance boats).
[0020] The global situation function constructs the potential function through type-weighted aggregation, and its specific formula is as follows:
[0021]
[0022] Among them, w k For type T k The contribution weights reflect the contribution ratios of different types of groups in the overall situation function and are dynamically updated through historical task success rates. The update rules are as follows:
[0023] w (t+1) k =w (t) k +η·(success rate k -Average success rate)
[0024] Among them, η is the learning rate, success rate k For type T kThe average success rate over the past 10 missions.
[0025] Preferably, the lower-level reinforcement learning behavior optimization module of the hierarchical game strategy engine adopts a double-delayed deep deterministic policy gradient (TD3) algorithm, independently trains differentiated strategy networks according to the type of unmanned boat, and shares a global evaluation network.
[0026] Preferably, the differentiation strategy network includes:
[0027] Interceptor boat movement space: a intercept ∈{encirclement angle, maneuvering speed};
[0028] Jammer action space: a jamming ∈{frequency band selection, transmit power};
[0029] Scout boat movement space: a recon ∈{detection radius, cruise path};
[0030] The policy network parameters of each type of unmanned boat are trained independently, and a global evaluation network is shared for value function evaluation. The action exploration noise of the lower-level reinforcement learning behavior optimization module adopts type-related Ornstein-Uhlenbeck (OU) noise, and its specific formula is:
[0031]
[0032] Among them, σT i For type T i The noise standard deviation is set according to the type of unmanned boat (interceptor boat: σ = 0.3, jammer boat: σ = 0.5), τ is the noise time constant, and ∈ is the noise coefficient.
[0033] Preferably, in the VCG auction model of the role dynamic adaptation module, the valuation function of task j to boat i is defined as:
[0034] υ ij =λ·Task Completion Rate+μ·E 剩余能力
[0035] Among them, λ and μ are weight coefficients. The following formula is used to maximize the social surplus of task allocation:
[0036]
[0037] Among them, x ij =1 indicates that task j is assigned to boat i, otherwise it is 0. The VCG auction model is used to maximize the social surplus of task allocation, and the role priority dynamic adjustment method is adopted. The steps are as follows:
[0038] ① Calculation of capability-demand matching:
[0039] ② Priority reordering: press p every 5 seconds i Refresh the task queue in descending order;
[0040] Among them, the current task requirement vector is composed of normalized multidimensional features; the capability vector is also composed of normalized multidimensional features.
[0041] Preferably, the hybrid solver architecture of the distributed equilibrium solver includes: a leader decision layer and a follower decision layer. The leader decision layer is used to maximize the total utility of the group, and its core formula is:
[0042]
[0043] The follower decision layer is used to solve the Nash equilibrium, and its core formula is:
[0044]
[0045] Among them, a L ,a F are the actions of the leader decision layer and the follower decision layer, namely Decision vectors of the leader’s decision-making layer; The decision vector of the follower decision layer; U best i is the optimal utility threshold, that is, the preset individual utility target value, which is used for convergence judgment.
[0046] Preferably, the communication optimization module adopts a distributed decision-making method under communication constraints, and the specific steps are as follows:
[0047] ① Divide the local game into sub-problems: decompose the global game into 2-hop sub-games within the domain;
[0048] ②Consensus protocol synchronization: Based on the Raft consensus protocol, domain policy consistency is ensured, and the maximum delay of distributed decision-making under communication constraints is ≤200ms.
[0049] Preferably, the specific steps of generating a forged strategy signal by presetting a probability in the countermeasure strategy selection module include:
[0050] ① False strategy generation: Send a fake strategy signal to the enemy with a probability of 10%. The specific formula is as follows:
[0051]
[0052] ② Strategy confusion verification: If the enemy response deviates from the expected value by more than a threshold, the confusion is terminated and the main strategy is switched;
[0053] Among them, afake is the generated forged strategy action, a real It is the real strategic action signal of the unmanned boat. is zero-mean Gaussian noise with a variance of 0.2.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1. The present invention provides a confrontation game collaborative control method and system for heterogeneous unmanned boat clusters. Through the comprehensive design of a layered game strategy engine, a dynamic role adaptation module, a hybrid equilibrium solution frame, a communication optimization module and a confrontation strategy selection module, the problem of poor coordination of heterogeneous strategies is solved and the task execution efficiency is improved; the game equilibrium solution time is reduced from 12 seconds to 1.5 seconds to meet high real-time requirements; dynamic role priority adjustment is supported, and task allocation takes only 0.3 seconds; under communication-restricted conditions, the delay is ≤200ms, ensuring the consistency of distributed decision-making. At the same time, the present invention provides an efficient and robust collaborative control solution for intelligent maritime confrontation tasks, with significant military and civilian value. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a principle block diagram of the present invention;
[0057] Figure 2 This is a block diagram of the core problem of the present invention and the technical effects after each module solves it;
[0058] Figure 3 This is a flow chart of the hybrid balance solution of the present invention.
[0059] The reference numerals and names in the figures are as follows:
[0060] 1. Hierarchical game strategy engine; 11. Upper-level potential game task allocation module; 12. Lower-level reinforcement learning behavior optimization module; 2. Role dynamic adaptation module; 3. Distributed equilibrium solver; 4. Communication optimization module; 5. Countermeasure strategy selection module; 6. Unmanned boat; 7. Dynamic training module. DETAILED DESCRIPTION
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0062] In the description of the embodiments of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0063] In the embodiments of the present invention, unless otherwise expressly specified or limited, the terms "installed," "connected," "connected," "fixed," etc. should be understood in a broad sense. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the embodiments of the present invention based on specific circumstances.
[0064] See also Figures 1 to 3, an embodiment provided by the present invention: a confrontation game collaborative control method and system for heterogeneous unmanned boat clusters, the system includes a hierarchical game strategy engine 1, a role dynamic adaptation module 2, a distributed equilibrium solver 3, a communication optimization module 4, a confrontation strategy selection module 5 and an unmanned boat 6, wherein the hierarchical game strategy engine 1 is used to coordinate the differentiated strategies of heterogeneous unmanned boats, and the hierarchical game strategy engine 1 is composed of an upper-layer potential game task allocation module 11 and a lower-layer reinforcement learning behavior optimization module 12, and the upper-layer potential game task allocation module 11 and the lower-layer reinforcement learning behavior optimization module 12 are communicated with each other; the role dynamic adaptation module 2 dynamically adjusts the unmanned boat task priority based on the real-time confrontation situation, and realizes the social maximization of task allocation based on the auction model of the VCG mechanism, and the role dynamic adaptation module 2 are respectively connected to the layered game strategy engine 1 and the distributed equilibrium solver 3; the distributed equilibrium solver 3 integrates a hybrid solution architecture of Nash equilibrium and Stackelberg equilibrium to support fast game decision-making; the communication optimization module 4 is respectively connected to the distributed equilibrium solver 3, the confrontation strategy selection module 5 and the unmanned boat 6, and the communication optimization module 4 realizes neighborhood strategy synchronization through the Raft consensus protocol to ensure that the maximum delay of distributed decision-making under communication constraints is ≤200ms; the confrontation strategy selection module 5 is used to generate a forged strategy signal with a preset probability. The system also includes a dynamic training module 7, which is based on a training model of confrontation evolution, and the dynamic training module 7 is connected to the unmanned boat 6. There can be multiple unmanned boats 6, and the multiple unmanned boats 6 are communicated with each other.
[0065] Specifically, in the hierarchical game strategy engine 1, the upper-level potential game task allocation module 11 realizes differentiated task allocation of heterogeneous unmanned boats based on the type-weighted global situation function, while the lower-level reinforcement learning behavior optimization module 12 optimizes the individual behavior strategies of each type of unmanned boat 6 through the TD3 algorithm and the differentiated strategy network. The upper-level potential game task allocation module 11 passes the task allocation results, such as reconnaissance area, interception target and other information to the lower-level reinforcement learning behavior optimization module 12. The lower-level reinforcement learning behavior optimization module 12 generates specific action instructions in real time according to the task objectives, such as speed, frequency band selection, etc., forming a closed-loop optimization of "global task → local behavior".
[0066] In the upper potential game task allocation module 11 of the hierarchical game strategy engine 1, for the i-th type of unmanned boat (type T i ), its task utility function is defined as:
[0067]
[0068] Among them, R task is the task benefit, C collabis the collaborative penalty term, and α, β, and γ are type weight coefficients (e.g., α = 0.7 for interceptor boats and β = 0.6 for reconnaissance boats).
[0069] The global situation function constructs the potential function through type-weighted aggregation, and its specific formula is as follows:
[0070]
[0071] Among them, w k For type T k The contribution weights reflect the contribution ratios of different types of groups in the overall situation function and are dynamically updated through historical task success rates. The update rules are as follows:
[0072] w (t+1) k =w (t) k +η·(success rate k -Average success rate)
[0073] Among them, η is the learning rate, success rate k For type T k The average success rate over the past 10 missions.
[0074] Specifically, the lower-level reinforcement learning behavior optimization module 12 of the hierarchical game strategy engine 1 adopts a double-delayed deep deterministic policy gradient (TD3) algorithm to independently train a differentiated strategy network according to the type of unmanned boat and share a global evaluation network. The differentiated strategy network includes:
[0075] Interceptor boat movement space: a intercept ∈{encirclement angle, maneuvering speed};
[0076] Jammer action space: a jamming ε{frequency band selection, transmit power};
[0077] Scout boat movement space: a recon ∈{detection radius, cruise path};
[0078] The policy network parameters of each type of unmanned boat are trained independently, and a global evaluation network is shared for value function evaluation. The action exploration noise of the lower-level reinforcement learning behavior optimization module adopts type-related Ornstein-Uhlenbeck (OU) noise, and its specific formula is:
[0079]
[0080] Among them, σT i For type T iThe noise standard deviation is set according to the type of unmanned boat (interceptor boat: σ = 0.3, jammer boat: σ = 0.5), τ is the noise time constant, and ∈ is the noise coefficient.
[0081] Specifically, the role dynamic adaptation module 2 adopts the VCG auction model to ensure efficient resource utilization and achieve dynamic priority adjustment by maximizing the social surplus task allocation. The collaboration mechanism is to receive the task allocation request of the hierarchical engine, combine real-time battlefield data such as enemy position and remaining energy, dynamically adjust the role priority of the unmanned boat, and feed the optimized task list back to the equilibrium solver. In the VCG auction model of the role dynamic adaptation module 2, the valuation function of task j for boat i is defined as:
[0082] υ ij =λ·Task Completion Rate+μ·E 剩余能力
[0083] Among them, λ and μ are weight coefficients. The following formula is used to maximize the social surplus of task allocation:
[0084]
[0085] Among them, x ij =1 indicates that task j is assigned to boat i, otherwise it is 0. The VCG auction model is used to maximize the social surplus of task allocation, and the role priority dynamic adjustment method is adopted. The steps are as follows:
[0086] ① Calculation of capability-demand matching:
[0087] ② Priority reordering: press p every 5 seconds i Refresh the task queue in descending order;
[0088] Among them, the current task requirement vector is composed of normalized multidimensional features; the capability vector is also composed of normalized multidimensional features.
[0089] Specifically, the distributed equilibrium solver 3 performs a hybrid equilibrium solution, that is, combining Stackelberg and Nash equilibrium, hierarchically optimizing the game strategies of the leader, command boat, follower, and execution boat, and compressing the equilibrium calculation time through an iterative solution algorithm. At the same time, based on the task list output by the role dynamic adaptation module, the command boat generates a global path planning leader layer, and the execution boat synchronously optimizes the speed and interference strategy follower layer, and realizes strategy synchronization through the communication optimization module 4. The hybrid solver architecture of the distributed equilibrium solver 3 includes: a leader decision layer and a follower decision layer. The leader decision layer is used to maximize the total utility of the group, and its core formula is:
[0090]
[0091] The follower decision layer is used to solve the Nash equilibrium, and its core formula is:
[0092]
[0093] Among them, a L ,a F are the actions of the leader decision layer and the follower decision layer, namely Decision vectors of the leader’s decision-making layer; The decision vector of the follower decision layer; U best i is the optimal utility threshold, that is, the preset individual utility target value, which is used for convergence judgment.
[0094] Specifically, the communication optimization module 4 decomposes the global game into local sub-games within the 2-hop communication range to reduce computational complexity, and then uses the Raft consensus protocol to ensure distributed decision consistency. The maximum communication delay is ≤ 200ms. The communication optimization module 4 adopts a distributed decision-making method under communication constraints. The specific steps are as follows:
[0095] ① Divide the local game into sub-problems: decompose the global game into 2-hop sub-games within the domain;
[0096] ②Consensus protocol synchronization: Based on the Raft consensus protocol, domain policy consistency is ensured, and the maximum delay of distributed decision-making under communication constraints is ≤200ms.
[0097] Specifically, the countermeasure strategy selection module 5 first generates a fake signal with a probability of 10% to confuse the enemy's strategy and interfere with the enemy's decision-making; then, after detecting an abnormal response from the enemy, it immediately switches to the historical optimal strategy in the backup strategy library. The specific steps of generating the fake strategy signal with a preset probability by the countermeasure strategy selection module 5 include:
[0098] ① False strategy generation: Send a fake strategy signal to the enemy with a probability of 10%. The specific formula is as follows:
[0099]
[0100] ② Strategy confusion verification: If the enemy response deviates from the expected value by more than a threshold, the confusion is terminated and the main strategy is switched;
[0101] Among them, a fake is the generated forged strategy action, a real It is the real strategic action signal of the unmanned boat. is zero-mean Gaussian noise with a variance of 0.2.
[0102] Specifically, the specific process of the dynamic training module 7 is as follows:
[0103] ① Offline strategy pre-training unit: simulates 1,000+ adversarial scenarios in the digital twin platform;
[0104] ② Online strategy fine-tuning unit: collects enemy strategy data in real time and updates network weights through meta-learning. The specific formula is as follows:
[0105]
[0106] ③ Historical strategy library backtracking unit: When a sudden change in enemy strategy is detected, it matches similar patterns from the strategy library and loads the corresponding parameters;
[0107] Among them, θ new is the updated neural network weight parameter; θ old is the neural network weight parameter before updating; η is the learning rate; is the loss function Gradient with respect to parameter θ; D online An online dataset consisting of enemy strategy data collected in real time.
[0108] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A method and system for cooperative control of a heterogeneous unmanned watercraft swarm using adversarial game theory, characterized by: The system comprises a hierarchical game strategy engine (1), a role dynamic adaptation module (2), a distributed equilibrium solver (3), a communication optimization module (4), a confrontation strategy selection module (5) and an unmanned boat (6), wherein the hierarchical game strategy engine (1) is used to coordinate the differentiated strategies of heterogeneous unmanned boats, and the hierarchical game strategy engine (1) is composed of an upper-layer potential game task allocation module (11) and a lower-layer reinforcement learning behavior optimization module (12), and the upper-layer potential game task allocation module (11) and the lower-layer reinforcement learning behavior optimization module (12) are connected in communication; the role dynamic adaptation module (2) dynamically adjusts the unmanned boat task priority based on the real-time confrontation situation, and is based on the auction model of the VCG mechanism. The social maximization of task allocation is achieved, and the role dynamic adaptation module (2) is connected to the hierarchical game strategy engine (1) and the distributed equilibrium solver (3) respectively; the distributed equilibrium solver (3) integrates a hybrid solution architecture of Nash equilibrium and Stackelberg equilibrium to support fast game decision-making; the communication optimization module (4) is connected to the distributed equilibrium solver (3), the confrontation strategy selection module (5) and the unmanned boat (6) respectively, and the communication optimization module (4) realizes neighborhood strategy synchronization through the Raft consensus protocol to ensure that the maximum delay of distributed decision-making under communication constraints is ≤200ms; the confrontation strategy selection module (5) is used to generate a forged strategy signal with a preset probability.
2. The method and system for cooperative control of a heterogeneous unmanned vehicle swarm based on adversarial game theory according to claim 1, characterized in that: It also includes a dynamic training module (7), which is based on a training model of confrontation evolution, and the dynamic training module (7) is connected to the unmanned boat (6).
3. The method and system for the confrontation game collaborative control of a heterogeneous unmanned vehicle swarm according to claim 2, characterized in that: The specific process of the dynamic training module (7) is as follows: ① Offline strategy pre-training unit: simulates 1,000+ adversarial scenarios in the digital twin platform; ② Online strategy fine-tuning unit: collects enemy strategy data in real time and updates network weights through meta-learning. The specific formula is as follows: ③ Historical strategy library backtracking unit: When a sudden change in enemy strategy is detected, it matches similar patterns from the strategy library and loads the corresponding parameters; Among them, θ new is the updated neural network weight parameter; θ old is the neural network weight parameter before updating; η is the learning rate; is the loss function Gradient with respect to parameter θ; D online An online dataset consisting of enemy strategy data collected in real time.
4. The method and system for cooperative control of a heterogeneous unmanned watercraft swarm based on adversarial game theory according to claim 1, characterized in that: In the upper potential game task allocation module (11) of the layered game strategy engine (1), for the i-th type of unmanned boat (type T i ), its task utility function is defined as: Among them, R task is the task benefit, C collab is the collaborative penalty term, and ɑ, β, and γ are type weight coefficients (e.g., α = 0.7 for interceptor boats and β = 0.6 for reconnaissance boats). The global situation function constructs the potential function through type-weighted aggregation, and its specific formula is as follows: Among them, w k For type T k The contribution weights reflect the contribution ratios of different types of groups in the overall situation function and are dynamically updated through historical task success rates. The update rules are as follows: w (t+1) k =w (t) k +η·(success rate k -Average success rate) Among them, η is the learning rate, success rate k For type T k The average success rate over the past 10 missions.
5. The method and system for cooperative control of a heterogeneous unmanned watercraft swarm based on adversarial game theory according to claim 1, characterized in that: The lower layer reinforcement learning behavior optimization module (12) of the hierarchical game strategy engine (1) adopts a double-delayed deep deterministic policy gradient (TD3) algorithm to independently train differentiated strategy networks according to the type of unmanned boat and share a global evaluation network.
6. The method and system for the confrontation game collaborative control of a heterogeneous unmanned watercraft swarm according to claim 5, characterized in that: The differentiation strategy network includes: Interceptor boat movement space: a intercept ∈{encirclement angle, maneuvering speed}; Jammer action space: a jamming ∈{frequency band selection, transmit power}; Scout boat movement space: a recon ∈{detection radius, cruise path}; The policy network parameters of each type of unmanned boat are trained independently, and a global evaluation network is shared for value function evaluation. The action exploration noise of the lower-level reinforcement learning behavior optimization module adopts type-related Ornstein-Uhlenbeck (OU) noise, and its specific formula is: Among them, σT i For type T i The noise standard deviation is set according to the type of unmanned boat (interceptor boat: σ = 0.3, jammer boat: σ = 0.5), τ is the noise time constant, and ∈ is the noise coefficient.
7. The method and system for cooperative control of a heterogeneous unmanned watercraft swarm based on adversarial game theory according to claim 1, characterized in that: In the VCG auction model of the role dynamic adaptation module (2), the valuation function of task j to boat i is defined as: υ ij =λ·Task Completion Rate+μ·E 剩余能力 Among them, λ and μ are weight coefficients. The following formula is used to maximize the social surplus of task allocation: Among them, x ij =1 indicates that task j is assigned to boat i, otherwise it is 0. The VCG auction model is used to maximize the social surplus of task allocation, and the role priority dynamic adjustment method is adopted. The steps are as follows: ① Calculation of capability-demand matching: ② Priority reordering: press p every 5 seconds i Refresh the task queue in descending order; Among them, the current task requirement vector is composed of normalized multidimensional features; the capability vector is also composed of normalized multidimensional features.
8. The method and system for the confrontation game collaborative control of a heterogeneous unmanned watercraft swarm according to claim 1, characterized in that: The hybrid solver architecture of the distributed equilibrium solver (3) includes: a leader decision layer and a follower decision layer. The leader decision layer is used to maximize the total utility of the group. Its core formula is: The follower decision layer is used to solve the Nash equilibrium, and its core formula is: Among them, a L ,a F are the actions of the leader decision layer and the follower decision layer, namely Decision vectors of the leader’s decision-making layer; The decision vector of the follower decision layer; U best i is the optimal utility threshold, that is, the preset individual utility target value, which is used for convergence judgment.
9. The method and system for cooperative control of a heterogeneous unmanned watercraft swarm based on adversarial game theory according to claim 1, characterized in that: The communication optimization module (4) adopts a distributed decision-making method under communication constraints, and its specific steps are as follows: ① Divide the local game into sub-problems: decompose the global game into 2-hop sub-games within the domain; ②Consensus protocol synchronization: Based on the Raft consensus protocol, domain policy consistency is ensured, and the maximum delay of distributed decision-making under communication constraints is ≤200ms.
10. The method and system for cooperative control of a heterogeneous unmanned watercraft swarm based on adversarial game theory according to claim 1, characterized in that: The specific steps of generating a forged strategy signal with a preset probability by the countermeasure strategy selection module (5) include: ① False strategy generation: Send a fake strategy signal to the enemy with a 10% probability. The specific formula is as follows: ② Strategy confusion verification: If the enemy response deviates from the expected value by more than a threshold, the confusion is terminated and the main strategy is switched; Among them, a fake is the generated forged strategy action, a real It is the real strategic action signal of the unmanned boat. is zero-mean Gaussian noise with a variance of 0.2.
Citation Information
Cited By
Counter control system and method for dynamic path correction of unmanned ship
CN121187311A
Dynamic cooperative control method and system for intelligent power plant with few people on duty
CN121763716A
Asymmetric information multi-boat cooperative defense simulation method based on Bayesian Stackelberg game
CN122021368A
Simulation method for asymmetric information multi-ship cooperative defense based on bayesian stackelberg game
CN122021368B