Cooperative reconnaissance and electronic countermeasure control system and method based on multiple unmanned aerial vehicles

By adopting multi-UAV collaborative control system in the drone electronic countermeasure system and using technologies such as reinforcement learning and self-organized networks, the problem that drone clusters are difficult to adjust their strategies in complex battlefield environments is solved, and efficient task allocation and coordinated operations are achieved.

CN120010549AInactive Publication Date: 2025-05-16SUZHOU LUYAO XINGCHEN TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510126633.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing drone electronic countermeasure system is difficult to quickly adjust its combat strategies in complex and dynamic battlefield environments, resulting in the inability of drone clusters to fully utilize the advantages of coordinated operations. At the same time, the diversity of mission goals and conflicts between missions complicate resource scheduling and task allocation.

Method used

The collaborative reconnaissance and electronic confrontation control system based on multi-unmanned aerial vehicles is adopted, including communication modules, task decomposition modules, reinforcement learning decision modules, collaborative optimization modules and task evaluation modules. Through technologies such as self-organized networks, deep reinforcement learning and multi-agent collaboration, real-time task adjustment and resource optimization are achieved.

Benefits of technology

It effectively resolves conflicts and resource allocation problems between tasks, adjusts strategies in real time to respond to changes in the battlefield environment, and improves the efficiency and collaborative combat capabilities of drone electronic countermeasures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010549A_ABST
    Figure CN120010549A_ABST
Patent Text Reader

Abstract

The invention provides a cooperative reconnaissance and electronic countermeasure control system and method based on multiple unmanned aerial vehicles. The cooperative reconnaissance and electronic countermeasure control system comprises five modules: a communication module, a task decomposition module, a reinforcement learning decision module, a cooperative optimization module and a task evaluation module. The communication module ensures data synchronization and safe communication between the unmanned aerial vehicles through a self-organizing network and an encryption anti-interference technology. And the task decomposition module is used for refining the electronic confrontation task into reconnaissance and interference subtasks according to the task type, and adjusting the priority in real time according to the battlefield environment. The reinforcement learning decision module adopts a DQN and PPO deep reinforcement learning algorithm to optimize a reconnaissance and interference strategy, and ensures optimal task execution. And the collaborative optimization module coordinates task allocation and flight paths of the unmanned aerial vehicle through multi-agent reinforcement learning and a bee colony algorithm, and solves the problem of task conflict. And the task evaluation module monitors a task state in real time and feeds back an adjustment strategy. According to the method, battlefield changes can be flexibly coped with, resource scheduling is optimized, task execution efficiency is improved, anti-interference capability is enhanced, and stability and safety of tasks are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of military electronic countermeasures and artificial intelligence technology, and more specifically, to a collaborative reconnaissance and electronic countermeasures control system and method based on multiple unmanned aerial vehicles. Background Art

[0002] With the rapid advancement of drone technology, especially breakthroughs in communication, navigation, and intelligent control, drone swarms have become an indispensable force in modern warfare. Drone swarms can efficiently perform tasks such as reconnaissance and electronic jamming through coordination and cooperation, and play an increasingly important role, especially in complex battlefield environments. For example, multiple drones can cooperate in combat to quickly carry out precise strikes on enemy targets through alternating reconnaissance, dynamic coverage, and rapid deployment, or suppress the enemy's information system through electronic countermeasures.

[0003] However, despite the broad application prospects of drone swarms, existing drone electronic countermeasure systems still face many challenges in practical applications. First, the complexity and dynamic changes of the battlefield environment make drone swarms prone to insufficient decision-making capabilities when performing tasks. The constant changes in the electromagnetic environment, the diversity of enemy interference methods, and the constant changes in mission objectives require the system to be able to adjust combat strategies in real time. However, existing control systems are often unable to respond to these changes quickly, resulting in drone swarms being unable to fully utilize the advantages of collaborative operations.

[0004] Secondly, the diversity of mission objectives and the conflicts between tasks are another important factor that restricts the efficient execution of tasks by drone swarms. For example, when performing reconnaissance sub-tasks, drones may need to stay in the enemy radar blind spot for a long time, while when performing electronic jamming sub-tasks, drones are required to intensively suppress the enemy's information system. This conflict between tasks makes the resource scheduling and task allocation of drone swarms more complicated. Existing control methods often cannot effectively balance the task load of each drone, resulting in waste of resources or inefficient task execution.

[0005] Furthermore, traditional UAV swarm control methods often rely on static preset task allocation and path planning algorithms, which cannot flexibly deal with complex decision-making problems in high-dimensional action spaces. When faced with rapidly changing battlefield environments and multi-dimensional task requirements, the optimization capabilities of these methods are stretched to the limit, making it difficult to achieve efficient collaborative operations. Summary of the invention

[0006] In view of the deficiencies in the prior art, the object of the present invention is to provide a collaborative reconnaissance and electronic countermeasure control system and method based on multiple UAVs.

[0007] To achieve the above-mentioned object, the present invention provides the following technical solution: a collaborative reconnaissance and electronic countermeasure control system based on multiple UAVs, characterized in that the system includes a communication module, a task decomposition module, a reinforcement learning decision module, a collaborative optimization module and a task evaluation module;

[0008] The communication module includes a self-organizing network unit, a data synchronization unit and an encryption and anti-interference unit;

[0009] The task decomposition module includes a task parsing unit, a task priority allocator and a resource allocation unit;

[0010] The reinforcement learning decision module includes a reconnaissance strategy unit, an interference strategy unit and a strategy fusion and regulation unit;

[0011] The collaborative optimization module includes a task collaboration management unit, a multi-machine action coordination unit, and a conflict detection and resolution unit;

[0012] The task evaluation module includes a task state monitoring unit, an environment state evaluation unit and a feedback generation unit.

[0013] Preferably, the task decomposition module transmits the task instructions through the communication module and cooperates with the reinforcement learning decision module to obtain the task decision result;

[0014] The reinforcement learning decision module makes decisions based on the task requirements and the status of the drone, and cooperates with the collaborative optimization module to adjust the drone task allocation and flight path. The communication module transmits the task information, status information and decision instructions to each module, and feeds back the status of each drone to the task evaluation module;

[0015] The task evaluation module provides feedback on task execution results and optimizes subsequent task allocation and strategy decisions.

[0016] A control method based on multi-UAV cooperative reconnaissance and electronic countermeasures, the method comprising the following steps:

[0017] Step S1: The self-organizing network unit communicates point-to-point through the self-organizing network protocol and dynamically adjusts the network topology according to the physical location of the UAV; the data synchronization unit ensures the consistency of battlefield data shared between UAVs through the time synchronization protocol; the encryption and anti-interference unit enhances the link anti-interference capability through frequency hopping communication;

[0018] Step S2: The task analysis unit receives the electronic countermeasure task from the external command system, sends it to the UAV cluster, and breaks it down into a reconnaissance subtask and a jamming subtask according to the task type;

[0019] Step S3: the task priority allocator adjusts the subtask priority according to the battlefield environment;

[0020] Step S4: the resource allocation unit allocates resources to the UAV to adapt to different mission objectives;

[0021] Step S5: The reconnaissance strategy unit combines DQN and PPO deep reinforcement learning to train the reconnaissance strategy, inputs state space and action space information during training, and outputs the best path or signal detection frequency;

[0022] Step S6: The interference strategy unit optimizes interference frequency band selection and interference power to perform interference operations based on deep reinforcement learning input of enemy frequency band usage according to enemy communication or radar signal characteristics;

[0023] Step S7: The task collaboration management unit is based on multi-agent reinforcement learning. Each UAV independently evaluates its own task adaptability and negotiates with other UAVs to complete task allocation. The multi-machine action coordination unit optimizes UAV actions based on the swarm algorithm, dynamically adjusts the flight path, and covers the target area.

[0024] Step S7: the conflict detection and resolution unit detects possible task conflicts in the system, analyzes the costs and benefits of the conflicts based on a game theory model, and adjusts task allocation or interference parameters in real time;

[0025] Step S8: The mission status monitoring unit dynamically monitors the reconnaissance coverage and interference effect index of the UAV to evaluate whether the mission has achieved the target requirements; the environment status evaluation unit collects the strength and frequency characteristics of the signal transmitted by the enemy equipment through the sensor to analyze whether the enemy equipment has been effectively interfered;

[0026] Step S9: Based on the evaluation results, return to step S3, where the task priority allocator adjusts the subtask priorities according to the battlefield environment and then repeats the subsequent steps.

[0027] Preferably, when the reconnaissance strategy of the reconnaissance subtask is trained, the input state space includes the location of the enemy signal and the current location information of the UAV, and the output action space includes flight path adjustment and reconnaissance signal frequency allocation.

[0028] Preferably, in the decision-making strategy in the interference subtask, the input state includes the signal frequency band and strength of the enemy radar, and the output action space includes the interference frequency, power selection and the selection of the target interference strength.

[0029] Preferably, the communication module performs communication initialization when each drone is started, discovers other drone nodes around it through broadcasting, and establishes a neighbor node list;

[0030] Based on the neighbor node information, the self-organizing network component dynamically generates the topology of the cluster communication network, and the identity authentication module of the encryption and security component ensures that only legitimate nodes can join the network.

[0031] Preferably, the reconnaissance subtask decision process is controlled by the state space and action space data of the drone and a set reward function.

[0032] Preferably, the position of the UAV in the state space is expressed using relative coordinates, where the origin is defined as the position where the UAV first detects radar information, and the direction of detection is expressed in absolute angles, starting from 0° in the north direction and calculated clockwise;

[0033] The action space includes approaching a target, hovering, and launching a reconnaissance payload;

[0034] The reward function includes approach reward and reconnaissance reward. When a multi-UAV team collaborates, each time a radar is captured and located, the team can obtain a reconnaissance reward, and each radar accumulates 1 point.

[0035] Preferably, the reconnaissance subtask and the interference subtask respectively decide on the combination of flight action and reconnaissance action, and the combination of flight action and interference action, and decide on the weight of repeated flight action according to the progress of the task, and reach a complete comprehensive decision with independent reconnaissance action and interference action.

[0036] Preferably, the sample data of each event in the deep learning training of the reconnaissance strategy unit is stored in the experience pool. After the event is over, the samples are processed by the Monte Carlo method, and the global state information is processed based on the Critic network to capture the collaborative relationship between the drones:

[0037] M1: Initialize the experience pool buffer, the capacity of which is the length of a single event episode;

[0038] M2: Randomly initialize Actor network parameters θ π and Critic network parameters θ π , is a random number in the range [-0.003, 0.003];

[0039] M3: Initialize the reconnaissance sub-strategy π scout and the interference substrategy π jam ;

[0040] M4: When the episode is less than the total number of episodes, execute steps M4 to M14;

[0041] M5: Initialization state s1;

[0042] M6: When the episode length (time step t = 1, 2, ..., K) is less than K, execute steps M6 to M13;

[0043] M7: For UAV i=1,…,m, execute steps M7 to M12;

[0044] M8: Based on local observations from drone i Select comprehensive action

[0045] M9: Call the reconnaissance substrategy π scout , according to observation Select the reconnaissance action a i-scout ;

[0046] M10: Call the interference sub-strategy π jam , according to observation Select the interference action a i-jam ;

[0047] M11: put a i-scout and a i-jam According to the synthesis of complete

[0048] M12: Performing a complete joint movement Get rewards t and the next time step state s t+1 ;

[0049] M13: The experience sample (s t ,a t ,r t ,s t+1 ) is stored in the buffer, time step t = t + 1;

[0050] M14: Calculate the discounted return G for each time step t , and the number of episodes is increased by 1;

[0051] M15: Randomly sample n groups of sample experiences from the buffer;

[0052] M16: Calculate the Actor loss function L(θ) and update the Actor network;

[0053] M17: Calculate the Critic loss function And update the Critic network;

[0054] Among them, Actor loss function

[0055] The reward function is

[0056]

[0057] The advantage function is

[0058]

[0059] The critic loss function is

[0060] The ideal discounted return is

[0061] G t =r t+1 +γr t+2 +…+γ T-t r T+1

[0062] Error δ t =r t +γV(s t +1)-V(s t ), E is the expectation, γ is the discount factor, V(s t ) is the state s t The value function of θ (a t |s t ) represents the strategy of transferring from state s to action s at time t, λ is the weight, is the predicted discounted return.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] 1. In the present invention, the multi-machine collaborative game theory model can effectively solve the conflicts between tasks and resource allocation problems, avoiding the low combat efficiency caused by task conflicts or insufficient resources in traditional methods.

[0065] 2. In the present invention, by optimizing the decision-making of multiple action dimensions such as flight path, interference frequency band selection, and interference intensity, the strategy can be adjusted in real time to cope with changes in the battlefield environment, alleviating the decision-making difficulties in high-dimensional action space in UAV electronic countermeasure tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The drawings described herein are used to provide further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0067] Figure 1 The decision-making process of multi-UAV deep reinforcement learning under subtask decomposition. DETAILED DESCRIPTION

[0068] The present invention provides a multi-UAV-based collaborative reconnaissance and electronic countermeasure control system and method, which includes five core modules: communication module, task decomposition module, reinforcement learning decision module, collaborative optimization module and task evaluation module. The functions of each module are relatively independent but mutually coordinated, forming a complete UAV cluster electronic countermeasure system to cope with complex battlefield environments.

[0069] The self-organizing network unit in the communication module realizes point-to-point communication through the self-organizing network protocol and dynamically adjusts the network topology according to the physical location of the drone.

[0070] The data synchronization unit ensures the consistency of battlefield data shared between UAVs through a time synchronization protocol.

[0071] The encryption and anti-interference unit uses encryption algorithm (AES) to protect communication data and prevent enemy eavesdropping; it enhances the link anti-interference capability through frequency hopping communication.

[0072] The task analysis unit in the task decomposition module receives data from the external command system and sends electronic countermeasure tasks to the drone cluster, such as "reconnaissance of enemy radar locations" and "interference of communication links". The task analysis unit breaks down the tasks into subtasks according to the task type, such as "signal detection in a certain area" and "interference with specific frequency bands".

[0073] The task priority allocator adjusts the priority in real time according to the battlefield environment (such as the strength of the enemy radar signal or the importance of the communication link), and the priority is adjusted dynamically. For example, when the enemy communication signal strength is high, the interference subtask priority will be higher than the reconnaissance subtask. The resource allocation unit allocates spectrum, energy and computing power resources to the drone to ensure resource optimization for different mission objectives.

[0074]

[0075] Furthermore, the reconnaissance strategy unit in the reinforcement learning decision module: combines DQN and PPO deep reinforcement learning to train the reconnaissance strategy, maximizes the coverage of the reconnaissance area and improves the efficiency of electromagnetic signal capture. During training, the enemy signal position, drone position state space data and flight path adjustment, interference intensity allocation action space data are input, and the optimal path or signal detection frequency recommended action is output.

[0076] The jamming strategy unit uses a deep reinforcement learning algorithm to input the enemy's frequency band, jamming frequency, and power selection usage, learn the optimal jamming operation, and maximize the interference with enemy communications or radar equipment.

[0077] The strategy fusion and control unit uses the weight adjustment mechanism to dynamically adjust the importance of the two strategies according to task requirements. For example, when the enemy's communication signal has been effectively interfered, the system will reduce the interference weight and increase the weight of the reconnaissance subtask. It will analyze the conflict points of the two strategies through game theory and generate a comprehensive optimal strategy.

[0078] The task collaboration management unit in the collaborative optimization module is based on a multi-agent reinforcement learning method. Each drone independently evaluates its own task adaptability and negotiates with other drones to complete task allocation. For example, if a drone detects that the signal interference area has been covered, it can actively adjust its target area to avoid repeated interference.

[0079] The multi-machine action coordination unit optimizes drone actions based on the swarm algorithm, ensures that drones maintain a reasonable physical distance, avoid mutual occlusion or collision, and dynamically adjusts flight paths so that drones can efficiently cover the target area.

[0080] The conflict detection and resolution unit detects possible task conflicts in the system, such as multiple drones trying to interfere with the same frequency. Through the game theory model, it analyzes the costs and benefits of the conflict and adjusts the task allocation or interference parameters in real time to ensure the overall efficiency of the system is optimized.

[0081] The mission status monitoring unit in the mission evaluation module dynamically monitors the drone's reconnaissance coverage, interference effect and other indicators to evaluate whether the mission meets the target requirements.

[0082] The environmental status assessment unit collects the strength, frequency and other characteristics of the signals transmitted by the enemy equipment through sensors to analyze whether the enemy equipment has been effectively interfered with.

[0083] The feedback generation unit quantifies the task completion status into a reinforcement learning reward value, such as a positive reward for successfully jamming the enemy's communication equipment, and a negative reward for mission failure due to resource conflicts among multiple drones. The feedback results are used to optimize the strategy training of the reinforcement learning model.

[0084] When multiple drones are in coordinated confrontation, communication initialization is first performed. When each drone is started, it will perform communication initialization, discover other drone nodes around it through broadcasting, and establish a list of neighbor nodes; based on the neighbor node information, the self-organizing network component dynamically generates the topology of the cluster communication network, and through the identity authentication module of the encryption and security component, it is ensured that only legitimate nodes can join the network, and then the task information is received and transmitted. The command center sends the electronic countermeasure task instruction to the drone cluster through the communication module, and the drone that receives the task (usually the node where the task decomposition module is located) broadcasts the task information to other drones through the self-organizing network; the task decomposition module decomposes the task received by the drone into reconnaissance or interference subtasks, and the task instruction is transmitted to the specific execution drone through the radio communication module; after receiving the task, each drone sends a task confirmation signal to the task decomposition module. Then real-time data sharing is required. During the task execution process, each drone will regularly broadcast its location, electromagnetic environment data, and task progress status information through the data transmission component. The data compression module reduces the amount of data transmission to ensure real-time performance. If the direct communication link between a drone and the command node is interrupted, its data will be forwarded through other drone relays. During the execution of the task, collaborative task adjustments are made. For example, when a drone detects an enemy target signal, its status information will be broadcast to all drones, and the task allocation module will immediately adjust the interference subtask allocation. If a frequency conflict or task resource competition is found in the cluster, the collaborative optimization module will coordinate the drone actions through the communication module to avoid conflicts. Each node regularly checks the communication status with neighboring nodes. If the communication link is found to be disconnected, the network topology is immediately readjusted through the adaptive topology control unit, and an attempt is made to restore the disconnected link; if the radio communication fails completely, the backup communication link is switched. After the task is completed, all drones will summarize the task results to the task evaluation module through the communication module and report to the command center; if there is an unfinished task or an emergency, the task evaluation module will feed back information to the task decomposition module through the communication module and readjust the task plan as needed.

[0085] Furthermore, the reconnaissance subtask identifies enemy radars or other military equipment (in this invention, radar is used as an example) through the flight direction and directional reconnaissance direction of the drone, and needs to directly or indirectly cooperate with other drones in the process. The decision of the reconnaissance subtask depends on the state space, action space, and reward function.

[0086] In the environmental exploration mission, the position of the drone is expressed in relative coordinates, where the origin is defined as the location where the drone first detects radar information. At the same time, the direction of the reconnaissance is expressed in absolute angles, starting at 0° in the north direction and calculated clockwise. This setting helps the drone to make more accurate orientation judgments and navigation when exploring unknown environments.

[0087] The necessary processes for executing the reconnaissance sub-task include: approaching the target, hovering, and launching the reconnaissance payload. Therefore, the dynamics of the UAV reconnaissance sub-task are broken down into three dimensions.

[0088] 1) The flight direction actions are east, south, west, north, and hover, a total of 5 actions;

[0089] 2) The flight speed is divided into low speed, medium speed and high speed, a total of 3 actions;

[0090] 3) Directional reconnaissance rotation direction actions are left turn, right turn and unchanged, and the angle of each rotation is fixed at 45°, which is the angle range of directional reconnaissance, with a total of 3 actions. The reward function of the reconnaissance subtask consists of two parts: approach reward and reconnaissance reward.

[0091] The approach reward function in the reconnaissance subtask is:

[0092]

[0093] where d s is the distance line that the UAV must approach to the area where reconnaissance information can be obtained in order to complete the reconnaissance subtask, x represents the distance traveled by the UAV at the current moment compared to the previous moment (if negative, it means retreat), and d represents the distance between the UAV and the enemy radar.

[0094] When multiple drones collaborate as a team, you can get a team reconnaissance reward for each radar captured and located, with 1 point added for each radar.

[0095] Furthermore, the jamming subtask suppresses the enemy radar’s detection capability through the UAV’s flight direction, jamming target, jamming intensity and frequency band. In the process, it needs to cooperate directly or indirectly with other UAVs. The jamming subtask decision depends on the state space, action space and reward function.

[0096] During the mission, multiple drones fly toward the target area along a straight path. When the enemy radar signal is detected for the first time, the system will determine a reference position and set that position as the origin of the relative coordinate system. The decision-making process of the drone mainly relies on the comprehensive evaluation of the following key indicators: its own position information relative to the origin, the Euclidean distance from the enemy radar, the absolute angle relative to the radar direction (measured in a clockwise direction with reference to the north direction), and the degree of suppression of the enemy radar detection capability (quantified by the ratio of the current detection distance to the maximum detection distance). Through the integrated analysis of the above data, the drone can accurately formulate interference strategies and optimize the interference effect to cope with the complex and changing battlefield environment, thereby effectively ensuring the safety and execution of the mission objectives.

[0097] The necessary process of executing the jamming subtask includes: reconnaissance to determine the target location, selection of jamming targets, and determination of jamming intensity and frequency band. Therefore, the dynamics of the drone reconnaissance subtask are broken down into three dimensions.

[0098] 1) The flight direction actions are east, south, west, north, and hover, a total of 5 actions;

[0099] 2) The flight speed is divided into low speed, medium speed and high speed, a total of 3 actions;

[0100] 3) The interference actions are divided into zero-intensity interference, that is, no interference action, and the combinations of interference targets ranked from the first to the seventh threat level and weak, medium, and strong interference intensity actions, a total of 22 actions. The reward function of the interference subtask consists of two parts: approach reward and interference reward.

[0101] The approach reward function in the interference subtask is:

[0102]

[0103] Among them, d i is the maximum effective radius of the UAV interference load, x represents the distance of the UAV at the current moment compared to the previous moment (if negative, it means retreat), and d represents the distance between the UAV and the enemy radar.

[0104] The ratio of the reduction in radar detection range compared to its maximum detection range at each time step is the interference reward. When the enemy radar is completely suppressed and maintained, an additional reward is given according to the holding time.

[0105] Furthermore, the comprehensive decision is completed in two steps. First, the reconnaissance subtask and the interference subtask decide on the combination of flight action and reconnaissance action, and the combination of flight action and interference action respectively. Secondly, the comprehensive sub-strategy decides the weight of repeated flight actions according to the progress of the task, and makes a complete comprehensive decision with independent reconnaissance actions and interference actions. During the training process, a sample storage and extraction mechanism based on the experience pool (buffer) is designed. The sample data of each event (episode) will be stored in the experience pool. The capacity of the experience pool is equal to the length of a single episode. After the episode ends, the samples are processed to achieve efficient training. The Critic network captures the collaborative relationship between drones by processing global state information, which helps to improve the collaborative efficiency of the overall system.

[0106] M1: Initialize the experience pool buffer, the capacity of which is the length of a single event episode;

[0107] M2: Randomly initialize Actor network parameters θ π and Critic network parameters θπ , is a random number in the range [-0.003, 0.003];

[0108] M3: Initialize the reconnaissance sub-strategy π scout and the interference substrategy π jam ;

[0109] M4: When the episode is less than the total number of episodes, execute steps M4 to M14;

[0110] M5: Initialization state s1;

[0111] M6: When the episode length (time step t = 1, 2, ..., K) is less than K, execute steps M6 to M13;

[0112] M7: For UAV i=1,…,m, execute steps M7 to M12;

[0113] M8: Based on local observations from drone i Select comprehensive action

[0114] M9: Call the reconnaissance substrategy π scout , according to observation Select the reconnaissance action a i-scout ;

[0115] M10: Call the interference sub-strategy π jam , according to observation Select the interference action a i-jam ;

[0116] M11: put a i-scout and a i-jam According to the synthesis of complete

[0117] M12: Performing a complete joint movement Get rewards t and the next time step state s t+1 ;

[0118] M13: The experience sample (s t ,a t ,r t ,s t+1 ) is stored in the buffer, time step t = t + 1;

[0119] M14: Calculate the discounted return G for each time step t , and the number of episodes is increased by 1;

[0120] M15: Randomly sample n groups of sample experiences from the buffer;

[0121] M16: Calculate the Actor loss function L(θ) and update the Actor network;

[0122] M17: Calculate the Critic loss function And update the Critic network;

[0123] Among them, Actor loss function

[0124] The reward function is

[0125]

[0126] The advantage function is

[0127]

[0128] The critic loss function is

[0129] The ideal discounted return is

[0130] G t =r t+1 +γr t+2 +…+γ T-t r T+1

[0131] Error δ t =r t +γV(s t +1)-V(s t ), E is the expectation, γ is the discount factor, V(s t ) is the state s t The value function of θ (a t |s t ) represents the strategy of transferring from state s to action a at time t, λ is the weight, is the predicted discounted return.

[0132] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any ordinary technician in the industry can smoothly implement the present invention as shown in the drawings and described above. However, any equivalent changes, modifications and evolutions made by technicians familiar with the profession without departing from the scope of the technical solution of the present invention using the technical content disclosed above are all equivalent embodiments of the present invention. At the same time, any equivalent changes, modifications and evolutions made to the above embodiments based on the essential technology of the present invention are still within the protection scope of the technical solution of the present invention.

Claims

1. A collaborative reconnaissance and electronic countermeasure control system based on multiple UAVs, characterized in that: The system includes a communication module, a task decomposition module, a reinforcement learning decision module, a collaborative optimization module and a task evaluation module; The communication module includes a self-organizing network unit, a data synchronization unit and an encryption and anti-interference unit; The task decomposition module includes a task parsing unit, a task priority allocator and a resource allocation unit; The reinforcement learning decision module includes a reconnaissance strategy unit, an interference strategy unit and a strategy fusion and regulation unit; The collaborative optimization module includes a task collaboration management unit, a multi-machine action coordination unit, and a conflict detection and resolution unit; The task evaluation module includes a task state monitoring unit, an environment state evaluation unit and a feedback generation unit.

2. A collaborative reconnaissance and electronic countermeasure control system based on multiple UAVs, characterized in that: The task decomposition module transmits task instructions through the communication module and cooperates with the reinforcement learning decision module to obtain task decision results; The reinforcement learning decision module makes decisions based on the task requirements and the status of the drone, and cooperates with the collaborative optimization module to adjust the drone task allocation and flight path. The communication module transmits the task information, status information and decision instructions to each module, and feeds back the status of each drone to the task evaluation module; The task evaluation module provides feedback on task execution results and optimizes subsequent task allocation and strategy decisions.

3. A control method based on multi-UAV cooperative reconnaissance and electronic countermeasures is carried out using the system described in claims 1-2, characterized in that: The method comprises the following steps: Step S1: The self-organizing network unit communicates point-to-point through the self-organizing network protocol and dynamically adjusts the network topology according to the physical location of the drone; the data synchronization unit ensures the consistency of battlefield data shared between drones through the time synchronization protocol; The encryption and anti-interference unit enhances the link's anti-interference capability through frequency hopping communication; Step S2: The task analysis unit receives the electronic countermeasure task from the external command system, sends it to the UAV cluster, and breaks it down into a reconnaissance subtask and a jamming subtask according to the task type; Step S3: the task priority allocator adjusts the subtask priority according to the battlefield environment; Step S4: the resource allocation unit allocates resources to the UAV to adapt to different mission objectives; Step S5: The reconnaissance strategy unit combines DQN and PPO deep reinforcement learning to train the reconnaissance strategy, inputs state space and action space information during training, and outputs the best path or signal detection frequency; Step S6: The interference strategy unit optimizes interference frequency band selection and interference power to perform interference operations based on deep reinforcement learning input of enemy frequency band usage according to enemy communication or radar signal characteristics; Step S7: The task collaboration management unit is based on multi-agent reinforcement learning. Each UAV independently evaluates its own task adaptability and negotiates with other UAVs to complete task allocation. The multi-machine action coordination unit optimizes UAV actions based on the swarm algorithm, dynamically adjusts the flight path, and covers the target area. Step S7: the conflict detection and resolution unit detects possible task conflicts in the system, analyzes the costs and benefits of the conflicts based on a game theory model, and adjusts task allocation or interference parameters in real time; Step S8: The mission status monitoring unit dynamically monitors the reconnaissance coverage rate and interference effect index of the UAV to evaluate whether the mission meets the target requirements; The environmental status assessment unit collects the strength and frequency characteristics of the signals transmitted by the enemy equipment through sensors, and analyzes whether the enemy equipment has been effectively interfered; Step S9: Based on the evaluation results, return to step S3, where the task priority allocator adjusts the subtask priorities according to the battlefield environment and then repeats the subsequent steps.

4. The method according to claim 3, characterized in that: When the reconnaissance strategy of the reconnaissance subtask is trained, the input state space includes the position of the enemy signal and the current position information of the UAV, and the output action space includes the flight path adjustment and the reconnaissance signal frequency allocation.

5. The method according to claim 3, characterized in that: In the decision-making strategy in the interference subtask, the input state includes the signal frequency band and strength of the enemy radar, and the output action space includes the interference frequency, power selection and the selection of the target interference strength.

6. The method according to claim 3, characterized in that: The communication module performs communication initialization when each drone is started, discovers other drone nodes around it through broadcasting, and establishes a neighbor node list; Based on the neighbor node information, the self-organizing network component dynamically generates the topology of the cluster communication network, and the identity authentication module of the encryption and security component ensures that only legitimate nodes can join the network.

7. The method according to claim 3, characterized in that: The reconnaissance subtask decision process is controlled by the state space and action space data of the UAV and the set reward function.

8. The method according to claim 3, characterized in that: The position of the UAV in the state space is expressed using relative coordinates, where the origin is defined as the position where the UAV first detects radar information, and the direction of detection is expressed in absolute angles, starting from 0° in the north direction and calculated clockwise; The action space includes approaching a target, hovering, and launching a reconnaissance payload; The reward function includes approach reward and reconnaissance reward. When a multi-UAV team collaborates, each time a radar is captured and located, the team can obtain a reconnaissance reward, and each radar accumulates 1 point.

9. The method according to claim 3, characterized in that: The reconnaissance subtask and the interference subtask respectively decide on the combination of flight action and reconnaissance action, and the combination of flight action and interference action, and decide on the weight of repeated flight action according to the progress of the task, and reach a complete comprehensive decision with independent reconnaissance action and interference action.

10. The method according to claim 9, characterized in that: The sample data of each event in the deep learning training of the reconnaissance strategy unit is stored in the experience pool. After the event, the samples are processed by the Monte Carlo method, and the global state information is processed based on the Critic network to capture the collaborative relationship between drones: M1: Initialize the experience pool buffer, the capacity of which is the length of a single event episode; M2: Randomly initialize Actor network parameters θ π and Critic network parameters θ π , is a random number in the range [-0.003, 0.003]; M3: Initialize the reconnaissance sub-strategy π scout and the interference substrategy π jam ; M4: When the episode is less than the total number of episodes, execute steps M4 to M14; M5: Initialization state s1; M6: When the episode length (time step t = 1, 2, ..., K) is less than K, execute steps M6 to M13; M7: For UAV i=1,…,m, execute steps M7 to M12; M8: Based on local observations from drone i Select comprehensive action M9: Call the reconnaissance substrategy π scout , according to observation Select the reconnaissance action a i-scout ; M10: Call the interference sub-strategy π jam , according to observation Select the interference action a i-jam ; M11: put a i-scout and a i-jam According to the synthesis of complete M12: Performing a complete joint movement Get rewards t and the next time step state s t+1 ; M13: The experience sample (s t ,a t ,r t ,s t+1 ) is stored in the buffer, time step t = t + 1; M14: Calculate the discounted return G for each time step t , and the number of episodes is increased by 1; M15: Randomly sample n groups of sample experiences from the buffer; M16: Calculate the Actor loss function L(θ) and update the Actor network; M17: Calculate the Critic loss function And update the Critic network; Among them, Actor loss function The reward function is The advantage function is The critic loss function is The ideal discounted return is G t =r t+1 +γr t+2 +…+γ T-t r T+1 Error δ t =r t +γV(s t +1)-V(s t ), E is the expectation, γ is the discount factor, V(s t ) is the state s t The value function of θ (a t |s t ) represents the strategy of transferring from state s to action a at time t, λ is the weight, is the predicted discounted return.

Citation Information

Cited By

  • Intelligent logistics management system based on large language model

    CN120806768A

  • Multi-unmanned aerial vehicle subsystem real-time integrated control method and system based on edge calculation

    CN121477982A

  • Multi-unmanned aerial vehicle subsystem real-time integrated control method and system based on edge computing

    CN121477982B

  • Path planning method based on double Q-Learning

    CN121540145A