Unmanned aerial vehicle electronic countermeasure method and device, unmanned aerial vehicle and storage medium
By decomposing the drone electronic countermeasure mission into reconnaissance and jamming subtasks, dynamically calculating task priorities and adjusting actions, the redundancy and strategy conflict problems of flight actions in drone clusters are solved, and the collaborative efficiency and mission success rate of drone clusters are improved.
Patent Information
- Application Number
- CN202510887410.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing technology, the decision-making of UAV swarm electronic countermeasure missions has the problems of repeated flight action decision redundancy and strategy conflict, which leads to the failure of mission goal tendency.
The electronic countermeasure task is decomposed into reconnaissance subtasks and interference subtasks, and the corresponding subtasks are trained in the reinforcement learning neural network. The task priority is dynamically calculated, and the task dominance is determined according to the priority. The reinforcement learning neural network gives priority to executing the dominance subtask, reuses the flight actions of non-dominance subtasks, detects the flight path and interference frequency band conflicts within the drone cluster, and dynamically adjusts the actions to generate global actions.
It reduces the redundancy of repeated decisions in flight actions, improves the collaborative efficiency of drone clusters, and increases the mission success rate and execution accuracy.
Smart Images

Figure CN120729463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and in particular relates to a method and device for electronic countermeasures against unmanned aerial vehicles (UAVs), a UAV, and a storage medium. Background Art
[0002] With the rapid development of drone technology, drone swarms are increasingly being used in electronic warfare. On modern battlefields, drone swarms can effectively suppress enemy information technology equipment and disrupt enemy communications and combat capabilities through coordinated reconnaissance, jamming, and deception. However, due to the complex battlefield environment, diverse mission objectives, and the dynamic electromagnetic environment, drone swarms present significant challenges in electronic warfare decision-making.
[0003] In the process of implementing the present invention, the inventors discovered the following problems in the prior art: Publication No. CN120010549 A discloses a technical solution for implementing multi-UAV collaborative reconnaissance and jamming decision-making through subtask decomposition and reinforcement learning strategy optimization. This solution divides the entire task into a reconnaissance subtask and a jamming subtask. The reconnaissance subtask is divided into flight action and reconnaissance action, and the jamming subtask is divided into flight action and jamming action. A reinforcement learning algorithm is used to make decisions for the reconnaissance subtask and the jamming subtask separately. However, the two subtasks make repeated decisions on flight action, resulting in decision redundancy. Moreover, the flight action in the comprehensive strategy result is biased towards a certain task objective, which can easily cause the action for the other task objective to be isolated or even ineffective. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a drone electronic countermeasure method and device, a drone, and a storage medium, which solve the problems of redundancy and strategy conflict caused by repeated decision-making of flight actions in traditional multi-agent reinforcement learning.
[0005] In a first aspect, an embodiment of the present application provides a method for electronic countermeasures against unmanned aerial vehicles (UAVs), the method comprising: decomposing an electronic countermeasure task into a reconnaissance subtask and an interference subtask, and training the corresponding subtasks in a reinforcement learning neural network; wherein the reconnaissance subtask is used to output flight action and reconnaissance action, and the interference subtask is used to output flight action and interference action; dynamically calculating task priority, and determining task dominance of the reconnaissance subtask or the interference subtask based on the task priority; based on the task dominance, the reinforcement learning neural network prioritizes executing the dominance subtask, the dominance subtask generates the flight action, reconnaissance action, and interference action of the current UAV, and when executing a non-dominance subtask, the non-dominance subtask reuses the flight action generated by the dominance subtask; splicing the previous action of the current UAV and the current observation vector into the reinforcement learning neural network, and determining the comprehensive action of the current UAV based on the previous action and the current observation vector; detecting flight path conflicts and interference frequency band overlap conflicts between multiple UAVs in a UAV cluster, and dynamically adjusting the flight action and interference action in the comprehensive action; generating a global action corresponding to the UAV cluster based on the flight action, interference action, and reconnaissance action in the dynamically adjusted comprehensive action, the global action being a collection of adjusted comprehensive actions of all UAVs in the UAV cluster.
[0006] In combination with the first aspect, in certain implementations of the first aspect, the task priority is dynamically calculated, and the task dominance of the reconnaissance subtask or the interference subtask is determined based on the task priority, including: obtaining the enemy interference threat level detected in real time by the sensors deployed on the current UAV; obtaining the current reconnaissance coverage and interference success rate; calculating the priority score of the interference subtask and the priority score of the reconnaissance subtask based on the enemy interference threat level, the reconnaissance coverage and the interference success rate respectively: comparing the priority score of the interference subtask and the priority score of the reconnaissance subtask; if the priority score of the interference subtask is greater than the priority score of the reconnaissance subtask, then determining that the interference subtask has task dominance, and the interference subtask takes over the right to generate flight actions; if the priority score of the interference subtask is less than or equal to the priority score of the reconnaissance subtask, then determining that the reconnaissance subtask has task dominance, and the reconnaissance subtask takes over the right to generate flight actions.
[0007] In combination with the first aspect, in certain implementations of the first aspect, before splicing the preceding action of the current UAV with the current observation vector and inputting it into the reinforcement learning neural network, it also includes: obtaining battlefield environment data collected in real time by sensors arranged on the current UAV to generate a current observation vector, wherein the current observation vector includes its own position coordinates, its own flight speed, enemy position coordinates, enemy interference threat level, and enemy interference intensity; obtaining the action of the current UAV in the previous time step to generate a preceding action; normalizing the current observation vector and encoding the preceding action into a vector; the reinforcement learning neural network includes an Actor network and a Critic network, splicing the preceding action of the current UAV with the current observation vector and inputting it into the reinforcement learning neural network , the comprehensive action of the current drone is determined according to the previous action and the current observation vector, including: for the Actor network: the input layer is used to input the current observation vector and the previous action and splice them; the network layer captures the temporal dependency between actions based on the LSTM module, receives the input results of the input layer and calculates the hidden state value; the output layer is used to output the flight action, reconnaissance action and interference action of the current drone generated by the fully connected layer; for the Critic network: the input layer is used to input the global state; among them, the global state includes the splicing of the current observation vector and the previous action of all drones; the network layer is used to calculate the weight matrix; the output layer is used to output the global value generated by the fully connected layer, as well as the randomly initialized Actor network parameters and Critic network parameters.
[0008] In combination with the first aspect, in certain implementations of the first aspect, when it is determined that the reconnaissance subtask takes the lead in the task, the reconnaissance subtask performs the following steps: input the vector obtained by concatenating the current observation vector and the preceding action; extract the timing features through the LSTM module of the network layer of the Actor network; output the flight action dominated by the reconnaissance subtask, the reconnaissance action generated by the Actor network; the interference subtask performs the following steps: input the vector obtained by concatenating the current observation vector and the preceding flight action; generate an interference action through the fully connected layer of the Actor network; and output the interference action of the reused flight action.
[0009] In combination with the first aspect, in some implementations of the first aspect, the comprehensive action of the current UAV is determined according to the previous action and the current observation vector, including: outputting the hidden state value through the LSTM module of the network layer of the Actor network; if the reconnaissance subtask dominated the task in the previous time step and the reconnaissance subtask currently dominates the task, the fully connected layer reuses the flight action generated by the reconnaissance subtask in the previous time step, and generates the reconnaissance action of the current reconnaissance subtask; if the interference subtask dominated the task in the previous time step and the reconnaissance subtask currently dominates the task, the fully connected layer reuses the flight action generated by the reconnaissance subtask in the previous time step ... The fully connected layer reuses the flight actions generated by the interfering subtask at the previous time step, and generates the reconnaissance actions of the current reconnaissance subtask; if the reconnaissance subtask dominated the task in the previous time step, and the interference subtask currently dominates the task, then the fully connected layer reuses the flight actions generated by the reconnaissance subtask at the previous time step, and generates the interference actions of the current interference subtask; if the interference subtask dominated the task in the previous time step, and the interference subtask currently dominates the task, then the fully connected layer reuses the flight actions generated by the interference subtask at the previous time step, and generates the interference actions of the current interference subtask; finally, the comprehensive action of the current UAV is output.
[0010] In combination with the first aspect, in certain implementation methods of the first aspect, flight path conflicts and interference frequency band overlap conflicts between multiple drones in a drone cluster are detected, the flight actions and interference actions in the comprehensive action are dynamically adjusted, and based on the flight actions and interference actions in the dynamically adjusted comprehensive action, as well as the reconnaissance action, a global action corresponding to the drone cluster is generated, including: for all drones in the drone cluster, inputting a set of flight directions and flight speeds of all drones; inputting a set of interference threat levels and interference intensities of all drones; calculating the Euclidean distance between each drone, and if the Euclidean distance between the drones does not meet the preset safety distance threshold, it is marked as a flight path conflict; counting the number of times the interference threat level and interference intensity are the same, and if the number of times is greater than 1, it is marked as an interference frequency band overlap conflict; outputting a conflict flag set, the conflict flag set including the conflict type and the ID of the drones involved; calculating the attention weight matrix, and reallocating the actions according to the attention weight, so as to dynamically adjust the flight actions and interference actions in the comprehensive action; for flight path conflicts, adjusting the flight direction; for interference frequency band overlap conflicts, assigning new interference threat levels and interference intensities; and outputting the coordinated global action.
[0011] In combination with the first aspect, in certain implementations of the first aspect, after generating the global action corresponding to the drone cluster based on the flight action and the interference action in the dynamically adjusted comprehensive action, as well as the reconnaissance action, it also includes: in the training of the reinforcement learning neural network, expanding the experience pool to store action-dependent data, and optimizing the reinforcement learning neural network through collaborative rewards and conflict penalties; expanding the experience pool to store action-dependent data, and optimizing the reinforcement learning neural network through collaborative rewards and conflict penalties, including: expanding the experience pool to store action-dependent data using five-tuple data, the five-tuple data including the global state, the action executed in the current time step, the original reward value of the environment feedback, the next time step The global state and preceding action of the inter-step; pre-processing operations on the five-tuple data and encapsulating it into a standardized format, the pre-processing operations include denoising and removing outliers; adding conflict flags and collaborative flags, allocating storage space, and storing extended samples in time step order; inputting the collaborative flag value, that is, counting the number of reuses of flight actions; inputting the conflict flag value, that is, counting the total number of all conflicts; calculating the collaborative reward increment based on the number of reuses of flight actions; using the collaborative reward increment to update the reward value; calculating the penalty term based on the total number of all conflicts and the penalty coefficient; based on the collaborative reward increment, reward value and penalty term, updating and outputting the final reward value to optimize the reinforcement learning neural network.
[0012] On the second aspect, an embodiment of the present application provides an unmanned aerial vehicle electronic countermeasure device, which includes: a decomposition module for decomposing the electronic countermeasure task into a reconnaissance subtask and an interference subtask, and training the corresponding subtasks in a reinforcement learning neural network; wherein the reconnaissance subtask is used to output flight actions and reconnaissance actions, and the interference subtask is used to output flight actions and interference actions; a calculation module for dynamically calculating the task priority, and determining the task dominance of the reconnaissance subtask or the interference subtask according to the task priority; a reuse module for, according to the task dominance, reinforcing the learning neural network to give priority to executing the dominance subtask, the dominance subtask generates the flight action, reconnaissance action and interference action of the current unmanned aerial vehicle, and when executing When performing a non-dominant subtask, the non-dominant subtask reuses the flight action generated by the dominant subtask; the determination module is used to splice the previous action of the current drone and the current observation vector into the reinforcement learning neural network, and determine the comprehensive action of the current drone based on the previous action and the current observation vector; the detection module is used to detect flight path conflicts and interference frequency band overlapping conflicts between multiple drones in the drone cluster, and dynamically adjust the flight actions and interference actions in the comprehensive action; the generation module is used to generate the global action corresponding to the drone cluster based on the flight actions and interference actions in the dynamically adjusted comprehensive action, as well as the reconnaissance action. The global action is the set of adjusted comprehensive actions of all drones in the drone cluster.
[0013] In a third aspect, an embodiment of the present application provides a drone, comprising: a processor and a memory, the memory storing a computer program that can be executed by the processor, and the processor executing the computer program to implement the drone electronic countermeasure method as described in the first aspect.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer program instructions, and the computer program instructions are used to execute the drone electronic countermeasure method of the first aspect.
[0015] Compared with the prior art, the embodiment of the present invention dynamically calculates task priorities and determines the task dominance of the reconnaissance subtask or the interference subtask based on the task priority; according to the task dominance, the reinforcement learning neural network prioritizes the execution of the dominance subtask, and the dominance subtask generates the flight action, reconnaissance action and interference action of the current UAV. When executing the non-dominance subtask, the non-dominance subtask reuses the flight action generated by the dominance subtask; it overcomes the redundancy and strategy conflict problems caused by repeated decision-making of flight actions in traditional multi-agent reinforcement learning, realizes the reuse of flight actions between reconnaissance and interference tasks, reduces computing resource consumption, and improves the collaborative efficiency of UAV clusters. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a flow chart of a UAV electronic countermeasure method provided by an exemplary embodiment of the present application.
[0017] Figure 2 The figure shows a schematic diagram of the execution flow of a reconnaissance subtask when the reconnaissance subtask takes the leading position in the task, provided by an exemplary embodiment of the present application.
[0018] Figure 3 The figure shows a schematic diagram of the execution flow of the interference subtask when the reconnaissance subtask takes the leading position in the task, provided by an exemplary embodiment of the present application.
[0019] Figure 4 Shown is a logical schematic diagram of a UAV electronic countermeasure method provided by an exemplary embodiment of the present application.
[0020] Figure 5 Shown is a schematic structural diagram of a drone electronic countermeasure device provided by an exemplary embodiment of the present application.
[0021] Figure 6 Shown is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] Figure 1 The figure shows a flow chart of a UAV electronic countermeasure method provided by an exemplary embodiment of the present application. Figure 1 The present invention provides an embodiment of a UAV electronic countermeasure method, comprising:
[0024] Step S100: decompose the electronic countermeasure task into a reconnaissance subtask and a jamming subtask, and train the corresponding subtasks in a reinforcement learning neural network.
[0025] Among them, the reconnaissance subtask is used to output flight actions and reconnaissance actions, and the interference subtask is used to output flight actions and interference actions.
[0026] For example, an electronic warfare mission refers to a task in which humans and / or computers collect battlefield information, analyze and make decisions, and then hand it over to computers for execution.
[0027] Step S200 : dynamically calculating the task priority and determining the task dominance of the reconnaissance subtask or the interference subtask according to the task priority.
[0028] In step S300, based on the task dominance, the reinforcement learning neural network prioritizes executing the dominance subtask. The dominance subtask generates the flight action, reconnaissance action, and interference action of the current drone. When executing the non-dominance subtask, the non-dominance subtask reuses the flight action generated by the dominance subtask.
[0029] In step S400, the preceding action of the current drone and the current observation vector are spliced and input into a reinforcement learning neural network, and the comprehensive action of the current drone is determined based on the preceding action and the current observation vector.
[0030] Specifically, the drone's previous actions and current observation vectors can be spliced and input into the reinforcement learning neural network through action dependency modeling.
[0031] For example, a reinforcement learning neural network can employ an actor-critic network architecture. The actor network includes a long short-term memory (LSTM) module to capture temporal dependencies. The actor network takes as input the concatenated observation vector and preceding actions and outputs flight actions and subtask parameters. The critic network uses a multi-head attention mechanism to calculate a weight matrix and output the collaborative value of multiple drones.
[0032] Exemplarily, the preceding action of the drone is the action in the previous time step, including flight, reconnaissance, and jamming actions.
[0033] Exemplarily, the current observation vector is battlefield environment data collected in real time by the drone sensor, including its own position coordinates, its own flight speed, enemy position coordinates, enemy interference threat level, and enemy interference intensity.
[0034] Step S500: Detect flight path conflicts and interference frequency band overlap conflicts among multiple drones in the drone cluster, and dynamically adjust the flight actions and interference actions in the comprehensive action.
[0035] Step S600: Based on the flight action, interference action, and reconnaissance action in the dynamically adjusted comprehensive action, a global action corresponding to the drone cluster is generated. The global action is a collection of the adjusted comprehensive actions of all drones in the drone cluster.
[0036] Specifically, the task priority can be dynamically calculated based on the enemy interference threat level, reconnaissance coverage rate and interference success rate, and the dominance of the reconnaissance or the interference subtask can be switched; according to the observation vector oti of UAV i and the preceding action Select comprehensive action The attention mechanism is used to detect flight path conflicts and interference frequency band overlap conflicts among multiple drones, and dynamically adjust flight actions and interference actions; a lightweight coordination network {a i-assemble}, receive the flight actions and interference actions of all drones, and generate the coordinated global action a global .
[0037] Furthermore, the goal of the reconnaissance subtask is for the UAV to identify enemy radars or other military equipment through flight direction and directional reconnaissance direction. In the embodiment of the present invention, radar is taken as an example, and direct or indirect collaboration with other UAVs is required during the process. The reconnaissance subtask includes a state space, an action space, and a reward function.
[0038] The state space of the reconnaissance subtask is defined as the position of the UAV, which can be expressed using relative coordinates. The origin is defined as the location where the UAV first detects radar information. The reconnaissance direction is expressed in absolute angles, starting at 0° in the north direction and counting clockwise.
[0039] The necessary process of the action space of the reconnaissance subtask includes: approaching the target - hovering - launching the reconnaissance payload. Therefore, the UAV reconnaissance subtask can be dynamically decomposed into the following three dimensions:
[0040] 1) The flight direction actions are east, south, west, north, and hover, a total of 5 actions;
[0041] 2) Flight speed is divided into three actions: low speed, medium speed and high speed;
[0042] 3) The rotation direction actions for directional reconnaissance are left turn, right turn and unchanged. The angle of each rotation is fixed at 45°, which is the angle range of directional reconnaissance. There are 3 actions in total.
[0043] The reward function of the reconnaissance subtask includes approach reward and reconnaissance reward.
[0044] Proximity reward: Set d s is the distance line that the UAV must approach to the area where reconnaissance information can be obtained when completing the reconnaissance mission, x represents the distance the UAV has traveled at the current moment compared to the previous moment (if negative, it means retreat), and d represents the distance between the UAV and the enemy radar. The proximity reward function of the reconnaissance subtask is:
[0045]
[0046] Reconnaissance Rewards: When a multi-drone team collaborates, each time a radar is captured and located, a team reconnaissance reward is obtained, with +1 point for each radar.
[0047] Furthermore, the goal of the jamming subtask is for the UAV to suppress the enemy radar’s detection capability through flight direction, jamming target, jamming intensity, and frequency band. In the process, it needs to collaborate directly or indirectly with other UAVs. The jamming subtask includes state space, action space, and reward function.
[0048] The state space of the jamming subtask is defined as follows: determining the origin of the coordinate system and calculating the evaluation indicators. The evaluation indicators include the position information relative to the origin, the Euclidean distance between the system and the enemy radar, the absolute angle relative to the radar direction (measured in a clockwise direction with reference to the north direction), and the degree of suppression of the enemy radar detection capability (quantified by the ratio of the current detection distance to the maximum detection distance).
[0049] The necessary process of the jamming subtask’s action space includes: reconnaissance to determine the target location – selecting the jamming target – determining the jamming intensity and frequency band. Therefore, the dynamics of the drone reconnaissance subtask can be broken down into the following three dimensions:
[0050] 1) The flight direction actions are east, south, west, north, and hover, a total of 5 actions;
[0051] 2) Flight speed is divided into three actions: low speed, medium speed and high speed;
[0052] 3) Interference actions are divided into zero-intensity interference (i.e., no interference action), and permutations and combinations of interference targets ranked from the first to the seventh threat level and weak, medium, and strong interference intensity actions, for a total of 22 actions;
[0053] The reward function of the interference subtask includes proximity reward and interference reward:
[0054] Proximity reward: Set d i is the maximum effective radius of the UAV jamming payload, then the proximity reward function of the jamming subtask is:
[0055]
[0056] Interference reward: It is defined as the ratio of the reduction in radar detection range compared to its maximum detection range at each time step. When the enemy radar is completely suppressed and maintained, an additional reward is given according to the holding time.
[0057] The UAV electronic countermeasure method provided by the embodiment of the present invention generates flight actions and subtask parameters by splicing the UAV's previous actions and current observation vectors into an LSTM neural network, thereby realizing the reuse of flight actions between reconnaissance and interference tasks, reducing computing resource consumption, and improving the collaborative efficiency of UAV clusters.
[0058] Furthermore, the specific execution steps of the action-dependent multi-agent reinforcement learning algorithm include:
[0059] Step 1: Initialize the experience pool (buffer) with a capacity equal to the length of a single episode; store the preceding action a in the buffer t-1 , that is, the sample format is (s t , a t , r t , s t+1 , a t-1 ), for temporal learning of action dependencies;
[0060] Step 2: The input layer of the Actor network takes the current observation vector o t i With the preceding action a t-1 Splicing, that is The network layer uses LSTM modules to capture the temporal dependencies between actions; the output layer is divided into two parts, one is the current flight action, where the dominance is dynamically switched between the reconnaissance and interference subtasks, and the other is the reconnaissance action and the interference action; the input layer of the critic network is the global state s t; The network layer adds an attention mechanism to calculate the contribution weight of different drone actions to the global value; Randomly initialize the Actor network parameters θ π and Critic network parameter Φ ν ,θ π and Φ ν is a random number in the range [-0.003, 0.003];
[0061] Step 3: Initialize the reconnaissance subtask π scout and the interference subtask π jam ;Reconnaissance subtask π scout The flight action after the decision will be interfered with by subtask π jam Multiplexing: If the interfering subtask requires adjusting the flight action, the flight action dominance is dynamically switched through the priority mechanism;
[0062] Step 4: When the episode is less than the total number of episodes, execute steps 4 to 14;
[0063] Step 5: Initialize state s1;
[0064] Step 6: When the episode length (time step, t = 1, 2, ..., K) is less than K, execute steps 6 to 13;
[0065] Step 7: For UAV i=1,...,m, execute steps 7 to 12;
[0066] Step 8: According to the current observation vector o of drone i t i With preceding action Select comprehensive action The comprehensive action includes flight action, reconnaissance action, and interference action. The flight action is generated by the reconnaissance subtask and includes flight direction and flight speed. In addition, the reconnaissance subtask supplements the reconnaissance action, and the interference subtask reuses the flight action and only supplements the interference action.
[0067] Step 9: Call the reconnaissance subtask π scout , according to the current observation vector o t i With preceding action Select the reconnaissance action of multiplexing flight action a i-scout ;
[0068] Step 10: Call the interference subtask π jam , according to the current observation vector o t i With preceding action Select the interference action a of multiplexed flight action i-jam ;
[0069] If the interference subtask π jam If flight actions need to be adjusted, the gating network is used to adjust the action weights and dynamically switch the dominant position.
[0070] Step 11: Place a i-scout and a i-jam According to a i-assmeble Synthesize into a complete Introducing lightweight coordination network {a i -assmeble}, receive action proposals from all drones, calculate the conflict costs between multiple drones, the conflict types include flight path conflicts and interference band overlap conflicts, and generate coordinated global actions a global ;
[0071] Step 12: Perform the complete joint Get rewards t and the next time step state s t+1 If the flight action of drone i is reused by other drones (e.g., a jamming mission uses a reconnaissance path), the reward increases by Δrcoop; if the flight action switches frequently (e.g., reconnaissance and jamming compete for dominance), the penalty term is -λ×the number of switches;
[0072] Step 13: The experience sample (s t ,a t ,r t ,s t+1 ) is stored in the buffer, time step t = t + 1;
[0073] Step 14: Calculate the discounted return G at each time step t , and the episode number +1;
[0074] Step 15: Randomly sample n groups of sample experiences from the buffer;
[0075] Step 16: Calculate the loss function L(θ) and update the Actor network;
[0076] Step 17: Calculate the loss function And update the Critic network;
[0077] Where E is the expectation; γ is the discount factor; V(s t ) is s t The value function of the state; π θ (a t |s t ) represents the strategy of transferring from state s to action a at time t; λ is the weight; ν st is the predicted discounted return;
[0078] The Actor loss function is:
[0079] award:
[0080] Advantage function:
[0081] Error: δ t =r t +γV(s t +1)-V(s t );
[0082] Critic loss function:
[0083] Ideal state discounted return: G t =r t+1 +γr t+2 +...+γ T-t r T+1 .
[0084] Furthermore, in step 1, the process of initializing the experience pool and storing action dependency extensions is as follows:
[0085] Get the global state of the drone cluster, that is, the splicing of all drones' observation data and execution actions t ;
[0086] Get the action a executed at the current time step t , namely flight actions and subtask parameters, subtask parameters include reconnaissance actions and jamming actions;
[0087] Get the original reward r of the environment feedback t ;
[0088] Get the new global state s after executing the action t+1 , that is, the next state;
[0089] Get the action of the previous time step, that is, the preceding action a t-1 ;
[0090] The five-tuple data (s t ,a t ,r t ,s t+1 ,a t-1 ) Perform preprocessing operations such as denoising and removing outliers, and encapsulate them into a standardized format;
[0091] Add conflict flag F conflict , used for flight path collision or interference threat and intensity overlap, add coordination flag F coop, used for flight actions to be reused by other UAVs or by the interference subtask of the current UAV;
[0092] Allocate storage space with a capacity of Nepisode×T, where Nepisode is the length of a single episode and T is the total number of episodes;
[0093] The extended samples are stored in time step order, with the index marked by timestamp t.
[0094] The UAV electronic countermeasure method provided by the embodiment of the present invention includes experience pool initialization and action-dependent extended storage, which accelerates the training of reinforcement learning neural networks and improves the robustness of training.
[0095] Furthermore, in step 2, the process of initializing and optimizing the Actor and Critic networks is as follows:
[0096] Get the battlefield environment data collected by the drone sensor in real time, that is, the current observation vector o t i ;
[0097] Get the action of the previous time step, that is, the preceding action a t-1 ;
[0098] For the current observation vector o t i Normalization, that is, converting the position coordinates into the offset of the coordinate origin;
[0099] The preceding action is encoded as a vector, the flight direction is encoded as binary [0,0,0], the flight speed is encoded as binary [0,0], the directional reconnaissance rotation direction is encoded as binary [0,0], and the interference action is encoded as binary [0,0,0,0,0]. Among them, the binary "0 to 4" of the flight direction represents the 5 actions of the flight direction, and the others are meaningless; the binary "0 to 2" of the flight speed represents the 3 actions of the flight speed, and the others are meaningless; the binary "0 to 2" of the directional reconnaissance rotation direction represents the 3 actions of the reconnaissance action, and the others are meaningless; the binary "0 to 21" of the interference action represents the 22 actions of the interference action, and the others are meaningless;
[0100] For Actor Networks:
[0101] Input layer: input the current observation vector o t i With the preceding action a i t-1 Splicing, that is
[0102] Network layer: LSTM module is used to capture the temporal dependency between actions, receiving xi t And calculate the hidden state value h i t =LSTM(h i t-1 ,x i t );
[0103] Output layer: Output fully connected layer to generate flight action The fully connected layer generates reconnaissance or jamming actions
[0104] For the Critic network:
[0105] Input layer: input global state s t ;
[0106] Network layer: Attention mechanism calculates weight matrix Where Q and K are trainable parameter matrices used to generate query and key vectors;
[0107] Output layer: Output the fully connected layer to generate the global value V(s t );
[0108] Output randomly initialized Actor network parameters θ π and Critic network parameter Φ ν .
[0109] Furthermore, in step 3, the process of subtask initialization and action dominance allocation is as follows:
[0110] Figure 2 The figure shows a schematic diagram of the execution flow of the reconnaissance subtask when the reconnaissance subtask takes the leading position in the task provided by an exemplary embodiment of the present application. Figure 2 As shown, when the reconnaissance subtask takes the lead in the task, the reconnaissance subtask is initialized:
[0111] Step S301: Input the current observation vector o t i With preceding action The concatenated vector
[0112] Step S302: Extracting time series features through the LSTM module of the network layer of the Actor network
[0113] Step S303: Output the flight action dominated by the reconnaissance subtask
[0114] Figure 3The figure shows a schematic diagram of the execution flow of the interference subtask when the reconnaissance subtask takes the leading position in the task provided by an exemplary embodiment of the present application. Figure 3 As shown in Figure 2, when the reconnaissance subtask takes the lead, the interference subtask is initialized:
[0115] Step S311: Input the current observation vector o t i With forward flight action Spliced
[0116] Step S312: Generate interference actions through the fully connected layer of the Actor network
[0117] Step S313: Outputting the interference action of the multiplexed flight action
[0118] Furthermore, in step 8, the process of comprehensive action generation and dependency injection is as follows:
[0119] Output from the LSTM module of the network layer of the Actor network
[0120] Generate flight actions through fully connected layers
[0121] If the current subtask is reconnaissance, the fully connected layer generates reconnaissance actions
[0122] If the current subtask is interference, the fully connected layer generates interference actions
[0123] Output comprehensive action
[0124] Furthermore, in step 9, the reconnaissance subtask is called as follows:
[0125] Enter the current observation vector With forward flight action Spliced
[0126] If the task that was executed in the previous time step was dominated by the reconnaissance subtask, and the task that is currently dominated is the reconnaissance subtask, then execute the following steps:
[0127] Generate reconnaissance actions through the fully connected layers of the Actor network
[0128] Output reconnaissance maneuvers that reuse flight maneuvers
[0129] If the interference subtask dominated the task in the previous time step, and the reconnaissance subtask currently dominates the task, execute the following steps:
[0130] Generate reconnaissance actions through the fully connected layers of the Actor network
[0131] Generate flight actions through the fully connected layer of the Actor network
[0132] Output reconnaissance actions including flying actions
[0133] Furthermore, in step 10, the process of interfering with the subtask call is as follows:
[0134] Input the current observation vector o t i With forward flight action Spliced
[0135] If the task that was executed at the previous time step was dominated by the interfering subtask, and the task that is currently dominated is the interfering subtask, then execute the following steps:
[0136] Generate interference actions through the fully connected layers of the Actor network
[0137] Output interference actions of multiplexed flight actions
[0138] If the task that was executed in the previous time step was dominated by the reconnaissance subtask, and the current task is dominated by the interference subtask, then execute the following steps:
[0139] Generate interference actions through the fully connected layers of the Actor network
[0140] Generate flight actions through the fully connected layer of the Actor network
[0141] Output interference actions including flight actions
[0142] Subtask initialization and action dominance allocation solve the redundancy of flight action decision-making and improve the execution efficiency of the algorithm.
[0143] Furthermore, in step 10, the process of dynamic dominance switching is as follows:
[0144] Obtain the interference threat level S of the enemy target detected by the sensor in real time, with interference threat levels ranging from 1 to 7;
[0145] Obtain the current reconnaissance coverage C and jamming success rate R;
[0146] Calculate the priority score:
[0147] P scout =α·C+(1-α)·(1-S / 7);
[0148] P jam =β·S / 7+(1-β)·R;
[0149] Among them, α and β are weight parameters, and the value range of α and β is [0,1], which are optimized through machine learning training;
[0150] If P jam >P scout +δ, δ is the switching threshold, set to δ = 0.1 to prevent frequent switching), the interference subtask takes over the flight action generation right and performs the following operations:
[0151] Enter o t i To interfere with subtask π jam , generate interference subtask dedicated flight action
[0152] Update comprehensive actions
[0153] Output dynamically adjusted flight actions
[0154] If P jam ≤P scout +δ, the subtask dominance is not switched, and the reconnaissance subtask remains as the dominant subtask, generating flight actions.
[0155] The UAV electronic countermeasure method provided by the embodiment of the present invention solves the problem that fixed priority allocation cannot adapt to the dynamic battlefield environment through dynamic dominance switching, realizes real-time optimization allocation of task weights, ensures priority execution of interference tasks in high-threat scenarios, and improves the mission success rate.
[0156] Furthermore, in step 11, the coordinated global action a is generated global The process is as follows:
[0157] Input the flight direction and speed of all drones
[0158] Input the set of planned interference threat and interference intensity of all drones
[0159]
[0160] Calculate the Euclidean distance d between dronesij =||x i -x j ||2, if d ij <d safe , it is marked as a flight path conflict, where d safe is the preset safety distance threshold, d safe The value range is [3,5m];
[0161] Count the number of times the interference threat level and intensity are the same If N k >1, it is marked as interference frequency band overlap conflict;
[0162] Output conflict flag set {conflict type, involved drone ID};
[0163] Calculate the attention weight matrix
[0164] Redistribute actions by weight:
[0165] Adjust flight direction for flight path conflicts
[0166] For overlapping interference band conflicts, assign new threat levels and intensities Where rank(W i ) is the priority ranking of drone i in the attention weight;
[0167] Output coordinated global actions
[0168] The UAV electronic countermeasure method provided by the embodiment of the present invention solves the problem of sub-task execution conflicts within the UAV cluster by detecting flight path conflicts and interference frequency band overlapping conflicts, thereby improving the accuracy of electronic countermeasure task implementation.
[0169] Furthermore, in step 12, the process of dynamic adjustment of the reward function is as follows:
[0170] Enter the collaborative flag F coop , count the number of flight action reuses N coop ;
[0171] Input conflict flag F conflict , count the total number of all conflicts N conflict ;
[0172] Calculate the incremental collaborative reward: Δr coop =η·N coop , η∈[0,1] where is the collaborative reward coefficient, which is set to 0.05 in the embodiment of the present invention;
[0173] Update reward value: rt '=r t +Δr coop ;
[0174] Calculate the penalty term: Δr con =-λ·N conflict , where λ∈[0,1] is the penalty coefficient, which is set to 0.1 in the embodiment of the present invention;
[0175] Update and output the final reward value: r t '=r t +Δr coop +Δr con .
[0176] The drone electronic countermeasure method provided by the embodiment of the present invention sets a dynamic adjustment of the reward function so that the optimization results of the reinforcement learning neural network training select actions that tend to increase coordination and reduce conflict, thereby improving the accuracy of the neural network model.
[0177] Figure 4 The figure shows a logic diagram of a drone electronic countermeasure method provided by an exemplary embodiment of the present application. Figure 4 As shown, the electronic countermeasure task includes a reconnaissance subtask 401 and a jamming subtask 402. The task priority scores are compared. The task with a higher priority score obtains task dominance and becomes the dominance subtask 403, while the task with a lower priority score becomes the non-dominance subtask 404. The dominance subtask generates subtask parameters 405 and flight actions 406. The subtask parameters 405 can be reconnaissance actions or jamming actions. The subtask parameters 405 depend on whether the dominance subtask 403 is the reconnaissance subtask 401 or the jamming subtask 402. The non-dominance subtask 404 generates subtask parameters 407 for the reused flight action 406.
[0178] Figure 5 The figure shows a schematic diagram of the structure of a drone electronic countermeasure device provided by an exemplary embodiment of the present application. Figure 5As shown, the UAV electronic countermeasure device provided in an embodiment of the present application includes a decomposition module 501, a calculation module 502, a reuse module 503, a determination module 504, a detection module 505, and a generation module 506. The decomposition module 501 is configured to decompose the electronic countermeasure task into a reconnaissance subtask and a jamming subtask, and train the corresponding subtasks in a reinforcement learning neural network. The reconnaissance subtask is configured to output flight maneuvers and reconnaissance maneuvers, while the jamming subtask is configured to output flight maneuvers and jamming maneuvers. The calculation module 502 is configured to dynamically calculate task priorities and determine the task dominance of the reconnaissance or jamming subtask based on the task priority. The reuse module 503 is configured to prioritize the execution of the dominance subtask by the reinforcement learning neural network based on the task dominance. The dominance subtask generates the flight maneuvers, reconnaissance maneuvers, and jamming maneuvers for the current UAV. When executing non-dominance subtasks, the non-dominance subtasks reuse the flight maneuvers generated by the dominance subtask. The determination module 504 is configured to concatenate the current UAV's previous maneuvers and the current observation vector into the reinforcement learning neural network, and determine the current UAV's comprehensive maneuver based on the previous maneuvers and the current observation vector. Detection module 505 is used to detect flight path conflicts and interference frequency band overlap conflicts between multiple drones in the drone cluster, and dynamically adjust the flight and interference actions in the integrated action. Generation module 506 is used to generate a global action corresponding to the drone cluster based on the flight and interference actions in the dynamically adjusted integrated action, as well as the reconnaissance action. The global action is the set of adjusted integrated actions for all drones in the drone cluster.
[0179] Exemplary electronic devices
[0180] Below, reference Figure 6 To describe the electronic device according to the embodiment of the present application. Figure 6 Shown is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application.
[0181] like Figure 6 As shown, the electronic device 60 includes one or more processors 601 and a memory 602 .
[0182] The processor 601 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 60 to perform desired functions.
[0183] The memory 602 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 601 may execute the program instructions to implement the drone electronic countermeasure method of each embodiment of the present application described above and / or other desired functions. Various contents such as data generated by the reinforcement learning neural network calculation process may also be stored in the computer-readable storage medium.
[0184] In one example, the electronic device 60 may further include an input device 603 and an output device 604 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0185] The input device 603 may include, for example, a keyboard, a mouse, and the like.
[0186] The output device 604 can output various information to the outside, including the determined line of sight estimation information, etc. The output device 604 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, etc.
[0187] Of course, to simplify, Figure 6 Only some of the components related to the present application in the electronic device 60 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device 60 may further include any other appropriate components according to specific application scenarios.
[0188] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the drone electronic countermeasure method according to various embodiments of the present application described above in this specification.
[0189] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0190] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the drone electronic countermeasure method according to various embodiments of the present application described above in this specification.
[0191] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0192] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.
[0193] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0194] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.
[0195] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0196] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A drone electronic countermeasure method, characterized in that: include: Decomposing the electronic countermeasure task into a reconnaissance subtask and a jamming subtask, and training the corresponding subtasks in a reinforcement learning neural network; wherein the reconnaissance subtask is used to output flight maneuvers and reconnaissance maneuvers, and the jamming subtask is used to output flight maneuvers and jamming maneuvers; Dynamically calculating a task priority, and determining a task dominance of the reconnaissance subtask or the interference subtask based on the task priority; According to the task dominance, the reinforcement learning neural network prioritizes executing the dominance subtask, which generates the flight action, reconnaissance action, and interference action of the current UAV. When executing the non-dominance subtask, the non-dominance subtask reuses the flight action generated by the dominance subtask; splicing the preceding action of the current drone and the current observation vector into the reinforcement learning neural network, and determining the comprehensive action of the current drone based on the preceding action and the current observation vector; Detect flight path conflicts and interference frequency band overlap conflicts among multiple drones in a drone cluster, and dynamically adjust the flight actions and interference actions in the comprehensive action; Based on the flight action and the interference action in the integrated action after dynamic adjustment, as well as the reconnaissance action, a global action corresponding to the drone cluster is generated. The global action is a collection of the adjusted integrated actions of all drones in the drone cluster.
2. The method according to claim 1, characterized in that The dynamically calculating task priority and determining the task dominance of the reconnaissance subtask or the interference subtask according to the task priority includes: Obtaining an enemy interference threat level detected in real time by sensors disposed on the current UAV; Obtain current reconnaissance coverage and jamming success rate; Calculating the priority score of the interference subtask and the priority score of the reconnaissance subtask respectively according to the enemy interference threat level, the reconnaissance coverage rate, and the interference success rate; comparing the priority score of the interference subtask with the priority score of the reconnaissance subtask; If the priority score of the interference subtask is greater than the priority score of the reconnaissance subtask, it is determined that the interference subtask occupies the task dominance, and the interference subtask takes over the right to generate the flight action; If the priority score of the interference subtask is less than or equal to the priority score of the reconnaissance subtask, it is determined that the reconnaissance subtask occupies the task dominance, and the reconnaissance subtask takes over the right to generate the flight action.
3. The method according to claim 1, characterized in that Before splicing the preceding action of the current drone and the current observation vector and inputting them into the reinforcement learning neural network, the method further includes: Obtain battlefield environment data collected in real time by sensors arranged on the current UAV to generate the current observation vector, wherein the current observation vector includes the current UAV's own position coordinates, the current UAV's own flight speed, the enemy's position coordinates, the enemy's interference threat level, and the enemy's interference intensity; Obtaining the action of the current UAV at the previous time step to generate the preceding action; Normalizing the current observation vector and encoding the preceding action into a vector; The reinforcement learning neural network includes an Actor network and a Critic network. The preceding action of the current UAV and the current observation vector are spliced and input into the reinforcement learning neural network. The comprehensive action of the current UAV is determined based on the preceding action and the current observation vector, including: For the Actor network: The input layer is used to input the current observation vector and the previous action and perform splicing; The network layer captures the temporal dependencies between actions based on the LSTM module, receives the input results from the input layer and calculates the hidden state value; The output layer is used to output the flight action, reconnaissance action and interference action of the current UAV generated by the fully connected layer; Regarding the Critic network: The input layer is used to input the global state; wherein the global state includes the concatenation of the current observation vector of all drones and the previous action; The network layer is used to calculate the weight matrix; The output layer is used to output the global value generated by the fully connected layer, as well as the randomly initialized Actor network parameters and Critic network parameters.
4. The method according to claim 3, characterized in that include: When it is determined that the reconnaissance subtask occupies the task dominance, the reconnaissance subtask performs the following steps: Input the vector obtained by concatenating the current observation vector and the preceding action; Extracting time series features through the LSTM module of the network layer of the Actor network; Outputting the flight action dominated by the reconnaissance subtask and the reconnaissance action generated by the Actor network; The interference subtask performs the following steps: Input the vector obtained by concatenating the current observation vector and the preceding flight action; Generating the interference action through a fully connected layer of the Actor network; Output the interference action that multiplexes the flying action.
5. The method according to claim 4, characterized in that The determining of the comprehensive action of the current UAV according to the preceding action and the current observation vector includes: Outputting hidden state values through the LSTM module of the network layer of the Actor network; If the reconnaissance subtask dominated the task in the previous time step and the reconnaissance subtask currently dominates the task, the fully connected layer reuses the flight action generated by the reconnaissance subtask in the previous time step and generates the reconnaissance action of the current reconnaissance subtask; If the interference subtask dominated the task in the previous time step, and the reconnaissance subtask currently dominates the task, the fully connected layer reuses the flight action generated by the interference subtask in the previous time step and generates the reconnaissance action of the current reconnaissance subtask; If the reconnaissance subtask dominated the task in the previous time step, and the interference subtask currently dominates the task, the fully connected layer reuses the flight action generated by the reconnaissance subtask in the previous time step and generates the interference action of the current interference subtask; If the interfering subtask dominated the task in the previous time step and the interfering subtask currently dominates the task, the fully connected layer reuses the flight action generated by the interfering subtask in the previous time step and generates the interfering action of the current interfering subtask; Finally, the comprehensive action of the current drone is output.
6. The method according to any one of claims 1 to 5, characterized in that The detecting of flight path conflicts and interference frequency band overlap conflicts among multiple drones in the drone cluster, dynamically adjusting the flight actions and interference actions in the comprehensive action, and generating a global action corresponding to the drone cluster based on the dynamically adjusted flight actions and interference actions in the comprehensive action and the reconnaissance action, includes: For all drones in the drone cluster, Input the set of flight directions and flight speeds of all the drones; Input a set of interference threat levels and interference intensities of all the UAVs; Calculating the Euclidean distance between each UAV, and marking the flight path conflict if the Euclidean distance between the UAVs does not meet a preset safety distance threshold; Counting the number of times the interference threat degree and the interference intensity are the same, and if the number of times is greater than 1, marking it as an interference frequency band overlap conflict; Outputting a conflict flag set, wherein the conflict flag set includes a conflict type and an ID of the involved drones; Calculating an attention weight matrix and redistributing actions according to the attention weights, thereby dynamically adjusting the flying actions and the interference actions in the integrated action; For the flight path conflict, adjusting the flight direction; For the overlapping conflict of the interference frequency bands, allocating a new interference threat degree and interference intensity; Output the coordinated global action.
7. The method according to any one of claims 1 to 5, characterized in that After generating a global action corresponding to the UAV cluster based on the flight action and the interference action in the integrated action after dynamic adjustment, and the reconnaissance action, the method further includes: During the training of the reinforcement learning neural network, the experience pool is expanded to store action-dependent data, and the reinforcement learning neural network is optimized through collaborative rewards and conflict penalties; The extended experience pool stores action-dependent data and optimizes the reinforcement learning neural network through collaborative rewards and conflict penalties, including: The extended experience pool stores action-dependent data in five-tuple form, which includes the global state, the action executed in the current time step, the original reward value fed back by the environment, the global state of the next time step, and the preceding action. Performing a preprocessing operation on the quintuple data and encapsulating it into a standardized format, wherein the preprocessing operation includes denoising and removing outliers; Add conflict flags and coordination flags, allocate storage space, and store extended samples in time step order; Input the collaborative flag value, that is, count the number of reuses of the flight action; Enter the conflict flag value to count the total number of all conflicts; Calculating a collaborative reward increment based on the number of reuses of the flight action; using the collaborative reward increment to update the reward value; Calculate the penalty term based on the total number of all conflicts and the penalty coefficient; Based on the collaborative reward increment, the reward value, and the penalty term, a final reward value is updated and output to optimize the reinforcement learning neural network.
8. An unmanned aerial vehicle electronic countermeasure device, characterized in that: include: a decomposition module for decomposing the electronic countermeasure task into a reconnaissance subtask and a jamming subtask, and training the corresponding subtasks in a reinforcement learning neural network; wherein the reconnaissance subtask is used to output a flight action and a reconnaissance action, and the jamming subtask is used to output a flight action and a jamming action; a calculation module, configured to dynamically calculate a task priority and determine a task dominance of the reconnaissance subtask or the interference subtask according to the task priority; a multiplexing module configured to prioritize, based on the task dominance, the execution of the dominance subtask by the reinforcement learning neural network, wherein the dominance subtask generates the flight maneuvers, reconnaissance maneuvers, and jamming maneuvers of the current UAV; and to reuse the flight maneuvers generated by the dominance subtask when executing the non-dominance subtask; a determination module, configured to concatenate the preceding action of the current drone and the current observation vector and input them into the reinforcement learning neural network, and determine the comprehensive action of the current drone based on the preceding action and the current observation vector; A detection module, configured to detect flight path conflicts and interference frequency band overlap conflicts among multiple drones in a drone cluster, and dynamically adjust the flight actions and interference actions in the integrated action; A generation module is used to generate a global action corresponding to the drone cluster based on the flight action and interference action in the comprehensive action after dynamic adjustment, as well as the reconnaissance action. The global action is a collection of the adjusted comprehensive actions of all drones in the drone cluster.
9. A drone, characterized in that: The system comprises a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the drone electronic countermeasure method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the unmanned aerial vehicle electronic countermeasure method according to any one of claims 1 to 7 is implemented.