A task cognitive driving low-altitude electromagnetic defense resource cooperative dynamic management method and system

By employing a task-aware-driven collaborative management method for low-altitude electromagnetic defense resources, and utilizing graph autoencoders and hierarchical actor-evaluator mechanisms to optimize resource scheduling, the threat of malicious sensing devices in low-altitude intelligent networks is addressed. This approach achieves efficient interference and resource management, while enhancing network security and privacy protection.

CN120050683BActive Publication Date: 2025-11-25NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510268799.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-11-25
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Existing technologies lack effective resource allocation methods when facing threats from malicious sensing devices in low-altitude intelligent networks, making it impossible to effectively interfere with these devices. This results in threats to network security and privacy, and existing optimization methods are not suitable for highly dynamic adversarial environments involving malicious sensing.

Method used

A task-cognitive-driven collaborative dynamic management method for low-altitude electromagnetic defense resources is adopted. This method uses a graph autoencoder to cluster unmanned aerial vehicle (UAV) malicious sensing devices, combines a hierarchical actor-evaluator approach to optimize resource scheduling, and constructs a semi-Markov decision process to achieve collaborative control and interference decision-making for multi-dimensional resources.

Benefits of technology

It significantly improves the overall performance gains and interference effectiveness of the system, effectively protects the network from malicious sensing devices, enhances the intelligence level of countermeasures devices, and flexibly responds to highly changing malicious sensing electromagnetic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050683B_ABST
    Figure CN120050683B_ABST
Patent Text Reader

Abstract

The application discloses a task cognition driven low-altitude electromagnetic defense resource cooperative dynamic management method, comprising the following steps: for the electromagnetic defense system of malicious sensing unmanned aerial vehicle, using the sensing equipment and jammer of the own side to sense the malicious sensing equipment carried by the unmanned aerial vehicle and implement task cognition based interference countermeasure; constructing a graph model based on the spatial position information of the malicious sensing equipment carried by the unmanned aerial vehicle, and using a graph autoencoder to realize clustering of dynamic multi-countermeasure targets; constructing the resource scheduling decision process of the sensing equipment and jammer of the own side into a semi-Markov decision process; combining the method of hierarchical actor-critic to optimize the joint detection and interference of the sensing equipment and jammer of the own side on the malicious sensing equipment carried by the unmanned aerial vehicle, and the management and decision of multi-dimensional resources such as electromagnetic spectrum, and output the optimal decision. The application can effectively extract high-dimensional features from dynamic countermeasure targets and realize task cognition based clustering, and significantly improve the overall performance benefit and interference efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless sensing and electronic countermeasure, and particularly relates to a task cognition driven low-altitude electromagnetic defense resource cooperative dynamic management method and system. BACKGROUND

[0002] With the rapid development of 5G communication and the Internet of Things, mobile flexible and easy-to-deploy unmanned aerial vehicles will play the role of air base stations in future wireless networks and edge computing systems, providing users with extensive coverage and additional computing power. The continuous development of unmanned aerial vehicles has also given birth to a new economic form with great potential - low-altitude economy. As the low-altitude economy continues to develop, the number of various low-altitude aircraft is also rapidly increasing, which will make the electromagnetic environment in low-altitude increasingly complex. In the future low-altitude intelligent networking, the fusion of sensing, communication and other technologies with mobile networks enables spectrum situation awareness to enter a new stage of development. In the future low-altitude intelligent networking, devices equipped with wireless networks can use the transmission and reflection of radio waves to sense the environment and obtain information such as distance, speed and angle of the sensing target from the wireless signal, laying the foundation for detection, tracking and other sensing capabilities. However, as the intelligent networking technology continues to advance, the privacy and security of devices in the network are increasingly significant, and devices equipped with radar-like sensing capabilities (including malicious sensing devices) can cause serious interference and security risks to devices in the network. Malicious sensing devices (such as unmanned aerial vehicle-mounted malicious sensing devices) can locate, track, and illegally obtain target activity information, determine the nature of the target terrain, and pose a threat to mobile network security.

[0003] Currently, in the field of wireless sensing and countermeasures, the resource allocation technology for system devices, including spectrum allocation, power allocation, etc., mostly uses iterative algorithms such as alternating optimization or successive convex approximation, and no work has focused on scenarios for malicious sensing electromagnetic defense. As low-altitude intelligent networking continues to develop, malicious sensing devices will pose a threat to network security. The optimization methods of the above research are no longer applicable to the highly dynamic malicious sensing countermeasure environment. To defend against the security risks brought about by future malicious electromagnetic sensing and protect the performance of our own network, the principles and ideas of artificial intelligence should be applied to electromagnetic defense against malicious sensing, and intelligent methods should be used to recognize highly dynamic countermeasure environments, improve the intelligence level of countermeasure devices, and intelligently optimize the management and decision-making of multi-dimensional resources such as detection, interference and electromagnetic spectrum, to flexibly respond to highly variable malicious sensing electromagnetic environments. To effectively interfere with malicious sensing devices and destroy their sensing capabilities, protect our own targets from malicious sensing, and protect the privacy and security of mobile networks, it is urgent to research wireless sensing and countermeasure methods for malicious sensing devices and explore multi-dimensional resource management methods for malicious sensing electromagnetic defense. SUMMARY

[0004] The application aims to provide a task cognition driven low-altitude electromagnetic defense resource cooperative dynamic management method and system, which can effectively extract high-dimensional features from dynamic confrontation targets and realize task cognition based clustering, thereby significantly improving the overall performance benefits and interference efficiency of the system.

[0005] To achieve the above technical purposes, the technical scheme adopted by the application is as follows:

[0006] In a first aspect, the application discloses a task cognition driven low-altitude electromagnetic defense resource cooperative dynamic management method, which comprises the following steps:

[0007] Step 1: For the electromagnetic defense system of the maliciously aware unmanned aerial vehicle, the own awareness device and the jammer are used to sense the maliciously aware device carried by the unmanned aerial vehicle and implement task cognition based jamming confrontation, thereby generating a maliciously aware electromagnetic defense scene;

[0008] Step 2: A graph model is constructed based on the spatial position information of the maliciously aware device carried by the unmanned aerial vehicle, a graph autoencoder is used to realize clustering of dynamic multiple confrontation targets, and multiple jamming tasks are formed;

[0009] Step 3: The resource scheduling decision process of the own awareness device and the jammer is constructed as a semi-Markov decision process;

[0010] Step 4: The joint detection and jamming of the maliciously aware device carried by the unmanned aerial vehicle by the own awareness device and the jammer and the control and decision of the multi-dimensional resources such as electromagnetic spectrum are optimized by combining the method of layered actor-critic; the iteration is repeated until the task cognition driven multi-dimensional resource cooperative control algorithm converges, and the action selection decision at this time is taken as the optimal decision.

[0011] Further, in step 1, the maliciously aware electromagnetic defense scene comprises a set of maliciously aware devices carried by unmanned aerial vehicles D={1,…,d,…,D}, a set of jammers W={1,…,w,…,W}, a set of own awareness devices S={1,…,s,…,S}, and a target to be protected M; with a spatial rectangular coordinate system as a reference, the position of the maliciously aware device carried by the unmanned aerial vehicle is represented as h d ={X d ,Y d ,Z d}, the position of the jammer is h w ={X w ,Y w ,0}, the position of the own awareness device is h s ={X s ,Y s ,0}, and the position of the target to be protected is h t ={X t ,Yt ,0};no drone-borne malicious awareness device moving speed vector V d , direction vector O d ;

[0012] There are N frequency domain channels in the frequency domain space of the malicious awareness electromagnetic defense scene; at time t, the drone-borne malicious awareness device transmits an awareness signal to detect the own target, and uses the frequency domain channel f d,t ; the own awareness device intercepts the frequency domain channel occupation of each frequency using device including the drone-borne malicious awareness device, jammer and own awareness device at time t; the jammer makes a decision according to the frequency domain channel occupation, and transmits an interference signal at the next time t+1 to interfere with the drone-borne malicious awareness device, when the interference frequency point f w,t decided by the jammer at time t coincides with the frequency point f d,t+1 occupied by the drone-borne malicious awareness device at the next time, it is considered that the interference is effective.

[0013] Step 2 further comprises:

[0014] In the malicious awareness electromagnetic defense scene, the graph convolutional neural network is used to extract the features of non-Euclidean space data, and the task cognitive driven method of graph autoencoder is used to cluster the drone-borne malicious awareness device based on the obtained parameters of the drone-borne malicious awareness device;

[0015] An adversarial target graph model is constructed according to the parameters of the awareness signal of the drone-borne malicious awareness device;

[0016] An adjacency matrix A is constructed based on whether the moving directions of the drone-borne malicious awareness devices are the same:

[0017]

[0018] If the drone-borne malicious awareness device d and the drone-borne malicious awareness device d' have the same direction, the elements a d,d' =1, a d',d =1 of the adjacency matrix A; if not, a d,d' =0, a d',d =0, and D is the number of drone-borne malicious awareness devices;

[0019] A feature matrix X is constructed:

[0020] X=[V d ,h d ,O d ]

[0021] Wherein, V d is the speed of the drone-borne malicious awareness device, h d is the position, and O d is the direction.

[0022] A high-dimensional feature vector Z of the UAV-borne malicious perception device perception signal parameter is extracted by using a graph convolutional neural network:

[0023] Z = GCN(X, A)

[0024]

[0025] wherein is a symmetric normalized adjacency matrix; W0 and W1 are parameters to be learned; an inner product is used as a decoder to reconstruct the original graph:

[0026]

[0027] wherein, is a reconstructed adjacency matrix;

[0028] In the training process of GAE, cross-entropy is used as a loss function:

[0029]

[0030] wherein y represents the value of an element in the adjacency matrix A, is a reconstructed adjacency matrix the value of the corresponding element in the matrix, and N is the number of agents;

[0031] Based on the generated high-dimensional feature vector Z, a k-means clustering algorithm is used to generate a clustering label vector G of the UAV-borne malicious perception device, as category information for guiding the jammer to interfere with the countermeasures.

[0032] Step 3 further comprises the following steps:

[0033] Different countermeasure clusters targeted by the jammer are determined, the transmission power selection is selected as a coarse time granularity, the spectrum decision is selected as a fine time granularity, a multi-dimensional resource collaborative management hierarchical decision framework based on two-layer time granularity is constructed, and a top-down three-layer decision is established: a node clustering layer based on the target, a transmission power selection layer, and a frequency device spectrum decision layer;

[0034] In the node clustering layer based on the target, the obtained perception information of the UAV-borne malicious perception device is used to obtain potential features by GCN, and the UAV-borne malicious perception device is clustered based on the target features; then a reinforcement learning method is used to decide the association of the countermeasure target and the device, and the association result is used as the macro action information of the top layer is expressed as:

[0035]

[0036] wherein, is the association action of the jammer w;

[0037] In the transmit power selection layer, observations of each device within a micro-time slot t. Including the number N of malicious detection devices on drones. x,t The relative distance d of the detected malicious sensing device on the drone x,d,t , Upper time slot decision action information And with top-level macro action information The observation information for the next layer of decision-making conditions is constructed as follows:

[0038]

[0039] Where x is a jammer When x is a friendly sensing device. The beam pointing selection action of each device is represented as follows:

[0040] In the spectrum decision-making layer of frequency-using devices, the observations of each frequency-using device x Including its own position h x Received signal strength at each frequency It also inherits macro actions from the feature-based target clustering and device association layer. Macro actions with transmit power selection layer Represented as:

[0041]

[0042] Among them, the frequency decision-making actions of each device are represented as follows:

[0043] In the spectrum decision-making layer of frequency-using equipment, each of the user's sensing devices s will receive a reward. Defined as the effectiveness F of our own sensing equipment in detecting whether a malicious sensing device on a UAV is successfully detected. s,t , is represented as:

[0044]

[0045] Assuming one time slot equals Δ microtime slots, in the transmit power selection layer, each jammer w receives a reward. Each friendly sensing device receives a reward. In the transmit power selection layer, the policy network for each frequency-using device x is defined. and value network In time slot Δ, the transmit power selection action of each device is represented as follows:

[0046]

[0047] Based on the reward obtained from device x and the value network output V(), the first time difference error is calculated as:

[0048]

[0049] where γ is a decay factor;

[0050] The value network updates the value network parameters with the goal of minimizing the first time difference error, and the loss function of the value network is defined as:

[0051]

[0052] where is the expected mean;

[0053] The update formula of the value network parameters is expressed as:

[0054]

[0055] where α is the learning rate;

[0056] The first time difference error calculated is used to update the policy network parameters along the gradient direction of the loss function J(θ), and the loss function of the policy network is defined as:

[0057]

[0058] The update formula of the policy network parameters at each time slot is expressed as:

[0059]

[0060] In the spectrum decision layer of the frequency using device, the policy network of each frequency using device x is defined as: and the value network In the micro time slot t, the frequency selection action of each device is expressed as:

[0061]

[0062] According to the obtained return of device x and the value network output, the second time difference error is calculated as:

[0063]

[0064] The loss function of the value network is defined as:

[0065]

[0066] The update formula of the value network parameters is expressed as:

[0067]

[0068] The policy network loss function is defined as:

[0069]

[0070] The update formula of the policy network parameters at each time slot is represented as:

[0071]

[0072] Step 4 further comprises:

[0073] S41, initializing the policy network and the value network of each agent;

[0074] S42, at the beginning of the micro time slot t, in the feature-based target clustering layer, based on the obtained perception information of the unmanned airborne malicious perception device, obtaining the potential features through the GCN, clustering the unmanned airborne malicious perception device based on the target features, and taking the clustering result as the macro action information of the top layer The macro action information of the upper layer is transmitted to each device x, and the transmission power selection in each time slot is based on the policy network is obtained, and the transmission power selection action of each device is generated The spectrum scheduling decision is based on the policy network is obtained, and the spectrum decision action of each frequency device is generated

[0075] At the end of the micro time slot, the frequency device x selects the action according to the obtained reward The second time difference error is calculated for updating the policy network and the value network .

[0076] After Δ micro time slots, the top layer macro action ends, and each device x selects the action according to the obtained reward The first time difference error is calculated for updating the policy network and the value network , and generating a new top layer macro action according to the new perception information of the new unmanned airborne malicious perception device

[0077] S43, repeating step S42 until the multi-dimensional resource collaborative management algorithm based on task cognition driving converges, and taking the action selection decision at this time as the optimal decision.

[0078] In a second aspect, the present application discloses a task cognition driven low-altitude electromagnetic defense resource collaborative dynamic management system, which comprises:

[0079] A scene construction module is used for the electromagnetic defense system of the malicious awareness unmanned aerial vehicle, and awareness of the malicious awareness equipment carried by the unmanned aerial vehicle is implemented by using the own awareness equipment and jammer, and task-aware jamming confrontation is implemented to generate a malicious awareness electromagnetic defense scene.

[0080] A clustering module is used for constructing a graph model based on the spatial position information of the malicious awareness equipment carried by the unmanned aerial vehicle, and clustering of dynamic multi-adversary targets is implemented by using a graph autoencoder to form a plurality of jamming tasks.

[0081] A resource scheduling module is used for constructing a semi-Markov decision process for the resource scheduling decision process of the own awareness equipment and jammer.

[0082] A decision module combines a hierarchical actor-critic method to optimize the joint detection and jamming of the malicious awareness equipment carried by the unmanned aerial vehicle by the own awareness equipment and jammer, and the management and decision of multi-dimensional resources such as electromagnetic spectrum; and the algorithm is repeatedly iterated until the multi-dimensional resource collaborative management algorithm based on task-aware driving converges, and the action selection decision at this time is taken as the optimal decision.

[0083] Compared with the prior art, the beneficial effects of the present application are as follows:

[0084] The task-aware low-altitude electromagnetic defense resource collaborative dynamic management method and system can effectively extract high-dimensional features from dynamic adversary targets and implement task-aware clustering, and significantly improve the overall performance benefits and jamming effectiveness of the system. BRIEF DESCRIPTION OF DRAWINGS

[0085] Figure 1 The task-aware multi-dimensional resource collaborative management method flowchart of the present application.

[0086] Figure 2 The task-aware multi-dimensional resource collaborative management system model diagram of the present application.

[0087] Figure 3 The task-aware clustering method diagram of the present application.

[0088] Figure 4 The hierarchical decision framework diagram of the task-aware multi-dimensional resource collaborative management method of the present application. DETAILED DESCRIPTION

[0089] The embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0090] A task-aware low-altitude electromagnetic defense resource collaborative dynamic management method and system, characterized in that the air-ground collaborative communication hierarchical decision method comprises the following steps:

[0091] Step 1: For the electromagnetic defense system against maliciously aware UAVs, to protect own targets, own awareness devices and jammers are used to sense and implement task-aware interference countermeasures against UAV-borne malicious awareness devices;

[0092] Step 2: Based on the spatial position and other information of UAV-borne malicious awareness devices, a graph model is constructed, and a graph autoencoder is used to realize clustering of dynamic multi-adversary targets, i.e., forming multiple interference tasks;

[0093] Step 3: The resource scheduling decision process of own awareness devices and jammers is constructed as a semi-Markov decision process, and the method of hierarchical actor-critic is used to optimize the joint detection and interference of own awareness devices and jammers against UAV-borne malicious awareness devices, as well as the management and decision of multi-dimensional resources such as electromagnetic spectrum; repeat iteration until the multi-dimensional resource coordination management algorithm based on task-aware driving converges, at which time the action selection decision is the optimal decision.

[0094] The embodiment proposes a task-aware driven low-altitude electromagnetic defense resource coordination dynamic management method and system, the main steps, system model and method architecture are as shown in Figures 1 to 4 , and specifically include the following:

[0095] Set the malicious awareness electromagnetic defense as the background, as shown in Figure 2 , the scene includes a set of UAV-borne malicious awareness devices D = {1,..., d,..., D}, a set of jammers W = {1,..., w,..., W}, a set of own awareness devices S = {1,..., s,..., S}, and a set of own targets to be protected M. Taking the spatial rectangular coordinate system as the reference, the position of the UAV-borne malicious awareness device is h d = {X d ,Y d ,Z d}, the position of the jammer is h w = {X w ,Y w ,0}, the position of the own awareness device is h s = {X s ,Y s ,0}, and the position of the target to be protected is h t = {X t ,Y t ,0}. The speed vector of the UAV-borne malicious awareness device is V d , and the direction vector is O d .

[0096] The unmanned aerial malicious sensing device senses and tracks the own side target to be protected. In order to protect the target, the own side uses the jammer to implement interference on the unmanned aerial malicious sensing device, so that the unmanned aerial malicious sensing device cannot successfully detect the target to be protected. The jammer beam can directly point to and locate the unmanned aerial malicious sensing device. The own side sensing device senses the unmanned aerial malicious sensing device. The sensing here refers to active detection similar to radar equipment, and it is assumed that the sensing of the own side sensing device on the unmanned aerial malicious sensing device can obtain accurate position information in the stage of strengthening the target. In order to ensure that the jammer and the own side sensing device work normally, the working frequency channel conflict of the decision should be avoided.

[0097] It is assumed that there are N frequency domain channels in the frequency domain space in the wireless sensing and confrontation scene. At time t, the unmanned aerial malicious sensing device transmits a sensing signal to detect the own side target, and uses a frequency domain channel f d,t . The own side sensing device can detect the frequency domain channel occupation of each frequency device (unmanned aerial malicious sensing device, jammer, own side sensing device) at time t. It is assumed that the own side sensing device has certain prior knowledge, can accurately distinguish the unmanned aerial malicious sensing device and the own side sensing device through sorting and identification of the wireless sensing signal, and can further analyze the working mode M d,t and other information of the unmanned aerial malicious sensing device. The jammer makes a decision according to the occupation of each frequency domain channel, and transmits an interference signal (the frequency point f d,t+1 of the unmanned aerial malicious sensing device has changed) at the next time t+1, to interfere with the unmanned aerial malicious sensing device. At least when the interference frequency point f w,t of the jammer at time t coincides with the frequency point f d,t+1 occupied by the unmanned aerial malicious sensing device at the next time, that is, f w,t =f d,t+1 , it is considered that the interference is effective.

[0098] The conversion rule of the working mode of the unmanned aerial malicious sensing device is as follows. In the search state, if the unmanned aerial malicious sensing device detects the target for 3 times in 4 detections, it enters the tracking state; otherwise, it remains in the search state. When the unmanned aerial malicious sensing device works in the tracking state, if the target is detected for 2 times in 3 detections, it enters the locking state from the tracking state; after entering the locking state, the locking state is maintained for n time slots, at this time the jammer cannot interfere with it, then it enters the termination state, it is considered that this round of simulation is ended, the working mode returns to the search state, and a new round of simulation starts; if the target is not detected for 3 times, it returns to the search state; otherwise, it remains in the tracking state.

[0099] The sensing of the own side sensing device on the unmanned aerial malicious sensing device can be calculated by the detection probability. The detection probability P s,tis modeled as:

[0100]

[0101] where P fa is the probability of false alarm, erfc(·) is the complementary error function, i.e.

[0102]

[0103] Similarly, the detection probability P d,t of the UAV-borne malicious sensing device d at time slot t is modeled as:

[0104]

[0105] First, the detection probability of the friendly sensing device s at time slot t to the UAV-borne malicious sensing device is calculated by the signal-to-noise ratio, and then a random number υ∈U(0,1) uniformly distributed in [0,1] is generated, and the calculated detection probability is compared with the random number: when the random number is less than the detection probability, i.e. υ s,t , the friendly sensing device can detect the UAV-borne malicious sensing device, i.e. the detection efficiency F s,t of the friendly sensing device s at time slot t is 1, indicating whether the friendly sensing device s at time slot t detects the UAV-borne malicious sensing device; when the random number is greater than the detection probability, the friendly sensing device cannot detect the UAV-borne malicious sensing device, i.e. F s,t = 0.

[0106] Whether the UAV-borne malicious sensing device detects the friendly target can be simulated by evaluating the detection probability P d,t . The detection probability of the UAV-borne malicious sensing device to the friendly target is calculated, and a random number υ'∈(0,1) uniformly distributed in [0,1] is generated to determine whether the target is detected. When the random number is less than the detection probability, i.e. υ' < P d,t , the UAV-borne malicious sensing device can sense the friendly target, i.e. the sensing efficiency F d,t of the UAV-borne malicious sensing device is 1; when the random number is greater than the detection probability, the UAV-borne malicious sensing device cannot sense the friendly target, i.e. the sensing efficiency F d,t of the UAV-borne malicious sensing device is 0. And when the UAV-borne malicious sensing device is interfered by the jamming device, it is considered that the target is not detected. And when the UAV-borne malicious sensing device is interfered by the jammer, i.e. the signal frequency point of the UAV-borne malicious sensing device coincides with the frequency point used by the jammer, it is considered that the friendly target is not detected.

[0107] If the UAV-borne malicious perception device working in search mode has detected the own target in the first two detections, the third detection is a critical moment, and the first confrontation occurs at this moment. In the third detection, if the target is not detected, the fourth detection is regarded as another critical moment, and the second confrontation occurs at this moment; otherwise, the UAV-borne malicious perception device will switch to the tracking mode. If the working mode of the UAV-borne malicious perception device remains in the search mode, or changes from the tracking mode to the search mode, it is considered that the confrontation is successful, that is, the jammer is successful in jamming. If the working mode of the UAV-borne malicious perception device changes from the search mode to the tracking mode, or changes from the tracking mode to the lock state, it is considered that the confrontation fails.

[0108] In the scene of malicious perception electromagnetic defense, first, the graph convolutional neural network is used to extract the features of non-Euclidean space data, and the task cognitive driven method of graph autoencoder is used to cluster the UAV-borne malicious perception device based on the obtained parameters of the UAV-borne malicious perception device. The adversarial target graph model is constructed according to the perception signal parameters of the UAV-borne malicious perception device. The adjacency matrix A is constructed based on whether the moving directions of the UAV-borne malicious perception devices are the same. If the UAV-borne malicious perception device d and the UAV-borne malicious perception device d' are in the same direction, the element a d,d' of the adjacency matrix A is a d',d =1;a d,d' =0, and if they are not in the same direction, a d',d =0, then the adjacency matrix A is represented as

[0109]

[0110] where D is the number of UAV-borne malicious perception devices.

[0111] In order to represent multiple adversarial targets from multiple characteristics of the UAV-borne malicious perception device such as speed, position, direction, etc., the feature matrix X is constructed as

[0112] X=[V d ,h d ,O d ],

[0113] where V d is the speed of the UAV-borne malicious perception device, h d is the position, and O d is the direction.

[0114] The graph convolutional neural network is used to extract the high-dimensional feature vector Z of the perception signal parameters of the UAV-borne malicious perception device,

[0115] Z=GCN(X,A),

[0116]

[0117] where is the symmetric normalized adjacency matrix. W0 and W1 are parameters to be learned. The inner product is used as the decoder to reconstruct the original graph:

[0118]

[0119] where, is the reconstructed adjacency matrix.

[0120] In the training process of GAE, cross-entropy is used as the loss function:

[0121]

[0122] where, y represents the value (0 or 1) of an element in the adjacency matrix A, is the reconstructed adjacency matrix corresponding element in the value (0~1), N is the number of agents.

[0123] Based on the generated high-dimensional feature vector Z, the k-means clustering algorithm is used to generate the clustering label vector G of the unmanned aerial malicious perception device, as the category information guiding the jammer to interfere with the confrontation.

[0124] In the detection and interference of the own perception device and the jammer once the fixed transmission power, due to the existence of failure probability, it is necessary to use different frequencies to implement multiple detection and interference on the unmanned aerial malicious perception device, so the transmission power selection is selected as the coarse time granularity, and the spectrum decision is selected as the fine time granularity to construct a multi-dimensional resource collaborative management hierarchical decision framework based on two-layer time granularity, and a top-down three-layer decision is established: target-based node clustering layer, transmission power selection layer, and frequency device spectrum decision layer.

[0125] In the target-based node clustering layer, based on the obtained perception information of the unmanned aerial malicious perception device, the potential features are obtained through GCN, and the unmanned aerial malicious perception device is clustered based on the target features; then the reinforcement learning method is used to decide the association of the target and the device, and the association result is used as the macro action information of the top layer is expressed as:

[0126]

[0127] where, is the associated action of the jammer w.

[0128] In the transmission power selection layer, the observation of each device in the micro time slot t includes the number of perceived unmanned aerial malicious perception devices N x,t , the relative distance d x,d,t, the upper time slot decision action information and the top layer macro action information as the next layer decision condition; the observation information of the transmit power selection layer is constructed as:

[0129]

[0130] wherein when x is a jammer, when x is a self-aware device, The device beam pointing selection action is represented as

[0131] In the spectrum decision layer of the frequency using device, the observation of each frequency using device x including its own position h x , the received signal strength of each frequency point and the macro action of the target clustering and device association layer based on features and the macro action of the transmit power selection layer is represented as:

[0132]

[0133] wherein the frequency using decision action of each device is represented as

[0134] In the spectrum decision layer of the frequency using device, the reward obtained by each self-aware device s is defined as the effectiveness F of the success of the self-aware device detecting the unmanned aerial malicious awareness device s,t is represented as:

[0135]

[0136] Assuming that one time slot is equal to Δ micro time slots, in the transmit power selection layer, the reward obtained by each jammer w the reward obtained by each self-aware device s In the transmit power selection layer, the policy network of each frequency using device x and the value network In time slot Δ, the transmit power selection action of each device is represented as:

[0137]

[0138] According to the obtained return of device x and the value network output V(), the first time difference error is calculated:

[0139]

[0140] wherein γ is the decay factor.

[0141] The value network updates the value network parameters with the first time difference error minimization as the goal, and the loss function of the value network is defined as:

[0142]

[0143] wherein is the expected mean value.

[0144] The update formula of the value network parameters is represented as:

[0145]

[0146] wherein α is the learning rate.

[0147] The first time difference error calculated is used to update the policy network parameters in the gradient direction of the loss function J(θ), and the loss function of the policy network is defined as:

[0148]

[0149] The update formula of the policy network parameters at each time slot is represented as:

[0150]

[0151] In the spectrum decision layer of the frequency using device, the policy network of each frequency using device x is defined as: and the value network is In the micro time slot t, the frequency selection action of each device is represented as:

[0152]

[0153] According to the obtained return of the device x and the value network output, the second time difference error is calculated:

[0154]

[0155] The loss function of the value network is defined as:

[0156]

[0157] The update formula of the value network parameters is represented as:

[0158]

[0159] The loss function of the policy network is defined as:

[0160]

[0161] The update formula of the policy network parameters at each time slot is represented as:

[0162]

[0163] Each agent adopts an actor-critic based method. First, the policy network and the value network of each agent are initialized. At the beginning of the micro time slot t, in the feature-based target clustering layer, based on the obtained perception information of the unmanned aerial malicious perception device, the potential features are obtained through the GCN, the unmanned aerial malicious perception device is clustered based on the target features, and the clustering result is used as the macro action information of the top layer The macro action information of the upper layer is transmitted to each device x, and the transmission power selection in each time slot is based on the policy network The action of transmission power selection of each device is obtained and generated The spectrum scheduling decision is based on the policy network The spectrum decision action of each frequency device is obtained and generated At the end of the micro time slot, the frequency device x obtains the reward TD error is calculated for the update of the policy network and the value network . After Δ micro time slots, the top layer macro action ends, and each device x obtains the reward TD error is calculated for the update of the policy network and the value network , and a new top layer macro action is generated according to the new perception information of the unmanned aerial malicious perception device A new round of decision making is started.

[0164] The above steps are repeated until the algorithm converges, and the action selection decision at this time is taken as the optimal decision.

[0165] Those skilled in the art will appreciate that embodiments of the application can be provided as methods, systems, or computer program products. Accordingly, the application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) having computer-usable program code embodied in the medium. The solutions in the embodiments of the application can be implemented in various computer languages, such as object-oriented programming languages Java and interpreted scripting language JavaScript.

[0166] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0167] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0168] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0169] While the preferred embodiments of the application have been described, additional variations and modifications can be employed by those skilled in the art. Therefore, the appended claims intend to cover all such modifications and variations as fall within the true spirit and scope of the application.

[0170] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A task-awareness-driven collaborative dynamic management method for low-altitude electromagnetic defense resources, characterized in that, The method includes the following steps: Step 1: For the electromagnetic defense system against malicious sensing drones, use friendly sensing equipment and jammers to sense the malicious sensing equipment on the drone and implement jamming countermeasures based on task cognition to generate a malicious sensing electromagnetic defense scenario. Step 2: Construct a graph model based on the spatial location information of the UAV-borne malicious detection device, and use a graph autoencoder to cluster dynamic multi-adversarial targets to form multiple interference tasks; Step 3: Construct the resource scheduling decision process for our own sensing equipment and jammers into a semi-Markov decision process; Step 4: Combine the hierarchical actor-evaluator method to optimize the joint detection and interference of our own sensing devices and jammers against UAV-borne malicious sensing devices and the control and decision-making of multi-dimensional resources such as electromagnetic spectrum; repeat the iteration until the multi-dimensional resource collaborative control algorithm based on task cognition converges, and take the action selection decision at this time as the optimal decision.

2. The task-awareness-driven collaborative dynamic management method for low-altitude electromagnetic defense resources according to claim 1, characterized in that, In step 1, the malicious sensing electromagnetic defense scenario includes a set of UAV-borne malicious sensing devices D = {1,...,d,...,D}, a set of jammers W = {1,...,w,...,W}, a set of friendly sensing devices S = {1,...,s,...,S}, and a friendly target M to be protected; with a spatial rectangular coordinate system as a reference, the position of the UAV-borne malicious sensing devices is represented as h. d ={X d ,Y d Z d The jammer's location is h. w ={X w ,Y w ,0}, the location of our sensing device is h s ={X s ,Y s The location of the target to be protected is h, 0}. t ={X t ,Y t The velocity vector of the UAV-borne malicious detection device is V. d The direction vector is O d ; In the frequency domain space of the malicious sensing electromagnetic defense scenario, there are N frequency domain channels; in time slot t, the UAV-borne malicious sensing device emits sensing signals to detect friendly targets, and the frequency domain channel it uses is f. d,t The friendly sensing equipment detects the frequency domain channel occupancy status of various frequency-using devices, including the UAV-borne malicious sensing equipment, the jammer, and the friendly sensing equipment, in time slot t. The jammer makes a decision based on the frequency domain channel occupancy status and transmits a jamming signal at the next time slot t+1 to interfere with the UAV-borne malicious sensing equipment. When the jammer determines the jamming frequency f in time slot t... w,t With the malicious sensing device on the drone occupying frequency point f in the next time slot d,t+1 When they overlap, the interference is considered effective.

3. The task-awareness-driven collaborative dynamic management method for low-altitude electromagnetic defense resources according to claim 1, characterized in that, Step 2 further includes: In the scenario of malicious electromagnetic defense, a graph convolutional neural network is used to extract features of non-Euclidean space data. Based on the acquired parameters of the UAV-borne malicious detection device, a task cognition-driven method using a graph autoencoder is adopted to cluster the UAV-borne malicious detection device. Construct an adversarial target graph model based on the signal parameters perceived by the malicious detection device on the UAV; Construct an adjacency matrix A based on whether the movement directions of the UAV-borne malicious detection devices are the same: If the direction of the UAV-borne malicious detection device d and the UAV-borne malicious detection device d' is the same, then the element a of the adjacency matrix A is... d,d' =1,a d',d =1; if they are not the same, then a d,d' =0,a d',d =0, where D is the number of malicious detection devices carried by the drone; Construct the feature matrix X: X=[V d ,h d ,O d ] Among them, V d h represents the speed at which the drone carries the malicious detection device. d For position, O d For direction; A graph convolutional neural network is used to extract the high-dimensional feature vector Z of the sensing signal parameters of the UAV-borne malicious detection device: Z = GCN(X,A) in It is a symmetric normalized adjacency matrix; W0 and W1 are parameters to be learned; the inner product is used as a decoder to reconstruct the original graph: in, This is the reconstructed adjacency matrix; During the training of GAE, cross-entropy is used as the loss function: Where y represents the value of one element in the adjacency matrix A. To reconstruct the adjacency matrix The value of the corresponding element in the table, where N is the number of agents; Based on the generated high-dimensional feature vector Z, the k-means clustering algorithm is used to generate cluster label vector G for UAV-borne malicious sensing devices, which serves as category information to guide jamming countermeasures.

4. The task-awareness-driven collaborative dynamic management method for low-altitude electromagnetic defense resources according to claim 1, characterized in that, Step 3 further includes the following steps: The jammer identifies different adversarial target clusters and constructs a multi-dimensional resource collaborative management and control hierarchical decision-making framework based on two time granularities: transmit power selection as coarse time granularity and spectrum decision as fine time granularity. It establishes a top-down three-layer decision-making layer: target-based node clustering layer, transmit power selection layer, and frequency equipment spectrum decision layer. In the target-based node clustering layer, latent features are obtained through GCN based on the perception information of the UAV-borne malicious detection devices. The UAV-borne malicious detection devices are then clustered based on the target features. Reinforcement learning methods are then used to determine the association between adversarial targets and devices, and the association results serve as the macro-action information at the top layer. Represented as: in, This refers to the associated actions of the jammer w; In the transmit power selection layer, observations of each device within a micro-time slot t. Including the number N of malicious detection devices on drones. x,t The relative distance d of the detected malicious sensing device on the drone x,d,t Information on decision-making actions in the upper time slot And with top-level macro action information The observation information for the next layer of decision-making conditions is constructed as follows: Where x is a jammer When x is a friendly sensing device. The beam pointing selection action of each device is represented as follows: In the spectrum decision-making layer of frequency-using devices, the observations of each frequency-using device x Including its own position h x Received signal strength at each frequency It also inherits macro actions from the feature-based target clustering and device association layer. Macro actions with transmit power selection layer Represented as: Among them, the frequency decision-making actions of each device are represented as follows: In the spectrum decision-making layer of frequency-using equipment, each of the user's sensing devices s will receive a reward. Defined as the effectiveness F of our own sensing equipment in detecting whether a malicious sensing device on a UAV is successfully detected. s,t , is represented as: Assuming one time slot equals Δ microtime slots, in the transmit power selection layer, each jammer w receives a reward. Each friendly sensing device receives a reward. In the transmit power selection layer, the policy network for each frequency-using device x is defined. and value network In time slot Δ, the transmit power selection action of each device is represented as follows: Based on the reward obtained from device x And the value network output V(), calculate the first time difference error: Where γ is the attenuation factor; The value network updates its parameters with the goal of minimizing the first-time difference error. The loss function of the value network is defined as: in It is the expected mean; The formula for updating the value network parameters is expressed as: Where α is the learning rate; Using the calculated first time difference error, the policy network parameters are updated along the gradient direction of the loss function J(θ). The loss function of the policy network is defined as: The update formula for the policy network parameters in each time slot is expressed as: In the spectrum decision layer of frequency-using devices, the policy network of each frequency-using device x is defined. and value network In a micro-time slot t, the frequency selection action of each device is represented as follows: Based on the reward obtained from device x And the value network output, calculate the second time difference error: The loss function of a value network is defined as: The formula for updating the value network parameters is expressed as: The policy network loss function is defined as: The update formula for the policy network parameters in each time slot is expressed as:

5. The task-awareness-driven collaborative dynamic management method for low-altitude electromagnetic defense resources according to claim 1, characterized in that, Step 4 further includes: S41, Initialize the policy network and value network of each agent; S42, at the start of micro-time slot t, in the feature-based target clustering layer, based on the acquired perception information of the UAV-borne malicious detection device, latent features are obtained through GCN. The UAV-borne malicious detection device is then clustered based on the target features, and the clustering result serves as the macro-action information at the top layer. upper-level macro action information Transmitted to each device x, the transmit power selection within each time slot is based on the policy network. Obtain and generate the transmit power selection action for each device. Spectrum scheduling decision-making based on strategy network Acquire and generate spectrum decision actions for each frequency-using device. At the end of the micro-timeslot, frequency-using device x is rewarded accordingly. Calculate the second time difference error For policy networks and value network Update; After Δ micro-slots, the top-level macro action ends, and each device x proceeds according to the rewards it received. Calculate the first time difference error For policy networks and value network The update generates new top-level macro actions based on the new information perceived by the drone-borne malware detection device. S43. Repeat step S42 until the multi-dimensional resource collaborative management and control algorithm based on task cognition converges, and take the action selection decision at this time as the optimal decision.

6. A task-awareness-driven collaborative dynamic management system for low-altitude electromagnetic defense resources, characterized in that, The system includes: The scenario construction module is used for electromagnetic defense systems against malicious sensing drones. It uses friendly sensing equipment and jammers to sense the malicious sensing equipment on the drone and implement jamming countermeasures based on task cognition to generate malicious sensing electromagnetic defense scenarios. The clustering module is used to construct a graph model based on the spatial location information of the UAV-borne malicious sensing device, and uses a graph autoencoder to cluster dynamic multi-adversarial targets to form multiple interference tasks. The resource scheduling module is used to construct the resource scheduling decision process of our own sensing equipment and jammers into a semi-Markov decision process; The decision-making module combines a hierarchical actor-evaluator approach to optimize the joint detection and interference of our own sensing devices and jammers against UAV-borne malicious sensing devices and the management and decision-making of multi-dimensional resources such as the electromagnetic spectrum; it repeats the iteration until the multi-dimensional resource collaborative management and control algorithm based on task cognition converges, and the action selection decision at this time is taken as the optimal decision.

Citation Information

Patent Citations

  • Unmanned aerial vehicle signal suppression method and device, electronic equipment and storage medium

    CN111669248A

  • Unmanned aerial vehicle auxiliary communication anti-interference method based on multi-agent reinforcement learning

    CN118921099A