Task cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method and system

Through the task cognitive-driven collaborative dynamic management method of low-altitude electromagnetic defense resources, high-dimensional features are extracted using graph autoencoder and graph convolutional neural network to realize the coordinated control of clustering and multi-dimensional resources of malicious perception devices, solving the resource allocation problem in wireless perception and confrontation scenarios, and significantly improving the system's performance and interference efficiency.

CN120050683AActive Publication Date: 2025-05-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510268799.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-27
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

In current wireless perception and confrontation scenarios, resource allocation technology is difficult to effectively deal with highly dynamic malicious perception electromagnetic environments, resulting in network security threats and performance guarantee problems.

Method used

The task cognitive-driven collaborative dynamic management method of low-altitude electromagnetic defense resources is adopted, and high-dimensional features are extracted through graph autoencoder and graph convolution neural network to achieve coordinated management of clustering and multi-dimensional resources of dynamic multi-adversarial targets.

Benefits of technology

It significantly improves the overall performance benefits and interference efficiency of the system, effectively deals with the perception and interference of malicious perception devices, and ensures the privacy and security of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050683A_ABST
    Figure CN120050683A_ABST
Patent Text Reader

Abstract

The invention discloses a task-cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method, which comprises the following steps of: for an electromagnetic defense system of a malicious perception unmanned aerial vehicle, perceiving malicious perception equipment on the unmanned aerial vehicle by using own perception equipment and a jammer and implementing task-cognition-based interference countermeasure; a graph model is constructed based on the spatial position information of the unmanned aerial vehicle-mounted malicious sensing equipment, and clustering of dynamic multi-confrontation targets is realized by using a graph auto-encoder; constructing a resource scheduling decision process of the own sensing device and the jammer into a semi-Markov decision process; a hierarchical action device-evaluator method is combined to optimize joint detection and interference of own sensing equipment and a jammer on unmanned aerial vehicle-mounted malicious sensing equipment and control and decision of multi-dimensional resources such as electromagnetic spectrum and the like, and an optimal decision is output. According to the method, high-dimensional features can be effectively extracted from the dynamic confrontation target, clustering based on task cognition is realized, and the overall performance benefit and interference efficiency of the system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of wireless sensing and electronic countermeasure, and particularly relates to a method and system for collaborative dynamic management of low-altitude electromagnetic defense resources driven by task cognition. Background Art

[0002] With the rapid development of the 5G communication and Internet of Things fields, drones that are mobile, flexible, and easy to deploy will play the role of aerial base stations in future wireless networks and edge computing systems, providing users with a wide coverage area and additional computing capabilities. The continuous development of drones has also given rise to a new economic form with great potential - the low-altitude economy. With the continuous development of the low-altitude economy, the number of various low-altitude aircraft has increased rapidly, which will inevitably make the low-altitude electromagnetic environment increasingly complex. In the future low-altitude intelligent Internet of Things, sensing technology is one of the main driving forces for its development. The integration of technologies such as sensing and communication with mobile networks has brought the spectrum situation sensing into a new stage of development. In the future low-altitude intelligent Internet of Things, devices equipped with wireless networks can use the transmission and reflection of radio waves to sense the environment, and obtain information such as the distance, speed, and angle of the sensing target from wireless signals, laying a foundation for detection, tracking, and other sensing capabilities. However, with the continuous progress of intelligent Internet of Things technology, the privacy and security issues of devices in the network have become increasingly prominent. Devices equipped with similar radar sensing capabilities (including malicious sensing devices) will cause serious interference and security risks to the devices in the network. Malicious sensing devices (such as unmanned aerial vehicle (UAV)-borne malicious sensing devices) can locate, track, illegally obtain target activity information, and determine the terrain nature of the target, posing a threat to the security of mobile networks.

[0003] Currently, in wireless sensing and countermeasure scenarios, for the resource allocation technology of system devices, including spectrum allocation, power allocation, etc., most adopt iterative algorithms such as alternating optimization or successive convex approximation, and no work focuses on the scenario of electromagnetic defense against malicious sensing. With the continuous development of the low-altitude intelligent Internet of Things, malicious sensing devices will pose a threat to the security of the network. The above-mentioned optimization methods are no longer applicable to the highly dynamic adversarial environment of malicious sensing. To defend against the security risks brought by future malicious electromagnetic sensing and ensure the performance of our own network, the principles and ideas of artificial intelligence should be considered and applied to electromagnetic defense against malicious sensing, use intelligent methods to recognize the highly dynamic adversarial environment, improve the intelligent level of adversarial devices, and intelligently optimize the management and decision-making of multi-dimensional resources such as detection, interference, and electromagnetic spectrum, and flexibly respond to the highly changing malicious sensing electromagnetic environment. To effectively interfere with malicious sensing devices and destroy their sensing capabilities, protect our own targets from being maliciously sensed, and ensure the privacy and security of mobile networks, it is urgent to study wireless sensing and countermeasure methods for malicious sensing devices and explore multi-dimensional resource management methods for electromagnetic defense against malicious sensing. Summary of the Invention

[0004] The purpose of the present invention is to provide a task-cognition-driven collaborative dynamic management method and system for low-altitude electromagnetic defense resources, which can effectively extract high-dimensional features from dynamic confrontation targets and achieve clustering based on task cognition, significantly improving the overall performance benefit and interference efficiency of the system.

[0005] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0006] In the first aspect, the present invention discloses a task-cognition-driven collaborative dynamic management method for low-altitude electromagnetic defense resources. The method includes the following steps:

[0007] Step 1: For the electromagnetic defense system against malicious sensing drones, use the own sensing devices and jammers to sense the malicious sensing devices on the drones and implement task-cognition-based interference countermeasures to generate a malicious sensing electromagnetic defense scenario.

[0008] Step 2: Construct a graph model based on the spatial position information of the malicious sensing devices on the drones, and use a graph autoencoder to achieve clustering of dynamic multi-confrontation targets, forming multiple interference tasks.

[0009] Step 3: Construct the resource scheduling decision-making process of the own sensing devices and jammers as a semi-Markov decision process.

[0010] Step 4: Combine the hierarchical actor-critic method to optimize the joint detection and interference of the own sensing devices and jammers on the malicious sensing devices on the drones and the management and decision-making of multi-dimensional resources such as the electromagnetic spectrum; repeat the iteration until the multi-dimensional resource collaborative management and control algorithm driven by task cognition converges, and use the action selection decision at this time as the optimal decision.

[0011] Further, in step 1, the malicious sensing electromagnetic defense scenario includes a set D = {1,..., d,..., D} of malicious sensing devices on the drones, a set W == {1,..., w,..., W} of jammers, a set S = {1,..., s,..., S} of the own sensing devices, and an own target M to be protected; with the spatial rectangular coordinate system as a reference, the position of the malicious sensing device on the drone is represented as h d ={X d , Y d , Z d}, the position of the jammer is h w ={X w , Y w , 0}, the position of the own sensing device is h s ={X s , Y s , 0}, and the position of the target to be protected is h t ={X t , Yt , 0}; The velocity vector of the malicious perception device on the unmanned aerial vehicle (UAV) is V d , and the direction vector is O d ;

[0012] There are N frequency domain channels in the frequency domain space of the malicious perception electromagnetic defense scenario; At time slot t, the malicious perception device on the UAV emits a perception signal to detect its own target, and the frequency domain channel it uses is f d,t , and the own perception device detects the frequency domain channel occupancy of each frequency-using device including the malicious perception device on the UAV, the jammer, and the own perception device at time slot t; The jammer makes a decision based on the frequency domain channel occupancy and emits a jamming signal at the next time slot t + 1 to jam the malicious perception device on the UAV. When the jamming frequency point f w,t decided by the jammer at time slot t coincides with the frequency point f d,t+1 occupied by the malicious perception device on the UAV in the next time slot, it is considered that the jamming is effective.

[0013] Step 2 further includes:

[0014] In the malicious perception electromagnetic defense scenario, use a graph convolutional neural network to extract the features of non-Euclidean space data, and cluster the malicious perception devices on the UAV using the task cognitive drive method of a graph autoencoder based on the obtained parameters of the malicious perception devices on the UAV;

[0015] Construct an adversarial target graph model according to the perception signal parameters of the malicious perception devices on the UAV;

[0016] Construct an adjacency matrix A based on whether the moving directions of the malicious perception devices on the UAV are the same:

[0017]

[0018] If the malicious perception device d and the malicious perception device d' on the UAV have the same direction, then the element a d,d' of the adjacency matrix A is 1, and a d',d is 1; If they are not the same, then a d,d' is 0, and a d',d is 0, where D is the number of malicious perception devices on the UAV;

[0019] Construct a feature matrix X:

[0020] X = [V d , h d , O d

[0021] where V d is the velocity of the malicious perception device on the UAV, h d is the position, and O d is the direction;​

[0022] Use a graph convolutional neural network to extract the high-dimensional feature vector Z of the perception signal parameters of the airborne malicious perception device:

[0023] Z = GCN(X, A)

[0024]

[0025] where is the symmetric normalized adjacency matrix; W 0 and W 1 are the parameters to be learned; use the inner product as the decoder to reconstruct the original graph:

[0026]

[0027] where, is the reconstructed adjacency matrix;

[0028] During the training process of GAE, use cross-entropy as the loss function:

[0029]

[0030] where y represents the value of one element in the adjacency matrix A, is the value of the corresponding element in the reconstructed adjacency matrix and N is the number of agents;

[0031] Based on the generated high-dimensional feature vector Z, use the k-means clustering algorithm to generate the clustering label vector G of the airborne malicious perception device, which is used as the category information to guide the jammer's interference countermeasure.

[0032] Step 3 further includes the following steps:

[0033] Determine the different countermeasure target clusters for the jammer, construct a hierarchical decision framework for multi-dimensional resource collaborative management and control based on two-layer time granularity by taking the transmit power selection as the coarse time granularity and the spectrum decision as the fine time granularity, and establish a three-layer decision from top to bottom: the node clustering layer based on the target, the transmit power selection layer, and the spectrum decision layer of the frequency-using equipment;

[0034] In the node clustering layer based on the target, based on the obtained perception information of the airborne malicious perception device, obtain the latent features through GCN, and cluster the airborne malicious perception device based on the target features; then use the reinforcement learning method to decide the association between the countermeasure target and the device, and the association result is used as the macro-action information at the top layer Expressed as:

[0035]

[0036] where, Is the associated action of jammer w;

[0037] In the transmit power selection layer, the observations of each device within micro-slot t Include the number N of malicious sensing devices carried by drones sensed x,t , the relative distance d of the malicious sensing devices carried by drones detected x,d,t , the decision action information of the previous time slot And use the macro-action information of the top layer As the decision condition for the next layer; The observation information of the transmit power selection layer is constructed as:

[0038]

[0039] Where when x is a jammer, When x is a friendly sensing device, The beam pointing selection action of each device is expressed as

[0040] In the spectrum decision layer of frequency-using devices, the observations of each frequency-using device x Include its own position h x , the received signal strength at each frequency point And inherit the macro-action of the target clustering and device association layer based on features And the macro-action of the transmit power selection layer Is expressed as:

[0041]

[0042] Among them, the frequency-using decision action of each device is expressed as

[0043] In the spectrum decision layer of frequency-using devices, the reward obtained by each friendly sensing device s Is defined as the effectiveness F of whether the friendly sensing device successfully detects the malicious sensing device carried by the drone s,t , is expressed as:

[0044]

[0045] Assume that one time slot is equal to Δ micro-slots. In the transmit power selection layer, each jammer w obtains a reward Each friendly sensing device s obtains a reward In the transmit power selection layer, define the policy network of each frequency-using device x And the value network In time slot Δ, the transmit power selection action of each device is expressed as:

[0046]

[0047] According to the obtained return of device x and the value network output V(), calculate the first temporal difference error:

[0048]

[0049] where γ is the attenuation factor;

[0050] The value network updates the value network parameters with the goal of minimizing the first temporal difference error, and the loss function of the value network is defined as:

[0051]

[0052] where is the expected mean;

[0053] The update formula for the value network parameters is expressed as:

[0054]

[0055] where α is the learning rate;

[0056] Using the calculated first temporal difference error, update the policy network parameters along the gradient direction of the loss function J(θ), and the loss function of the policy network is defined as:

[0057]

[0058] The update formula for the policy network parameters in each time slot is expressed as:

[0059]

[0060] In the spectrum decision layer of the frequency-using device, define the policy network and the value network of each frequency-using device x. In the micro time slot t, the frequency-using selection action of each device is expressed as:

[0061]

[0062] According to the obtained return of device x and the value network output, calculate the second temporal difference error:

[0063]

[0064] The loss function of the value network is defined as:

[0065]

[0066] The update formula for the value network parameters is expressed as:

[0067]

[0068] The loss function of the policy network is defined as:

[0069]

[0070] The update formula of the policy network parameters at each time slot is expressed as:

[0071]

[0072] Step 4 further includes:

[0073] S41, initialize the policy network and value network of each agent;

[0074] S42, at the beginning of the micro time slot t, in the feature-based target clustering layer, based on the acquired sensing information of the UAV-borne malicious sensing device, obtain potential features through GCN, and cluster the UAV-borne malicious sensing device based on the target features. The clustering result is used as the macro-action information at the top layer The macro-action information of the upper layer is transmitted to each device x, and the transmission power selection within each time slot is based on the policy network to obtain and generate the transmission power selection actions of each device The spectrum scheduling decision is based on the policy network to obtain and generate the spectrum decision actions of each frequency-using device

[0075] At the end of the micro time slot, the frequency-using device x calculates the second temporal difference error according to the obtained reward for updating the policy network and the value network and the value network update;

[0076] After Δ micro time slots, when the top-level macro-action ends, each device x calculates the first temporal difference error according to the obtained reward for updating the policy network and the value network and the value network and generate a new top-level macro-action according to the new sensing information of the UAV-borne malicious sensing device

[0077] S43, repeat step S42 until the multi-dimensional resource collaborative control algorithm driven by task cognition converges, and take the action selection decision at this time as the optimal decision.

[0078] In a second aspect, the present invention discloses a task-cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management system, and the system includes:

[0079] A scenario construction module for an electromagnetic defense system against malicious perception drones, which uses its own perception devices and jammers to perceive the malicious perception devices on the drones and implement interference countermeasures based on mission cognition, and generates a malicious perception electromagnetic defense scenario;

[0080] A clustering module for constructing a graph model based on the spatial position information of the malicious perception devices on the drones, and using a graph autoencoder to achieve clustering of dynamic multi-countermeasure targets to form multiple interference tasks;

[0081] A resource scheduling module for constructing the resource scheduling decision-making process of its own perception devices and jammers into a semi-Markov decision-making process;

[0082] A decision-making module that combines the hierarchical actor-critic method to optimize the joint detection and interference of its own perception devices and jammers on the malicious perception devices on the drones and the management and decision-making of multi-dimensional resources such as the electromagnetic spectrum; repeat the iteration until the multi-dimensional resource collaborative management algorithm driven by mission cognition converges, and use the action selection decision at this time as the optimal decision

[0083] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0084] The mission cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method and system of the present invention can effectively extract high-dimensional features from dynamic countermeasure targets and achieve clustering based on mission cognition, significantly improving the overall performance benefit and interference efficiency of the system. Brief Description of the Drawings

[0085] Figure 1 It is a flowchart of the mission cognition-driven multi-dimensional resource collaborative management method of the present invention.

[0086] Figure 2 It is a system model diagram of the mission cognition-driven multi-dimensional resource collaborative management of the present invention.

[0087] Figure 3 It is a clustering method diagram of the mission cognition-driven of the present invention.

[0088] Figure 4 It is a hierarchical decision-making framework diagram of the mission cognition-driven multi-dimensional resource collaborative management method of the present invention. Detailed Embodiment

[0089] The following further describes the embodiments of the present invention in detail with reference to the drawings.

[0090] A mission cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method and system, characterized in that the air-ground collaborative communication hierarchical decision-making method includes the following steps:

[0091] Step 1: For the electromagnetic defense system against malicious perception drones, to protect its own targets, use its own perception devices and jammers to perceive the malicious perception devices on the drones and implement interference countermeasures based on mission awareness;

[0092] Step 2: Based on information such as the spatial location of the malicious perception devices on the drones, construct a graph model, and use a graph autoencoder to achieve clustering of dynamic multi-countermeasure targets, that is, form multiple interference tasks;

[0093] Step 3: Construct the resource scheduling decision-making process of its own perception devices and jammers as a semi-Markov decision-making process, and combine the hierarchical actor-critic method to optimize the joint detection and interference of its own perception devices and jammers on the malicious perception devices on the drones and the management and decision-making of multi-dimensional resources such as the electromagnetic spectrum; Repeat the iteration until the multi-dimensional resource collaborative management algorithm driven by mission awareness converges, and at this time, the action selection decision is the optimal decision.

[0094] This embodiment proposes a method and system for collaborative dynamic management of low-altitude electromagnetic defense resources driven by mission awareness. Its main steps, system model, and method architecture are as Figures 1 to 4 shown, specifically including the following:

[0095] Set the malicious perception electromagnetic defense as the background. As Figure 2 shown, the scenario includes a set D = {1,..., d,..., D} of malicious perception devices on drones, a set W == {1,…, w,…, W} of jammers, a set S = {1,..., s,..., S} of its own perception devices, and a target M to be protected by itself. With the spatial rectangular coordinate system as a reference, the position of the malicious perception device on the drone is represented as h d ={X d , Y d , Z d}, the position of the jammer is h w ={X w , Y w , 0}, the position of its own perception device is h s ={X s , Y s , 0}, and the position of the target to be protected is h t ={X t , Y t , 0}. The velocity vector of the malicious perception device on the drone is V d , and the direction vector is O d .

[0096] The malicious perception device on the unmanned aerial vehicle (UAV) perceives and tracks the target to be protected on one's own side. To protect the target, one's own side uses a jammer to interfere with the malicious perception device on the UAV, making it unable to successfully detect the target to be protected. The jammer beam can directly point to and locate the malicious perception device on the UAV. The perception device on one's own side conducts wireless perception of the malicious perception device on the UAV. Here, the perception refers to active detection similar to a radar device, and it is assumed that when the perception device on one's own side perceives the malicious perception device on the UAV during the stage of strengthening the target, accurate position information can be obtained as feedback. To ensure the normal operation of the jammer and the perception device on one's own side, the working frequency channels of their decisions should be avoided from conflicting.

[0097] Assume that there are N frequency channels in the frequency domain space in the wireless perception and countermeasure scenario. At time slot t, the malicious perception device on the UAV emits a perception signal to detect the target on one's own side, and the frequency channel used is f d,t . The perception device on one's own side can detect the occupancy of the frequency channels of each frequency-using device (the malicious perception device on the UAV, the jammer, and the perception device on one's own side) at time slot t. Assume that the perception device on one's own side has certain prior knowledge and can accurately distinguish the malicious perception device on the UAV from the perception device on one's own side through the sorting and recognition of wireless perception signals, and can further analyze the working mode M d,t and other information of the malicious perception device on the UAV. The jammer makes a decision based on the occupancy of each frequency channel and emits a jamming signal at the next time t + 1 (the frequency channel f d,t+1 of the malicious perception device on the UAV has changed), interfering with the malicious perception device on the UAV. At least when the jamming frequency point f w,t decided by the jammer at time slot t coincides with the frequency channel occupied by the malicious perception device on the UAV in the next time slot f d,t+1 , that is, f w,t = f d,t+1 , it is considered that the jamming is effective.

[0098] The conversion rule of the working mode of the malicious perception device on the UAV is as follows. In the search state, if the malicious perception device on the UAV detects the target 3 times out of 4 detections, it enters the tracking state; otherwise, it remains in the search state. When the malicious perception device on the UAV is working in the tracking state, if it detects the target 2 times out of 3 detections, it enters the locking state from the tracking state; after entering the locking state, the locking state lasts for n time slots, at this time the jammer cannot interfere with it, and then it enters the termination state, considering this round of simulation ended, and the working mode returns to the search state to start a new round of simulation; if it fails to detect the target in all 3 detections, it returns to the search state; otherwise, it remains in the tracking state.

[0099] The perception of the malicious perception device on the UAV by the perception device on one's own side can be calculated through the detection probability. The detection probability P of the perception device s on one's own side at time slot t s,tIs modeled as:

[0100]

[0101] Where P fa Is the false alarm probability, erfc(·) is the complementary error function, that is

[0102]

[0103] Similarly, the detection probability P d,t Of the airborne malicious sensing device d in time slot t is modeled as:

[0104]

[0105] First, calculate the detection probability of the own sensing device s for the airborne malicious sensing device in time slot t through the signal-to-noise ratio, and then generate a random number υ ∈ U(0,1) uniformly distributed in [0,1], and compare the calculated detection probability with the random number: when the random number is less than the detection probability, that is, υ < P s,t , the own sensing device can detect the airborne malicious sensing device, that is, the detection efficiency F s,t Of the own sensing device s in time slot t is 1, which is used to indicate whether the own sensing device s detects the airborne malicious sensing device in time slot t; when the random number is greater than the detection probability, the own sensing device cannot detect the airborne malicious sensing device, that is, F s,t = 0.

[0106] Whether the airborne malicious sensing device detects the own target can be simulated by evaluating the detection probability P d,t . Calculate the detection probability of the airborne malicious sensing device for the own target, and generate a random number υ' ∈ (0,1) uniformly distributed in [0,1], and determine whether the target is detected. When the random number is less than the detection probability, that is, υ' < P d,t , the airborne malicious sensing device can sense the own target, that is, the sensing efficiency F d,t Of the airborne malicious sensing device is 1; when the random number is greater than the detection probability, the airborne malicious sensing device cannot sense the own target, that is, the sensing efficiency F d,t Of the airborne malicious sensing device is 0. And when the airborne malicious sensing device is interfered by the interference device, it is considered that the target is not detected. And when the airborne malicious sensing device is interfered by the jammer, that is, when the signal frequency point of the airborne malicious sensing device coincides with the frequency point used by the jammer, it is considered that the own target is not detected.

[0107] If the airborne malicious perception device operating in search mode has detected its own target in the first two detections, the third detection is a critical moment, and the first confrontation occurs at this time. In the third detection, if the target is not detected, the fourth detection is regarded as another critical moment, and the second confrontation occurs at this time; otherwise, the airborne malicious perception device switches to the tracking mode. If the working mode of the airborne malicious perception device remains in the search mode or changes from the tracking mode to the search mode, it is considered that the confrontation is successful, that is, the jammer has successfully jammed. If the working mode of the airborne malicious perception device changes from the search mode to the tracking mode or from the tracking mode to the locked state, it is considered that the confrontation fails.

[0108] In the scenario of malicious perception electromagnetic defense, first use the graph convolutional neural network to extract the features of non-Euclidean space data, and cluster the airborne malicious perception device by using the task cognitive drive method of the graph autoencoder based on the obtained parameters of the airborne malicious perception device. Construct an adversarial target graph model according to the perception signal parameters of the airborne malicious perception device. Construct an adjacency matrix A based on whether the moving directions of the airborne malicious perception devices are the same. If the airborne malicious perception device d and the airborne malicious perception device d' have the same direction, the element a d,d' of the adjacency matrix A is 1, a d',d is 1; if they are not the same, then a d,d' is 0, a d',d is 0, then the adjacency matrix A is expressed as

[0109]

[0110] where D is the number of airborne malicious perception devices.

[0111] To characterize multiple adversarial targets from multiple features such as the speed, position, and direction of the airborne malicious perception device, construct a feature matrix X as

[0112] X = [V d , h d , O d ,

[0113] where V d is the speed of the airborne malicious perception device, h d is the position, and O d is the direction.

[0114] Use the graph convolutional neural network to extract the high-dimensional feature vector Z of the perception signal parameters of the airborne malicious perception device,

[0115] Z = GCN(X, A),

[0116]

[0117] where is the symmetric normalized adjacency matrix. W 0 and W 1 are the parameters to be learned. The inner product is used as the decoder to reconstruct the original graph:

[0118]

[0119] where is the reconstructed adjacency matrix.

[0120] During the training process of GAE, the cross-entropy is used as the loss function:

[0121]

[0122] where y represents the value (0 or 1) of an element in the adjacency matrix A, is the corresponding element value in the reconstructed adjacency matrix with a value range of 0 to 1, and N is the number of agents.

[0123] Based on the generated high-dimensional feature vector Z, the k-means clustering algorithm is used to generate the clustering label vector G of the UAV-borne malicious perception devices, which is used as the category information to guide the jammer's interference countermeasure.

[0124] In the detection and interference of the own perception devices and the jammer with a fixed transmission power once, due to the existence of a failure probability, it is necessary to use different frequencies to detect and interfere with the UAV-borne malicious perception devices multiple times. Therefore, the transmission power selection is used as the coarse time granularity, and the spectrum decision is used as the fine time granularity to construct a hierarchical decision framework for multi-dimensional resource collaborative management and control based on two-layer time granularity, and a three-layer decision from top to bottom is established: the node clustering layer based on the target, the transmission power selection layer, and the spectrum decision layer of the frequency-using devices.

[0125] In the node clustering layer based on the target, based on the obtained perception information of the UAV-borne malicious perception devices, the potential features are obtained through GCN, and the UAV-borne malicious perception devices are clustered based on the target features; then the reinforcement learning method is used to decide the association between the countermeasure target and the device, and the association result is used as the macro-action information at the top layer is expressed as:

[0126]

[0127] where is the association action of the jammer w.

[0128] In the transmission power selection layer, the observations of each device in the micro-slot t include the number N of the UAV-borne malicious perception devices sensed x,t, the relative distance d of the detected malicious perception device on the UAV x,d,t , the decision action information of the previous time slot and use the macro-action information at the top layer as the decision condition for the next layer; the observation information of the transmission power selection layer is constructed as:

[0129]

[0130] where when x is a jammer, when x is one's own perception device, the beam pointing selection actions of each device are expressed as

[0131] In the spectrum decision layer of the frequency-using device, the observation of each frequency-using device x includes its own position h x , the received signal strength at each frequency point and inherit the macro-actions of the target clustering and device association layer based on features and the macro-actions of the transmission power selection layer are expressed as:

[0132]

[0133] wherein, the frequency-using decision actions of each device are expressed as

[0134] In the spectrum decision layer of the frequency-using device, the reward obtained by each own perception device s is defined as the effectiveness F of whether the own perception device successfully detects the malicious perception device on the UAV s,t , expressed as:

[0135]

[0136] Assume that one time slot is equal to Δ micro time slots. In the transmission power selection layer, each jammer w obtains a reward each own perception device s obtains a reward In the transmission power selection layer, define the policy network and value network of each frequency-using device x. In the time slot Δ, the transmission power selection actions of each device are expressed as:

[0137]

[0138] According to the obtained return of the device x and the output V() of the value network, calculate the first-time difference error:

[0139]

[0140] where γ is the attenuation factor.

[0141] The value network updates the value network parameters with the goal of minimizing the first temporal difference error. The loss function of the value network is defined as:

[0142]

[0143] where is the expected mean.

[0144] The update formula for the value network parameters is expressed as:

[0145]

[0146] where α is the learning rate.

[0147] Using the calculated first temporal difference error, update the policy network parameters along the gradient direction of the loss function J(θ). The loss function of the policy network is defined as:

[0148]

[0149] The update formula for the policy network parameters at each time slot is expressed as:

[0150]

[0151] In the spectrum decision layer of the frequency-using device, define the policy network and the value network of each frequency-using device x. In the micro time slot t, the frequency-using selection action of each device is expressed as:

[0152]

[0153] According to the obtained reward of device x and the output of the value network, calculate the second temporal difference error:

[0154]

[0155] The loss function of the value network is defined as:

[0156]

[0157] The update formula for the value network parameters is expressed as:

[0158]

[0159] The loss function of the policy network is defined as:

[0160]

[0161] The update formula for the policy network parameters in each time slot is expressed as:

[0162]

[0163] Each agent adopts an actor-critic method. First, the policy network and value network of each agent are initialized. At the beginning of micro time slot t, in the feature-based target clustering layer, based on the obtained sensing information of the airborne malicious sensing devices, potential features are obtained through GCN, and the airborne malicious sensing devices are clustered based on the target features. The clustering result is used as the macro-action information at the top layer The macro-action information at the upper layer is transmitted to each device x, and the transmission power selection in each time slot is based on the policy network to obtain and generate the transmission power selection actions of each device The spectrum scheduling decision is based on the policy network to obtain and generate the spectrum decision actions of each frequency-using device At the end of the micro time slot, the frequency-using device x calculates the TD error according to the obtained reward for updating the policy network and the value network After Δ micro time slots, the top-level macro-action ends. Each device x calculates the TD error according to the obtained reward for updating the policy network and the value network and generates a new top-level macro-action according to the new sensing information of the airborne malicious sensing devices to start a new round of decision-making Repeat the above steps until the algorithm converges, and use the action selection decision at this time as the optimal decision

[0164] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript, etc

[0165]

[0166] ​​This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions run by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.

[0167] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.

[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are run on the computer or other programmable device to generate a computer-implemented process, so that the instructions run on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.

[0169] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0170] Obviously, those skilled in the art can make various changes and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A task cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method, characterized in that: The method comprises the following steps: Step 1: For the electromagnetic defense system of malicious sensing drones, use your own sensing equipment and jammers to sense the malicious sensing equipment on the drone and implement jamming countermeasures based on task cognition to generate a malicious sensing electromagnetic defense scenario; Step 2: Build a graph model based on the spatial location information of the malicious sensing device on the drone, and use the graph autoencoder to cluster dynamic multiple adversarial targets to form multiple interference tasks; Step 3: Construct the resource scheduling decision process of the sensing equipment and jammer as a semi-Markov decision process; Step 4: Combine the hierarchical actor-evaluator method to optimize the joint detection and interference of one's own sensing equipment and jammers on malicious sensing equipment on drones, as well as the management and decision-making of multi-dimensional resources such as the electromagnetic spectrum; repeat the iteration until the multi-dimensional resource collaborative management algorithm driven by task cognition converges, and take the action selection decision at this time as the optimal decision.

2. The task cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method according to claim 1 is characterized in that: In step 1, the malicious sensing electromagnetic defense scenario includes a set of unmanned aerial vehicle malicious sensing devices D = {1, ..., d, ..., D}, a jammer set W = = {1, ..., w, ..., W}, a self-sensing device set S = {1, ..., s, ..., S}, and a self-protected target M; with the spatial rectangular coordinate system as a reference, the position of the unmanned aerial vehicle malicious sensing device is represented as h d ={X d ,Y d ,Z d }, the jammer position is h w ={X w ,Y w ,0}, the position of the own sensing device is h s ={X s ,Y s ,0}, the target position to be protected is h t ={X t ,Y t ,0}; The speed vector of the malicious sensing device on the drone is V d , the direction vector is O d ; There are N frequency domain channels in the frequency domain space of the malicious perception electromagnetic defense scenario. In time slot t, the malicious perception device on the drone transmits a perception signal to detect its own target, and the frequency domain channel used is f d,t , the sensing device of the own side detects the frequency domain channel occupancy of various frequency-using devices including the malicious sensing device on the unmanned aircraft, the jammer, and the sensing device of the own side in the time slot t; the jammer makes a decision based on the occupancy of each frequency domain channel, and transmits a jamming signal at the next time t+1 to interfere with the malicious sensing device on the unmanned aircraft. When the jammer decides the interference frequency point f at the time slot t w,t The malicious sensing device on the drone occupies the frequency point f in the next time slot d,t+1 When they overlap, the interference is considered effective.

3. The task cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method according to claim 1 is characterized in that: Step 2 further includes: In the malicious perception electromagnetic defense scenario, a graph convolutional neural network is used to extract the features of non-Euclidean spatial data. Based on the acquired parameters of the drone-borne malicious perception devices, the task cognition-driven method of the graph autoencoder is used to cluster the drone-borne malicious perception devices. Construct an adversarial target graph model based on the signal parameters sensed by the malicious sensing device onboard the drone; Construct the adjacency matrix A based on whether the moving directions of the malicious sensing devices on drones are the same: If the drone-mounted malicious sensing device d and the drone-mounted malicious sensing device d' are in the same direction, then the element a of the adjacency matrix A d,d' =1,a d',d =1; if they are not the same, then a d,d' =0,a d',d =0, D is the number of malicious sensing devices onboard drones; Construct the feature matrix X: X=[V d ,h d ,O d ] Among them, V d is the speed of the UAV-mounted malicious sensing device, h d is the position, O d For direction; A graph convolutional neural network is used to extract the high-dimensional feature vector Z of the sensing signal parameters of the malicious sensing device on the drone: Z=GCN(X,A) in is a symmetric normalized adjacency matrix; W0 and W1 are parameters that need to be learned; the inner product is used as a decoder to reconstruct the original graph: in, is the reconstructed adjacency matrix; During GAE training, cross entropy is used as the loss function: Among them, y represents the value of one of the elements in the adjacency matrix A, To reconstruct the adjacency matrix The value of the corresponding element in , N is the number of agents; Based on the generated high-dimensional feature vector Z, the k-means clustering algorithm is used to generate the clustered label vector G of the malicious sensing device on the drone as the category information to guide the jammer's interference countermeasures.

4. The task cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method according to claim 1 is characterized in that: Step 3 further comprises the following steps: Determine the clusters of different adversarial targets targeted by the jammer, use the transmit power selection as the coarse time granularity and the spectrum decision as the fine time granularity to build a hierarchical decision framework for multi-dimensional resource collaborative management and control based on two layers of time granularity, and establish a top-down three-layer decision-making: target-based node clustering layer, transmit power selection layer, and frequency equipment spectrum decision layer; In the target-based node clustering layer, based on the perception information of the drone-borne malicious perception devices, the potential features are obtained through GCN, and the drone-borne malicious perception devices are clustered based on the target features; then the reinforcement learning method is used to decide the association between the adversarial targets and devices, and the association results are used as the macro action information of the top layer. It is expressed as: in, is the associated action of the jammer w; In the transmission power selection layer, the observation of each device in the mini-time slot t Including the number of malicious sensing devices on drones N x,t , the relative distance d of the detected malicious sensing device on the drone x,d,t , upper time slot decision action information And with the top-level macro action information is the decision condition for the next layer; the observation information of the transmission power selection layer is constructed as: When x is a jammer, When x is the own sensing device, The beam pointing selection action of each device is expressed as In the spectrum decision layer of frequency-using equipment, the observation of each frequency-using equipment x Including its own position h x , Receive signal strength at each frequency And inherit the macro actions of feature-based target clustering and device association layer Macro Action with Transmit Power Selection Layer It is expressed as: Among them, the frequency decision action of each device is expressed as In the spectrum decision layer of frequency-using equipment, each self-sensing device s is rewarded Defined as the effectiveness of the self-sensing device in detecting the malicious sensing device on the drone. s,t , expressed as: Assuming that one time slot is equal to Δ micro-time slots, in the transmission power selection layer, each jammer w receives a reward Each of your own sensing devices receives a reward In the transmission power selection layer, define the strategy network of each frequency-using device x and value network In the time slot Δ, the transmit power selection action of each device is expressed as: Based on the return of device x And the value network output V(), calculate the first time difference error: Where γ is the attenuation factor; The value network updates the value network parameters with the goal of minimizing the first time difference error. The loss function of the value network is defined as: in is the expected mean; The update formula of the value network parameters is expressed as: Where α is the learning rate; Using the calculated first time difference error, update the policy network parameters along the gradient direction of the loss function J(θ). The loss function of the policy network is defined as: The update formula of the policy network parameters in each time slot is expressed as: In the spectrum decision layer of frequency-using equipment, define the strategy network of each frequency-using equipment x and value network In the mini-time slot t, the frequency selection action of each device is expressed as: Based on the return of device x And the value network output, calculate the second time difference error: The loss function of the value network is defined as: The update formula of the value network parameters is expressed as: The policy network loss function is defined as: The update formula of the policy network parameters in each time slot is expressed as:

5. The task cognition-driven low-altitude electromagnetic defense resource collaborative dynamic management method according to claim 1 is characterized in that: Step 4 further includes: S41, initializing the strategy network and value network of each agent; S42, at the beginning of the micro-time slot t, in the feature-based target clustering layer, based on the acquired perception information of the drone-borne malicious perception device, the potential features are obtained through GCN, and the drone-borne malicious perception devices are clustered based on the target features. The clustering results are used as the macro action information of the top layer. Upper-level macro action information Transmitted to each device x, the transmission power selection in each time slot is based on the strategy network Get and generate the transmit power selection action for each device Spectrum scheduling decisions based on strategic networks Acquire and generate spectrum decision actions for each frequency-using device At the end of the micro-slot, the frequency device x receives the reward Calculate the second time difference error For policy networks and value network Updates; After Δ micro-time slots, the top-level macro action ends, and each device x receives the reward Calculate the first time difference error For policy networks and value network and generate new top-level macro actions based on the new drone-borne malicious perception device perception information S43, repeat step S42 until the multi-dimensional resource collaborative management and control algorithm driven by task cognition converges, and the action selection decision at this time is taken as the optimal decision.

6. A task-cognition driven low-altitude electromagnetic defense resource collaborative dynamic management system, characterized in that: The system comprises: The scenario construction module is used for the electromagnetic defense system against malicious sensing drones. It uses its own sensing equipment and jammers to sense the malicious sensing equipment on drones and implement jamming countermeasures based on task cognition to generate malicious sensing electromagnetic defense scenarios. The clustering module is used to build a graph model based on the spatial location information of the malicious sensing device on the drone, and use the graph autoencoder to cluster dynamic multiple adversarial targets to form multiple interference tasks; Resource scheduling module, used to construct the resource scheduling decision process of the own sensing equipment and jammer into a semi-Markov decision process; The decision-making module combines the hierarchical actor-evaluator method to optimize the joint detection and interference of one's own sensing equipment and jammers on malicious sensing equipment on drones, as well as the management and decision-making of multi-dimensional resources such as the electromagnetic spectrum; iterates repeatedly until the multi-dimensional resource collaborative management algorithm driven by task cognition converges, and the action selection decision at this time is taken as the optimal decision.

Citation Information

Patent Citations

  • Unmanned aerial vehicle signal suppression method and device, electronic equipment and storage medium

    CN111669248A

  • Method and system for implementing electromagnetic spectrum interference by unmanned aerial vehicle cluster cooperating with machine learning

    CN117539270A

  • Unmanned aerial vehicle auxiliary communication anti-interference method based on multi-agent reinforcement learning

    CN118921099A

  • Air-ground cooperative communication hierarchical decision-making method based on environmental cognition

    CN119071817A

  • Intelligent communication countermeasure method and system

    WO2025020552A1