An unmanned aerial vehicle autonomous jamming decision method and device based on historical update impulse

CN117335859BActive Publication Date: 2026-08-28NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311151730.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2026-08-28
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

[0002]目前无人机干扰技术领域内,采用如马尔可夫决策、强化学习模型等干扰决策模型,针对目标任务无人机进行通信干扰,传统的强化学习模型存在样本效率低下的问题,近年来,通过联邦强化学习模型可以解决大样本需求和低采样效率之间的矛盾,利用传统的联邦强化学习模型(Federated Reinforcement Learning,FRL)采用联邦平均(FederatedAveraging,FedAvg)的方法,虽然可以将不同无人机的策略权重通过简单平均的方式进行策略聚合和更新,但是无人机在面对连续动作、状态空间等复杂任务环境时,由于对无人机所处地理环境、通信链路情况、干扰信号馈送信息、受干扰无人机的状态和动作等探索信息量较大、探索信息样本分布差异显著、无人机异构等原因,干扰决策模型更新梯度会引入巨大的方差,使得无人机与服务器之间的信息传输量变大,基于无人机的历史信息得到的干扰策略容易陷入局部最优的问题,干扰链路过程中一旦被截断或者信息缺失,将使得整个干扰决策过程鲁棒性较差、信息获取能力受限等问题,导致无人机自主干扰决策失准,干扰效果欠佳

Benefits of technology

[0034]The aforementioned method and apparatus for UAV autonomous interference decision-making based on historical update impulses constructs a UAV interference network and substitutes the global interference decision model into the sub-interference decision models of the working nodes. This ensures the robustness of the process of obtaining the overall interference decision scheme and avoids getting trapped in local optima. Furthermore, in addition to transmitting sub-interference strategies between the central server and several working nodes, the working nodes also need to feed back their sub-model parameters to the central server. Performance testing is then conducted on the central server to further verify the accuracy of the training process of the sub-interference decision models of the working nodes. This, combined with historical update impulses, calibrates the global interference decision model, resulting in a more realistic and reliable global interference strategy obtained after training on UAV interference resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117335859B_ABST
    Figure CN117335859B_ABST
Patent Text Reader

Abstract

The application relates to an unmanned aerial vehicle autonomous interference decision method and device based on historical update impulse. The method comprises the following steps: acquiring a training sample data set; inputting the training sample data set into a sub-interference decision model of a working node for training, obtaining a sub-interference strategy, iteratively updating the sub-interference decision model according to the sub-interference strategy and a preset sub-sampling threshold, and obtaining sub-model parameters. The performance of the sub-interference strategy corresponding to the working node is tested by a central server according to the sub-model parameters, and an interference reward value of the working node is obtained. Global interference decision average model parameters are generated according to the interference reward value and the sub-model parameters, the parameters of the global interference decision model are updated based on the historical update impulse and the global interference decision average model parameters, the global interference decision model is optimized according to the updated parameters of the global interference decision model, and a global interference strategy is obtained. The method can obtain reliable and accurate unmanned aerial vehicle interference strategies in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) jamming technology, and in particular to a method and apparatus for autonomous UAV jamming decision-making based on historical update impulses. Background Technology

[0002] Currently, in the field of UAV jamming technology, jamming decision models such as Markov decision-making and reinforcement learning models are used to jam the communication of target UAVs. Traditional reinforcement learning models suffer from low sample efficiency. In recent years, federated reinforcement learning models have been developed to resolve the contradiction between the need for large samples and low sampling efficiency. While Federated Averaging (FedAvg) in FRL (Functional Learning) can aggregate and update policy weights from different UAVs through simple averaging, it faces challenges when UAVs encounter complex environments such as continuous actions and state spaces. These challenges arise from the large amount of information to explore, including the geographical environment, communication links, interference signal feeds, and the state and actions of the interfered UAVs. Furthermore, the significant differences in the distribution of these information samples and the heterogeneity of the UAVs themselves contribute to the high variance introduced into the gradient update of the interference decision model. This increases the amount of information transmitted between the UAVs and the server, making interference policies based on historical UAV information prone to getting trapped in local optima. Furthermore, if the interference link is interrupted or information is missing, the entire interference decision-making process suffers from poor robustness and limited information acquisition capabilities, leading to inaccurate autonomous interference decisions and suboptimal interference effectiveness. Summary of the Invention

[0003] Therefore, it is necessary to address the aforementioned technical problems by providing an autonomous UAV interference decision-making method based on historically updated impulses that can improve the accuracy of real-time UAV interference decision-making.

[0004] A method for autonomous UAV interference decision-making based on historical update impulse is applied to a UAV interference network architecture, which includes a central server and several working nodes. The method includes:

[0005] Obtain the training sample dataset, which includes: global interference decision model parameters and UAV interference resources.

[0006] The training sample dataset is input into the sub-interference decision model of the working node for training to obtain the sub-interference policy. The sub-interference decision model is iteratively updated according to the sub-interference policy and the preset sub-sampling threshold to obtain the sub-model parameters.

[0007] The central server performs performance tests on the sub-interference strategies corresponding to the worker nodes based on the sub-model parameters, and obtains the interference reward value of the worker nodes.

[0008] The interference reward value and sub-model parameters are calculated to generate the global interference decision average model parameters. The parameters of the global interference decision model are updated based on the historical update impulse and the global interference decision average model parameters. The global interference decision model is optimized based on the updated global interference decision model parameters to obtain the global interference strategy.

[0009] In one embodiment, the central server includes a control drone used to broadcast parameters of the global interference decision model to several working nodes in the drone interference network architecture, collect sub-model parameters from the working nodes using a synchronous communication mechanism, generate average global interference decision model parameters based on the average interference reward value of the sub-model parameters, and calculate the historical update impulse of the average global interference decision model parameters. The working nodes include multiple interference drones, which are used to input collected interference resources as training samples into the sub-interference decision models. The sub-interference decision models use reinforcement learning strategies to train the training samples and update their parameters, resulting in sub-interference strategies and sub-model parameters of the sub-interference decision models.

[0010] In one embodiment, the drone interference resources include: target drone coordinate information, target drone attitude information, image information of the geographical environment where the target drone is located, and the target drone's communication protocol.

[0011] In one embodiment, the method further includes: inputting a training sample dataset into the sub-interference decision model of the working node; obtaining target mission schemes using Markov decision methods in the UAV interference environment; storing the current target mission scheme in the experience replay library of the working node to obtain historical target mission schemes; training the historical target mission schemes and global interference decision model parameters using the sub-interference decision model to obtain a sub-interference strategy; iteratively updating the sub-interference decision model according to the sub-interference strategy and preset sub-sampling thresholds in the experience replay library to obtain the sub-model parameters of the working node.

[0012] In one embodiment, the method further includes: using a central server to perform performance tests on the sub-interference strategies corresponding to different working nodes according to a preset test count threshold and sub-model parameters, and saving the obtained interference reward values ​​of the working nodes.

[0013] In one embodiment, the method further includes: calculating the mean of the interference reward value.

[0014]

[0015] in, The interference reward value is E, which is the preset threshold for the number of tests, and k is the index identifier of the worker node. This represents the mean of the interference reward values. The mean of the interference reward values ​​is weighted by the sub-model parameters of different working nodes to obtain the global interference decision average model parameters:

[0016]

[0017] in, These are the parameters of the global disturbance decision-making average model. Let be the sub-model parameters for iteration t, and K be the number of worker nodes. Let k be the average of the interference reward values, and k be the index of the working node. Based on historical update impulses and the average parameters of the global interference decision model, the parameter update amount of the global interference decision model is updated using a block update filtering algorithm:

[0018]

[0019] Where, Δ t-1 Δ represents the parameter update amount of the global disturbance decision model in the previous iteration. t μ represents the parameter update amount of the global disturbance decision model in the current iteration round. t Let ξ be the impulse coefficient for the current iteration round. t Let θ be the learning rate for the current iteration. t Let θ be the global disturbance decision average model parameter for the current iteration round. t-1 These are the parameters for the global disturbance decision in the previous iteration.

[0020] Summing the parameters of the global disturbance decision model in the current iteration with the parameter update amount of the global disturbance decision model yields the updated parameters of the global disturbance decision model in the current iteration.

[0021] θ t+1 =θ t +Δ t

[0022] Where, Δ t θ represents the parameter update amount of the global disturbance decision model in the current iteration round. t Let θ be the global disturbance decision average model parameter for the current iteration round. t+1 These are the parameters of the global disturbance decision model after the current iteration. The global disturbance decision model is optimized based on the updated parameters to obtain the global disturbance strategy.

[0023] In one embodiment, the method further includes: acquiring global interference decision model parameters by controlling the drone, and the interference drone acquiring drone interference resources by exploring the environmental space.

[0024] An autonomous interference decision-making device for unmanned aerial vehicles based on historical update impulse, the device comprising:

[0025] The training sample dataset acquisition module is used to acquire the training sample dataset, which includes: global interference decision model parameters and UAV interference resources.

[0026] The sub-model parameter acquisition module is used to input the training sample dataset into the sub-interference decision model of the working node for training, obtain the sub-interference policy, and iteratively update the sub-interference decision model according to the sub-interference policy and the preset sub-sampling threshold to obtain the sub-model parameters.

[0027] The interference reward value acquisition module is used to perform performance tests on the sub-interference strategies corresponding to the worker nodes based on the sub-model parameters through the central server, and obtain the interference reward value of the worker nodes.

[0028] The interference decision module is used to generate global interference decision average model parameters based on interference reward value and sub-model parameters, update global interference decision model parameters based on historical update impulse and global interference decision average model parameters, optimize global interference decision model based on updated global interference decision model parameters, and obtain global interference strategy.

[0029] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:

[0030] Obtain the training sample dataset, which includes: global interference decision model parameters and UAV interference resources.

[0031] The training sample dataset is input into the sub-interference decision model of the working node for training to obtain the sub-interference policy. The sub-interference decision model is iteratively updated according to the sub-interference policy and the preset sub-sampling threshold to obtain the sub-model parameters.

[0032] The central server performs performance tests on the sub-interference strategies corresponding to the worker nodes based on the sub-model parameters, and obtains the interference reward value of the worker nodes.

[0033] The global interference decision average model parameters are generated based on the interference reward value and the sub-model parameters. The parameters of the global interference decision model are updated based on the historical update impulse and the global interference decision average model parameters. The global interference decision model is optimized based on the updated global interference decision model parameters to obtain the global interference strategy.

[0034] The aforementioned method and apparatus for UAV autonomous interference decision-making based on historical update impulses constructs a UAV interference network and substitutes the global interference decision model into the sub-interference decision models of the working nodes. This ensures the robustness of the process of obtaining the overall interference decision scheme and avoids getting trapped in local optima. Furthermore, in addition to transmitting sub-interference strategies between the central server and several working nodes, the working nodes also need to feed back their sub-model parameters to the central server. Performance testing is then conducted on the central server to further verify the accuracy of the training process of the sub-interference decision models of the working nodes. This, combined with historical update impulses, calibrates the global interference decision model, resulting in a more realistic and reliable global interference strategy obtained after training on UAV interference resources. Attached Figure Description

[0035] Figure 1 This is an application scenario diagram of an autonomous interference decision-making method for UAVs based on historical update impulse, as shown in one embodiment.

[0036] Figure 2 This is a flowchart illustrating an autonomous interference decision-making method for unmanned aerial vehicles (UAVs) based on historical update impulse in one embodiment.

[0037] Figure 3 The following are performance curves of the proposed method and two baseline algorithms in five environments, as shown in one embodiment; wherein, Figure 3 (a) is the HalfCheetah-v2 environment. Figure 3 (b) is the Swimmer-v2 environment. Figure 3 (c) is the InvertedDoublePendulum-v2 environment. Figure 3 (d) represents the Walker2d-v2 environment. Figure 3 (e) represents the Ant-v2 environment;

[0038] Figure 4 Here are the performance curves of the MFRL method under different impulse coefficients in one embodiment; where, Figure 4 (a) is the HalfCheetah-v2 environment. Figure 4 (b) is the Swimmer-v2 environment;

[0039] Figure 5 This is a performance comparison chart of the MFRL method and the FedAvg algorithm under different communication intervals in one embodiment; where, Figure 5 (a) is the HalfCheetah-v2 environment. Figure 5 (b) is the Swimmer-v2 environment;

[0040] Figure 6 This is a structural diagram of an autonomous interference decision-making device for unmanned aerial vehicles based on historical update impulse in one embodiment.

[0041] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0043] This application provides a method for autonomous interference decision-making of unmanned aerial vehicles (UAVs) based on historically updated impulse, which can be applied to, for example... Figure 1 The diagram illustrates an autonomous interference decision-making network environment for unmanned aerial vehicles (UAVs). This environment includes a central server and several UAV worker nodes, with worker nodes A, B, and C belonging to different network domains. The central server comprises a global agent, and each worker node corresponds to a local agent A, local agent B, and local agent C. Multiple local agents can be connected. Worker nodes A, B, and C upload data to the central server via network links, and the central server broadcasts data to each worker node via network links. The interference decision-making sub-models among the worker nodes are independent. Data interactions between worker nodes must be processed by the central server before being broadcast to the corresponding worker nodes.

[0044] In one embodiment, such as Figure 2 As shown, a method for autonomous interference decision-making of UAVs based on historical update impulse is provided, and this method is applied to... Figure 1 Taking the autonomous interference decision-making network environment of unmanned aerial vehicles (UAVs) as an example, the following steps are included:

[0045] Step 102: Obtain the training sample dataset, which includes: global interference decision model parameters and UAV interference resources.

[0046] Step 104: Input the training sample dataset into the sub-interference decision model of the working node for training to obtain the sub-interference strategy. Iteratively update the sub-interference decision model according to the sub-interference strategy and the preset sub-sampling threshold to obtain the sub-model parameters.

[0047] Step 106: The central server performs performance testing on the sub-interference strategy corresponding to the working node based on the sub-model parameters to obtain the interference reward value of the working node.

[0048] Step 108: Generate the global interference decision average model parameters based on the interference reward value and the sub-model parameters; update the parameters of the global interference decision model based on the historical update impulse and the global interference decision average model parameters; optimize the global interference decision model based on the updated global interference decision model parameters to obtain the global interference strategy.

[0049] In the aforementioned UAV autonomous interference decision-making method based on historical update impulses, a UAV interference network is constructed, and the global interference decision-making model is substituted into the sub-interference decision-making models of the working nodes. This ensures the robustness of the process of obtaining the overall interference decision-making scheme and avoids getting trapped in local optimum decision-making problems. Furthermore, in addition to transmitting sub-interference strategies between the central server and several working nodes, the working nodes also need to feed back their sub-model parameters to the central server. Performance testing is then conducted on the central server to further verify the accuracy of the training process of the working nodes' sub-interference decision-making models. This, combined with historical update impulses, calibrates the global interference decision-making model, resulting in a more realistic and reliable global interference strategy obtained after training the UAV interference resources.

[0050] In one embodiment, the central server includes a control drone used to broadcast parameters of the global interference decision model to several working nodes in the drone interference network architecture, collect sub-model parameters from the working nodes using a synchronous communication mechanism, generate average global interference decision model parameters based on the average interference reward value of the sub-model parameters, and calculate the historical update impulse of the average global interference decision model parameters. The working nodes include multiple interference drones, which are used to input collected interference resources as training samples into the sub-interference decision models. The sub-interference decision models use reinforcement learning strategies to train the training samples and update their parameters, resulting in sub-interference strategies and sub-model parameters of the sub-interference decision models.

[0051] In one embodiment, the drone interference resources include: target drone coordinate information, target drone attitude information, image information of the geographical environment where the target drone is located, and the target drone's communication protocol.

[0052] In one embodiment, the training sample dataset is input into the sub-interference decision model of the working node. The UAV interference environment uses a Markov decision method to obtain the target mission plan, and the target mission plan at the current moment is stored in the experience replay library of the working node to obtain the historical target mission plan. After the historical target mission plan and the global interference decision model parameters are trained by the sub-interference decision model, the sub-interference strategy is obtained. The sub-interference decision model is iteratively updated according to the sub-interference strategy and the preset sub-sampling threshold in the experience replay library to obtain the sub-model parameters of the working node.

[0053] In one embodiment, the central server performs performance tests on the sub-interference strategies corresponding to different working nodes according to a preset test count threshold and sub-model parameters, and saves the interference reward value of the working nodes.

[0054] In one embodiment, the mean of the interference reward value is calculated:

[0055]

[0056] in, The interference reward value is E, which is the preset threshold for the number of tests, and k is the index identifier of the worker node. This represents the mean of the interference reward values. The mean of the interference reward values ​​is weighted by the sub-model parameters of different working nodes to obtain the global interference decision average model parameters:

[0057]

[0058] in, These are the parameters of the global disturbance decision-making average model. Let be the sub-model parameters for iteration t, and K be the number of worker nodes. Let k be the average of the interference reward values, and k be the index of the working node. Based on historical update impulses and the average parameters of the global interference decision model, the parameter update amount of the global interference decision model is updated using a block update filtering algorithm:

[0059]

[0060] Where, Δ t-1 Δ represents the parameter update amount of the global disturbance decision model in the previous iteration. t μ represents the parameter update amount of the global disturbance decision model in the current iteration round. t Let ξ be the impulse coefficient for the current iteration round. t Let θ be the learning rate for the current iteration. t Let θ be the global disturbance decision average model parameter for the current iteration round. t-1 These are the parameters for the global disturbance decision in the previous iteration.

[0061] Summing the parameters of the global disturbance decision model in the current iteration with the parameter update amount of the global disturbance decision model yields the updated parameters of the global disturbance decision model in the current iteration.

[0062] θ t+1 =θ t +Δ t

[0063] Where, Δ tθ represents the parameter update amount of the global disturbance decision model in the current iteration round. t Let θ be the global disturbance decision average model parameter for the current iteration round. t+1 These are the parameters of the global disturbance decision model after the current iteration. The global disturbance decision model is optimized based on the updated parameters to obtain the global disturbance strategy.

[0064] In one embodiment, the parameters of the global interference decision model are obtained by controlling the drone, and the interference drone obtains drone interference resources by exploring the environmental space.

[0065] In one embodiment, it is assumed that there is a central server and K worker nodes, where the global agent model parameters of the central server are denoted as .

[0066] θ t t∈{0,1,…,T-1}

[0067] The superscript t indicates the number of iterations.

[0068] The parameters of the local agent model deployed on each of the K working nodes are denoted as follows:

[0069]

[0070] In this context, the subscript k represents the working node number.

[0071] Iteration first requires adjusting the global agent model parameters θ 0 The number of worker node agents K, the number of iteration rounds T, the communication interval M, and the number of model testing rounds E are initialized. Each iteration mainly consists of three stages, which are executed alternately by the central server and worker nodes. We will introduce the different stages in detail below.

[0072] In the first phase, the central server will assign global agent model parameters θ t Broadcast;

[0073] In the second stage, K worker nodes execute in parallel. Taking worker node agent k as an example, it first receives the global agent model parameters θ. t and using global model parameters θ t Replace local model parameters After parameter replacement, according to the new agent policy Interact with the local environment to obtain new experience samples (s, a, s′, r, done).

[0074] It's important to note that unlike traditional FL frameworks where worker nodes have local training datasets, in the FRL framework, the local agent lacks historical samples for model training beforehand. Therefore, it needs to accumulate experience samples through continuous interaction with the local environment after training begins. Thus, after each environment step, it's necessary to determine if the number of samples in the local experience replay library has reached the batch size for model updates. If it has reached the update threshold, then... size Then, a batch of samples with a batch size of 100 is sampled from the experience sample replay library. size The sample set is used to update the local policy model parameters according to the specific reinforcement learning strategy adopted.

[0075]

[0076] If the communication interval M between the central server and the worker nodes is greater than 1, the worker node will restart the environment and repeat the process of interacting with the environment and updating the local policy. This process continues until after M rounds of the game, at which point the worker node agent will update the local model parameters. Uploaded to the central server;

[0077] In the third stage, the central server will use a synchronous communication mechanism to transmit the agent model parameters of the K working nodes. All data was collected and sent to the server. Then, in the server test environment, different model strategies were used on different worker nodes. Perform performance tests, testing each model E times, and save the cumulative reward value obtained throughout the entire game.

[0078] In the model aggregation phase, the average score of the agents is first calculated based on the model test results.

[0079]

[0080] Then, the model parameters for different working nodes are weighted according to the average score of the agents, and the weighted average model parameters are calculated.

[0081]

[0082] Similar to the BMUF (Block-wise Model Update Filtering) framework, the update amount Δ t The value consists of two parts: the historical model update amount and the current update amount.

[0083]

[0084] In the experiment presented in this paper, the impulse coefficient μ t and learning rate ξ tAll values ​​are fixed and do not change over time. Finally, the global agent model parameters after this round of iteration are obtained.

[0085] θ t+1 =θ t +Δ t

[0086] Repeat the above three stages until the model performance meets the training requirements or t=T.

[0087] The pseudocode for the MFRL (Friction Based on Historical Model Update) method is as follows:

[0088]

[0089]

[0090] In one embodiment, such as Figure 3 As shown, the performance of the MFRL method is evaluated in five real-world continuous control tasks within the rllab framework: Figure 3 (a) is the HalfCheetah-v2 environment. Figure 3 (b) is the Swimmer-v2 environment. Figure 3 (c) is the InvertedDoublePendulum-v2 environment. Figure 3 (d) represents the Walker2d-v2 environment. Figure 3 (e) represents the Ant-v2 environment. For all these tasks, the game length is set to 500, assuming multiple worker node devices and a central server, where each edge computing platform runs an independent environment with the same environment dynamics, and the number of worker node devices is set to 3. The MFRL method and all baseline algorithms are implemented using PyTorch. Specifically, this paper selects the SAC algorithm as the concrete implementation of the reinforcement learning algorithm in the MFRL method. This algorithm includes one Actor network, two Critic networks, and two Target Critic networks. All networks use the same feedforward neural network structure, containing two hidden layers, each containing 256 neurons, and ReLU is used as the activation function. Furthermore, the output of the policy neural network is a Gaussian distribution N(μ(s), σ... 2 ).

[0091] For the other default settings of the MFRL method, the number of inner loops per training epoch is set to 1. Each worker node drone first executes 1000 environmental steps using a random policy and stores historical samples locally before starting to train a local model and execute subsequent actions using a local policy. The aim is to reduce the impact of policy initialization quality on model performance and to collect more extensive experience samples. We set the batch size for local updates of all worker nodes to 256, and the experience replay library size to 1e6. The neural network uses the Adam optimizer for updates, with a learning rate of 3e-4, a cumulative reward discount rate γ of 0.99, and in the SAC algorithm, the initial temperature parameter α is set to 0.05, and the soft update coefficients for the Critic and Target Critic network parameters are set to τ = 0.005.

[0092] Specifically, the standard SAC algorithm and the FedAvg algorithm were selected as the baseline algorithms for the comparative experiment. To ensure experimental fairness, all algorithm implementation details were kept consistent with the MFRL method, with the FedAvg algorithm also employing an architecture of three worker nodes and one central server. Figure 3 As can be seen, the MFRL method and FedAvg algorithm, employing the FRL architecture, are significantly faster in learning speed compared to the standard SAC algorithm. The MFRL method, due to the addition of impulse, results in a faster rate of performance improvement in the early stages of training compared to FedAvg. This is because in the early stages of policy exploration, the model update gradient variance is large, and the update magnitude is small. The MFRL method appropriately retains historical model updates as impulse, increasing the model update magnitude and accelerating the model's escape from local optima that are easily trapped in the early stages of exploration. Furthermore, the introduction of impulse makes the MFRL method more stable in training than the FedAvg algorithm, and it can converge to higher score levels in more complex continuous action control tasks.

[0093] Furthermore, sensitivity analysis was conducted by setting the impulse coefficient and communication interval in the model as variables.

[0094] The magnitude of the impulse coefficient μ determines the impact of historical model updates on the current training amplitude. Generally, both excessively large and excessively small values ​​will affect algorithm performance. To more accurately and quantitatively analyze the degree to which algorithm performance is affected by the impulse coefficient, experiments were conducted on the HalfCheetah-v2 and Swimmer-v2 environments to test the model performance under different impulse coefficients. The experimental results are as follows: Figure 4 As shown in Table 1, Figure 4 (a) is the HalfCheetah-v2 environment. Figure 4 (b) is the Swimmer-v2 environment, as shown in Table 1 below:

[0095] Table 1. Average scores of the MFRL method under different impulse coefficients.

[0096]

[0097]

[0098] Experimental results show that excessively large impulse coefficients can lead to gradient explosion, causing the loss values ​​of both the policy and value networks to overflow. Gradient explosion occurs in both environments when μ ∈ [0.5, 0.9], indicating that excessive impulse coefficients severely impact model convergence in the early stages of training, resulting in highly unstable training. In the HalfCheetah-v2 environment, performance degrades when μ = 0.01, while the convergence results are essentially the same for μ values ​​between 0.05 and 0.3. In the Swimmer-v2 environment, the best results are achieved when μ = 0.2, while model performance remains relatively consistent for other μ values. Overall, there is a threshold range for μ. When the impulse coefficient exceeds this range, gradient explosion leads to catastrophic results. However, when μ is within this range, model performance is not sensitive to the impulse coefficient, and models with different impulse coefficients can achieve good convergence speeds and results.

[0099] Furthermore, due to hardware limitations, FRL suffers from communication constraints, especially in mobile scenarios. The FL framework places higher demands on communication; for example, worker nodes must be connected to a wireless network and charging to participate in the global model training on the central server. Therefore, this method conducts a sensitivity analysis to address whether the communication interval between worker nodes and the server affects model performance.

[0100] like Figure 5 The figure shows a performance comparison under different communication intervals, where... Figure 5 (a) is the HalfCheetah-v2 environment. Figure 5 (b) shows the average reward obtained by the MFRL method and the FedAvg algorithm in the Swimmer-v2 environment. As the interval gradually increases, the performance of both algorithms fluctuates to some extent, but the MFRL method achieves better model performance than FedAvg across all intervals.

[0101] In the HalfCheetah-v2 environment, the MFRL method achieved an average score of 4184.71 across four communication intervals, while the FedAvg algorithm averaged 3657.02. The MFRL method outperformed the FedAvg algorithm by 14.43% on average, with the performance difference reaching 37.81% when the communication interval M=3. In the Swimmer-v2 environment, the MFRL method achieved an average score of 55.11 across four communication intervals, while the FedAvg algorithm averaged 49.35. The MFRL method outperformed the FedAvg algorithm by 11.67% on average, with the performance difference reaching 17.11% when the communication interval M=50.

[0102] Furthermore, in the HalfCheetah-v2 environment, the performance of the FedAvg algorithm degrades significantly when M=3. We believe this is because the excessively small communication interval leads to frequent aggregation of model parameters, which exacerbates training instability to some extent. Our proposed method effectively compensates for the performance degradation caused by drastic parameter changes in the early stages of training by increasing the historical update impulse. Therefore, our method exhibits good robustness to changes in communication intervals, achieving performance superior to the FedAvg algorithm under various communication intervals. Even in situations where information is missing or ambiguous during UAV autonomous interference decision-making, it can still yield a reliable global interference strategy.

[0103] Therefore, it is evident that traditional FRL algorithms suffer from problems such as unstable model training, limited exploratory capabilities in different environments, large model update gradient variance, and susceptibility to local optima when facing complex task environments such as continuous actions and state spaces. Our proposed method, an FRL algorithm based on historical model update impulse, addresses these issues. Building upon traditional FRL, this algorithm introduces the accumulated impulse from historical model updates during model aggregation, effectively mitigating risks such as excessive model update gradient variance, training instability, and susceptibility to local optima. Experimental results on several classic continuous control tasks show that our method outperforms the FedAvg algorithm and the baseline SAC algorithm in terms of convergence speed and average score. In sensitivity analysis experiments on the key hyperparameters impulse coefficient μ and communication interval M, our method's average score surpasses that of the FedAvg algorithm, demonstrating better robustness to hyperparameters and resulting in higher stability and accuracy in the UAV's autonomous interference decision-making process.

[0104] It should be understood that, although Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0105] In one embodiment, such as Figure 6 As shown, an autonomous interference decision-making device for unmanned aerial vehicles (UAVs) based on historical update impulse is provided, comprising: a training sample dataset acquisition module 602, a sub-model parameter acquisition module 604, an interference reward value acquisition module 606, and an interference decision-making module 608, wherein:

[0106] The training sample dataset acquisition module 602 is used to acquire the training sample dataset, which includes: global interference decision model parameters and UAV interference resources.

[0107] The sub-model parameter acquisition module 604 is used to input the training sample dataset into the sub-interference decision model of the working node for training, obtain the sub-interference strategy, and iteratively update the sub-interference decision model according to the sub-interference strategy and the preset sub-sampling threshold to obtain the sub-model parameters.

[0108] The interference reward value acquisition module 606 is used to perform performance testing on the sub-interference strategy corresponding to the working node through the central server based on the sub-model parameters, and obtain the interference reward value of the working node.

[0109] The interference decision module 608 is used to generate global interference decision average model parameters based on interference reward value and sub-model parameters, update global interference decision model parameters based on historical update impulse and global interference decision average model parameters, optimize global interference decision model based on updated global interference decision model parameters, and obtain global interference strategy.

[0110] For specific limitations regarding the UAV autonomous interference decision-making device based on historical update impulse, please refer to the limitations of the UAV autonomous interference decision-making method based on historical update impulse mentioned above, which will not be repeated here. Each module in the aforementioned UAV autonomous interference decision-making device based on historical update impulse can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0111] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements an autonomous interference decision-making method for unmanned aerial vehicles based on historical update impulses. The display screen can be an LCD screen or an e-ink display screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0112] Those skilled in the art will understand that Figures 6-7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0113] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to perform the following steps:

[0114] Obtain the training sample dataset, which includes: global interference decision model parameters and UAV interference resources.

[0115] The training sample dataset is input into the sub-interference decision model of the working node for training to obtain the sub-interference policy. The sub-interference decision model is iteratively updated according to the sub-interference policy and the preset sub-sampling threshold to obtain the sub-model parameters.

[0116] The central server performs performance tests on the sub-interference strategies corresponding to the worker nodes based on the sub-model parameters, and obtains the interference reward value of the worker nodes.

[0117] The global interference decision average model parameters are generated based on the interference reward value and the sub-model parameters. The parameters of the global interference decision model are updated based on the historical update impulse and the global interference decision average model parameters. The global interference decision model is optimized based on the updated global interference decision model parameters to obtain the global interference strategy.

[0118] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0119] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0120] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for autonomous interference decision-making of unmanned aerial vehicles based on historical update impulse, characterized in that, This is applied to a drone jamming network architecture, which includes a central server and several working nodes. The method includes: Obtain a training sample dataset, which includes: global interference decision model parameters and UAV interference resources; The training sample dataset is input into the sub-interference decision model of the working node for training to obtain the sub-interference strategy. The sub-interference decision model is iteratively updated according to the sub-interference strategy and the preset sub-sampling threshold to obtain the sub-model parameters. The central server performs performance testing on the sub-interference strategy corresponding to the working node based on the sub-model parameters to obtain the interference reward value of the working node. The global interference decision average model parameters are generated based on the interference reward value and the sub-model parameters. The parameters of the global interference decision model are updated based on the historical update impulse and the global interference decision average model parameters. The global interference decision model is optimized based on the updated global interference decision model parameters to obtain the global interference strategy.

2. The method according to claim 1, characterized in that, The central server includes a drone control system for broadcasting the parameters of the global interference decision model to several working nodes in the drone interference network architecture, collecting the sub-model parameters of the working nodes using a synchronous communication mechanism, generating the global interference decision average model parameters based on the average interference reward value of the sub-model parameters, and calculating the historical update impulse of the global interference decision average model parameters. The working node includes multiple interference drones, wherein the drones are used to input the collected interference resources as training samples into the sub-interference decision model. The sub-interference decision model uses a reinforcement learning strategy to train the training samples and updates the parameters of the sub-interference decision model to obtain the sub-interference strategy and the sub-model parameters of the sub-interference decision model.

3. The method according to claim 1, characterized in that, The UAV interference resources include: target UAV coordinate information, target UAV attitude information, image information of the geographical environment where the target UAV is located, and the target UAV's communication protocol.

4. The method according to claim 1, characterized in that, The training sample dataset is input into the sub-interference decision model of the working node for training to obtain a sub-interference strategy. The sub-interference decision model is iteratively updated according to the sub-interference strategy and a preset sub-sampling threshold to obtain sub-model parameters, including: The training sample dataset is input into the sub-interference decision model of the working node. The UAV interference environment uses the Markov decision method to obtain the target mission scheme. The target mission scheme at the current moment is stored in the experience replay library of the working node to obtain the historical target mission scheme. After the historical target task scheme and the global interference decision model parameters are trained by the sub-interference decision model, a sub-interference strategy is obtained. The sub-interference decision model is iteratively updated according to the sub-interference strategy and the preset sub-sampling threshold in the experience replay library to obtain the sub-model parameters of the working node.

5. The method according to claim 4, characterized in that, The central server performs performance testing on the sub-interference strategy corresponding to the working node based on the sub-model parameters to obtain the interference reward value of the working node, including: The central server performs performance tests on the sub-interference strategies corresponding to different working nodes according to a preset test count threshold and the sub-model parameters, and saves the interference reward value of the working node.

6. The method according to claim 5, characterized in that, The process involves calculating the interference reward value and the sub-model parameters to generate global interference decision average model parameters, updating the parameters of the global interference decision model based on historical update impulses and the global interference decision average model parameters, and optimizing the global interference decision model according to the updated parameters to obtain a global interference strategy, including: Calculate the mean of the interference reward values: in, The interference reward value, The preset threshold number of tests, This serves as the index identifier for the working node. The mean of the interference reward values; The mean of the interference reward value is weighted by the sub-model parameters of different working nodes to obtain the global interference decision average model parameters: in, The parameters of the global disturbance decision-making average model are... for The sub-model parameters of the iteration rounds, The number of the working nodes. The mean of the interference reward values, This serves as the index identifier for the working node; The parameter update amount of the global disturbance decision model is updated using a block update filtering algorithm based on the historical update impulse and the average parameters of the global disturbance decision model: in, This represents the parameter update amount of the global disturbance decision model from the previous iteration. This represents the parameter update amount of the global disturbance decision model in the current iteration round. This is the impulse coefficient for the current iteration round. The learning rate for the current iteration round. The global disturbance decision average model parameters for the current iteration round are... The parameters for the global disturbance decision in the previous iteration; The parameters of the global disturbance decision model in the current iteration are summed with the parameter update amount of the global disturbance decision model to obtain the parameters of the global disturbance decision model after the update in the current iteration: in, This represents the parameter update amount of the global disturbance decision model in the current iteration round. The global disturbance decision average model parameters for the current iteration round are... These are the parameters of the global disturbance decision model after the current iteration round update; The global interference decision model is optimized based on the updated parameters to obtain the global interference strategy.

7. The method according to claim 6, characterized in that, Obtain the training sample dataset, which includes: global interference decision model parameters and UAV interference resources, including: The parameters of the global interference decision model are obtained by controlling the drone, and the interference drone obtains the drone interference resources by exploring the environmental space.

8. A UAV autonomous interference decision-making device based on historical update impulse, characterized in that, The device includes: The training sample dataset acquisition module is used to acquire the training sample dataset, which includes: global interference decision model parameters and UAV interference resources; The sub-model parameter acquisition module is used to input the training sample dataset into the sub-interference decision model of the working node for training, obtain the sub-interference strategy, and iteratively update the sub-interference decision model according to the sub-interference strategy and the preset sub-sampling threshold to obtain the sub-model parameters. The interference reward value acquisition module is used to perform performance testing on the sub-interference strategy corresponding to the working node through the central server based on the sub-model parameters, and obtain the interference reward value of the working node. The interference decision module is used to generate global interference decision average model parameters based on the interference reward value and the sub-model parameters, update the parameters of the global interference decision model based on the historical update impulse and the global interference decision average model parameters, optimize the global interference decision model based on the updated global interference decision model parameters, and obtain a global interference strategy.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Federal learning-based ad hoc network frequency hopping anti-interference method and device special for unmanned aerial vehicle

    CN113890564A

  • Unmanned aerial vehicle small target detection method and system and storable medium

    CN114067225A