Physical layer authentication enhanced unmanned aerial vehicle federal learning method

By adopting a combination of physical layer authentication and LSTM-SDRL algorithms in drone federated learning, the security challenges of malicious node attacks in open wireless environments are solved, and efficient and reliable authentication of edge devices and multiple performance guarantees for federated learning are achieved.

CN119946630AActive Publication Date: 2025-05-06BEIHANG UNIV

Patent Information

Application Number
CN202510157709.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-06
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

In open wireless environments, federated learning has the security challenge of malicious nodes masquerading as qualified nodes, intercepting or attacking uploaded model parameters, especially how to perform efficient and reliable continuous authentication on resource-constrained edge devices is a key issue.

Method used

The drone federated learning method enhanced by physical layer authentication is used to determine whether the identity verification of the edge device is qualified through the binary false test model of signal-to-noise ratio and missed rate. If it is qualified, the updated local model will be received, and if it is not qualified, the reception will be refused. At the same time, the LSTM-SDRL algorithm is used to build a drone orchestration resource model to determine the orchestration strategies for safety performance, communication performance, energy-saving performance and model accuracy performance.

Benefits of technology

It realizes identification and isolation of malicious devices, ensures the security of federated learning, communication performance, energy-saving performance, and model accuracy, and reduces communication overhead to meet the efficient and reliable certification of edge devices by drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946630A_ABST
    Figure CN119946630A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of federated learning, and provides a physical layer authentication enhanced unmanned aerial vehicle federated learning method, which comprises the following steps: judging whether the identity verification of edge equipment is qualified or not through a communication model based on physical layer authentication enhancement, if so, receiving a federated learning local model updated by the edge equipment by an unmanned aerial vehicle, if the communication model is not qualified, the unmanned aerial vehicle refuses to receive the local model, and the communication model is a binary false test model adopting a signal-to-noise ratio and a missing report rate; and constructing an unmanned aerial vehicle arrangement resource model based on the missing report rate, and solving the unmanned aerial vehicle arrangement resource model through an LSTM-SDRL algorithm to generate an unmanned aerial vehicle arrangement strategy for determining the safety performance, the communication performance, the energy-saving performance and the model accuracy performance. According to the method, malicious equipment with unqualified identity authentication is identified through the physical layer, and the security, communication performance, energy-saving performance, model accuracy performance, high privacy and low communication overhead of federal learning are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning, and in particular to a federated learning method for unmanned aerial vehicles with enhanced physical layer authentication. Background Art

[0002] In order to promote the evolution of next-generation networks towards ubiquitous intelligence and endogenous security, using unmanned aerial vehicles (UAVs) to collect data from edge devices on the ground and extract valuable information based on machine learning algorithms has become a very promising method. However, the centralized training method of traditional machine learning has the risk of privacy leakage and network traffic congestion, making it impractical for edge devices to transmit raw data to UAVs.

[0003] Based on the paradigm of federated learning (FBL), edge devices can be allowed to collaborate with drones (UAVs) to train models by transmitting updated parameters instead of raw data to them, thereby significantly improving privacy protection capabilities and communication efficiency. However, in an open wireless environment, the FBL paradigm still faces security challenges. For example, malicious nodes may disguise themselves as qualified nodes and use various attacks to intercept or attack uploaded model parameters.

[0004] At present, most of the research on ensuring the security of federated learning focuses on various cryptography-based technologies, which usually have large computational overhead and are not suitable for widespread deployment in IoT devices. How to efficiently and reliably perform continuous authentication on resource-constrained edge devices during the training process of federated learning is a key issue.

[0005] Existing federated learning security technologies include authentication protocols, homomorphic encryption, blockchain, and differential privacy. These commonly used technologies to ensure the security of federated learning can usually only be used to achieve encryption and authentication between high-computing nodes and servers, and cannot be applied to access massive edge devices in open wireless environments with limited communication and computing resources.

[0006] In addition, existing physical layer authentication methods include fingerprint recognition: establishing a unique device fingerprint by analyzing the inherent RF characteristics of communication equipment, such as signal strength, phase and spectrum characteristics; channel response: utilizing the time-varying characteristics of the channel such as multipath propagation, delay spread, etc.; phase characteristics: achieving authentication by analyzing signal phase changes.

[0007] While the above technologies provide different approaches to achieve physical layer authentication, they are not applicable to the UAV environment, especially for establishing high-capacity secure links between mobile UAVs and ground edge devices.

[0008] Therefore, how drones can identify malicious devices through physical layer authentication technology to ensure the performance of federated learning is a technical problem that needs to be solved. Summary of the invention

[0009] To this end, the present invention provides a UAV federated learning method with enhanced physical layer authentication, which identifies malicious devices that fail identity authentication through the physical layer, thereby ensuring the security, communication performance, energy saving performance, model accuracy, high privacy and low communication overhead of federated learning.

[0010] To achieve the above object, the present invention proposes a UAV federated learning method with enhanced physical layer authentication, in which the UAV acts as a mobile server and performs federated learning with a ground edge device, including:

[0011] Determine whether the identity authentication of the edge device is qualified through a communication model based on physical layer authentication enhancement, if qualified, the drone receives the local model of federated learning updated by the edge device, if not qualified, the drone refuses to receive the local model, wherein the communication model is a binary false detection model using signal-to-noise ratio and false negative rate;

[0012] A drone scheduling resource model is constructed based on the false negative rate, and the drone scheduling resource model is solved by the LSTM-SDRL algorithm to generate a drone scheduling strategy for determining safety performance, communication performance, energy saving performance and model accuracy performance.

[0013] Further, the local optimization problem of the communication model enhanced by physical layer authentication aims to minimize the local loss function of the model change between the global model and the local model of the edge device in the time slot, and the local optimization problem includes a local change loss term, a local loss function gradient term and a global loss function gradient term;

[0014] The local change loss term is a local loss function with the sum of the global model and the model change as variables;

[0015] The local loss function gradient term includes the gradient of the local loss function in the global model;

[0016] The global loss function gradient term includes the gradient of the global loss function in the global model.

[0017] Furthermore, the communication model performs a binary hypothesis test based on the difference in signal-to-noise ratios of adjacent global rounds of federated learning to determine whether the identity authentication of the edge device is qualified;

[0018] The probability density function of the binary hypothesis test is calculated based on the signal-to-noise ratio, Rayleigh fading and channel deterministic variables.

[0019] Further, the channel deterministic variable is determined based on the reference channel power gain, the distance between the UAV and the edge device, the edge device transmit power, and the noise power at the UAV;

[0020] The probability density function is derived through the communication relationship between the signal-to-noise ratio, the Rayleigh fading and the channel deterministic variable.

[0021] Furthermore, the drone scheduling resource model includes an objective function and constraints;

[0022] The objective function is to minimize the safety learning cost function based on the UAV trajectory, the federated learning participation variables, and the delay energy consumption variables;

[0023] The constraint conditions include a comparison of a false negative rate and a security threat threshold, a federated learning minimum participating device constraint based on the federated learning participating variables, and a drone motion constraint based on the drone trajectory.

[0024] Furthermore, the false negative rate is determined by calculating a probability function based on a Rayleigh fading rate, a signal-to-noise ratio difference of a malicious edge device, and a channel deterministic variable.

[0025] Further, it is characterized in that the LSTM-SDRL algorithm includes an LSTM sub-algorithm and an SDRL sub-algorithm;

[0026] The LSTM sub-algorithm is used to strengthen the expected cumulative reward network and the expected degree of violation of the security constraint network to generate a strengthened expected cumulative reward network and a strengthened expected degree of violation of the security constraint network;

[0027] The SDRL sub-algorithm constructs a strategy function based on the enhanced expected cumulative reward network and the enhanced safety constraint violation expected degree network to solve it and generate the drone scheduling strategy.

[0028] Furthermore, the output layer hidden state of the LSTM sub-algorithm is connected to the SDRL sub-algorithm through a fully connected neural network.

[0029] Furthermore, a policy distribution function of the SDRL sub-algorithm is constructed according to the expected cumulative reward network, the expected degree of violation of the safety constraint network and the reward function.

[0030] Furthermore, the SDRL sub-algorithm constructs a reward function based on the negative penalty constraints of the security learning cost and the false negative rate.

[0031] In the above solution, the UAV can efficiently and reliably authenticate the edge devices and achieve the Pareto optimality of training accuracy, latency, and energy consumption, thus providing four-fold guarantees for improving the safety performance, communication performance, energy-saving performance, and model accuracy performance of UAVs in carrying out edge device data mining tasks.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] 1. Identify malicious devices that fail identity authentication through the physical layer to ensure the security, communication performance, energy saving performance, model accuracy, high privacy and low communication overhead of federated learning.

[0034] 2. It can meet the requirements of efficient and reliable authentication of edge devices by drones, achieve Pareto optimality in training accuracy, latency, and energy consumption, and provide quadruple guarantees for improving safety performance, communication performance, energy-saving performance, and model accuracy for drones to carry out edge device data mining tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of the general flow of the UAV federated learning method enhanced by physical layer authentication according to an embodiment of the present invention;

[0036] Figure 2 A schematic diagram of the structure of a UAV federated learning method enhanced by physical layer authentication according to an embodiment of the present invention;

[0037] Figure 3 The figure is a schematic diagram of an algorithm of a federated learning method for UAVs enhanced by physical layer authentication according to an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0039] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0040] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0041] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0042] like Figures 1 to 3 As shown, the present invention provides a UAV federated learning method with enhanced physical layer authentication, which identifies malicious devices that fail identity authentication through the physical layer, thereby ensuring the security, communication performance, energy saving performance, model accuracy, high privacy and low communication overhead of federated learning.

[0043] like Figures 1 to 3 As shown, this embodiment proposes a UAV federated learning method with enhanced physical layer authentication, where the UAV acts as a mobile server and performs federated learning with a ground edge device, including:

[0044] Determine whether the identity authentication of the edge device is qualified through a communication model based on physical layer authentication enhancement, if qualified, the drone receives the local model of federated learning updated by the edge device, if not qualified, the drone refuses to receive the local model, wherein the communication model is a binary false detection model using signal-to-noise ratio and false negative rate;

[0045] A drone scheduling resource model is constructed based on the false negative rate, and the drone scheduling resource model is solved by the LSTM-SDRL algorithm to generate a drone scheduling strategy for determining safety performance, communication performance, energy saving performance and model accuracy performance.

[0046] It should be noted that federated learning is a distributed machine learning method that trains models on local devices and only shares model parameters rather than raw data, thereby achieving multi-party collaborative training while protecting data privacy and security. Physical layer authentication is a technology that uses the random characteristics of wireless channels and device hardware features to achieve identity authentication. It distinguishes qualified users from illegal intruders by analyzing physical layer information such as channel characteristics or radio frequency fingerprints, thereby enhancing the security of the communication system. Compared with cryptography-based authentication technology, it has more advantages on resource-constrained IoT devices.

[0047] Specifically, Figure 2 As shown, the drone u stays in the urban environment and acts as a mobile parameter server for the federated learning (FLT) task, co-training with the ground edge device (ED), denoted as In the formula, represents all qualified edge devices, i represents qualified edge devices that have passed identity authentication, and I represents the total number of qualified edge devices.

[0048] However, due to the openness of the wireless environment, malicious devices can disguise themselves as qualified devices and hinder the task through various attacks such as backdoor attacks and poisoning attacks, which are expressed as In the formula, J represents all malicious edge devices, and j represents illegal edge devices / malicious edge devices that fail identity authentication.

[0049] Specifically, the backbone steps of the federated learning include: defining the private data set of each edge device i as Where x and y represent the feature vector and label of each training data respectively. The edge device performs local training of federated learning by minimizing its local loss function. The local loss function is:

[0050]

[0051] Where, L i (w g ) is the local loss function of the edge device, w g represents the global model, D i represents a private dataset, x i,n ,y i,n are the feature vector and label of the training data of the i-th edge device, and l() is the local loss function of the n-th private dataset of the i-th edge device.

[0052] Upload the local model to the global model. The training of the global model can be expressed as:

[0053]

[0054] In the formula, w g represents the global model, L i (w g ) represents the local loss function of the i-th edge device, D i represents a private data set, and D represents the entire data set.

[0055] in, It indicates that the entire data set is the entire private data set of all qualified devices that have passed the identity authentication.

[0056] Furthermore, the local optimization problem of the communication model enhanced by physical layer authentication aims to minimize the local loss function of the model change between the global model and the local model of the edge device in the time slot, and the local optimization problem includes a local change loss term, a local loss function gradient term and a global loss function gradient term; the local change loss term is a local loss function with the sum of the global model and the model change as variables; the local loss function gradient term includes the gradient of the local loss function in the global model; the global loss function gradient term includes the gradient of the global loss function in the global model.

[0057] Specifically, the algorithm flow of this embodiment is generally as follows: Step S1, start: each qualified edge device reports its location and transmits a pilot signal to the drone for registration and system startup.

[0058] Step S2, local training and model upload: Each participating edge device trains and uploads a local model based on a private dataset to solve the following local optimization problems for physical layer certification:

[0059]

[0060] In the formula, G i (w g,t ,h i,t ) is the authentication performance loss function, h i,t is the change between the global model and the local model of the ith edge device in the tth time slot, which serves as a local adjustment term, w g,t represents the global model of the tth time slot, L i (w g,t +h i,t ) is the local change loss term, which represents the local loss function of the i-th edge device, which depends on w g,t and h i,t ; is the local loss function gradient term, which represents the local loss function L i In w g,t gradient; is the global loss function gradient term, which represents the gradient of the global loss function L. Therefore, a local optimization problem of physical layer authentication is defined with the global model and local model of federated learning as variables. i,t is the change between the global and local models of the edge device in the tth time slot, w g,t represents the global model of the tth time slot, L i () is the local loss function of the edge device, δ represents the update step size, G i Represents the local optimization problem at the edge device.

[0061] Step S3, UAV-assisted physical layer authentication (PLA), the UAV authenticates each updated parameter to distinguish unauthorized Byzantine nodes from qualified participating edge devices, and the physical layer authentication is based on wireless characteristics, such as SNR difference. The detailed method is described in detail in Physical layer authentication based on signal-to-noise ratio difference.

[0062] Step S4, edge device selection enhanced by physical layer authentication: The drone makes a judgment based on the physical layer authentication, and refuses to receive models from nodes that are judged to be malicious nodes or a few easily confused qualified nodes. The security device selection set is defined as in represents all qualified edge devices, K t represents all binary choice variables s i.t , k i.t =1 indicates that device i is determined to be a qualified edge device in time slot t.

[0063] Step 5, global aggregation and distribution: After collecting local models from all selected edge devices, the drone server merges them into a new global model iteration and uses a specific aggregation method. Subsequently, the drone broadcasts the updated global model w of round t. g,t And distributed to all qualified edge devices to facilitate local model optimization in the subsequent t+1 round, which can be expressed as:

[0064]

[0065] In the formula, w g,t represents the global model of the tth round, h i,t is the change between the global and local models of the edge device in the tth time slot, η represents the local model accuracy, represents the local optimal gradient. This formula represents h i,t The conditions for achieving local accuracy η.

[0066] Therefore, once a given global accuracy ∈ is reached, the whole process can be terminated, i.e.:

[0067]

[0068] In the formula, L(w g ) represents the global loss function of the tth round, ∈ represents the global accuracy, represents the global optimal model.

[0069] In this embodiment, both the drone and the edge device are equipped with a single antenna, and it can be easily extended to a multi-antenna configuration by applying physical layer authentication in multiple-input and multiple-output MIMO communications.

[0070] Furthermore, the communication model performs a binary hypothesis test based on the difference in signal-to-noise ratios of adjacent global rounds of federated learning to determine whether the identity authentication of the edge device is qualified; wherein the probability density function of the binary hypothesis test is calculated based on the signal-to-noise ratio, Rayleigh fading and channel deterministic variables.

[0071] Furthermore, the channel deterministic variable is determined based on a reference channel power gain, a distance between the UAV and the edge device, a transmit power of the edge device, and a noise power at the UAV.

[0072] The probability density function is derived through the communication relationship between the signal-to-noise ratio, the Rayleigh fading and the channel deterministic variable, and the communication relationship is specifically:

[0073] The latency of key performance of federated learning is affected by the communication model. In the drone scenario, due to the continuous change of motion trajectory, the communication model directly determines the latency of federated learning. In the t+1th time slot / round of federated learning, the channel gain between the i-th qualified device and the drone is The channel gain can be expressed as

[0074]

[0075] Where α represents the reference channel power gain and the value when the preferred distance is 1 meter. represents the distance between the drone and the i-th qualified edge device, represents the Rayleigh fading between the UAV and the i-th qualified edge device, and β is the air-to-ground path loss exponent, which is usually between 2.5 and 4, depending on the environment.

[0076] in, Represents the distance calculation process, where: represents the distance between the drone and the i-th qualified edge device, Indicates the position of the drone. represents the location of the i-th qualified edge device. The location of the edge device is preferably defined as That is, the position of the edge device preferably includes an x-axis position, a y-axis position, and a z-axis position.

[0077] The Rayleigh fading follows an exponential distribution, and the rate of Rayleigh fading is given by λ i Characterization, that is, the function value of Rayleigh fading Follow this formula: Therefore, the signal-to-noise ratio (SNR) received by the i-th qualified edge device at the drone can be expressed as:

[0078]

[0079] in, is the signal-to-noise ratio, is the channel deterministic variable, where α represents the reference channel power gain, preferably at a distance of 1 meter, represents the distance between the drone and the i-th qualified edge device, p i,t represents the transmission power of device i in round t, Represents the noise power at the drone.

[0080] Similarly, the signal-to-noise ratio received by malicious device j at the drone can also be represented by the same signal-to-noise ratio formula, denoted as

[0081] The bandwidth from qualified edge device i to the drone is b iThe communication data rate can be expressed as:

[0082]

[0083] In the formula, Indicates the communication data rate, b i represents bandwidth, is the channel deterministic variable, represents the Rayleigh fading between the UAV and the i-th qualified edge device.

[0084] Specifically, the physical layer authentication model based on the signal-to-noise ratio difference: In order to determine the signal source, the drone deploys a physical layer authentication module. The drone can use the signal-to-noise ratio difference between the tth time slot and the t+1th time slot (in federated learning, it can be regarded as the global t round and the global t+1 round), which are respectively expressed as the above Γ t and Γ t+1 Assume that the edge device for the tth transmission has been authenticated as a qualified device, that is, In the formula represents all qualified edge devices, based on the observed Γ t+1 , a binary hypothesis test is formulated to determine whether the source of the t+1th transmission is a qualified device. The binary hypothesis test is specifically:

[0085]

[0086] Among them, the null hypothesis H0 indicates that the parameters transmitted in the t+1th time slot come from the i-th qualified edge device, while the alternative hypothesis H1 indicates that the parameters transmitted in the t+1th time slot come from the Byzantine device (malicious device) j, and Γ is the signal-to-noise ratio.

[0087] Therefore, the binary hypothesis test of the drone on the edge device is divided into two cases. The probability density function of the binary false test includes the probability density function of qualified edge devices and the probability density function of unqualified edge devices. The derivation process of the probability density function of qualified edge devices is:

[0088] In the t+1 round of global training, if the parameters are uploaded by a qualified edge device, the SNR difference between adjacent transmissions can be re-derived as follows:

[0089] Where v is the signal-to-noise ratio difference of qualified edge devices, are the channel deterministic variables of the t+1th round and the tth round, They represent the Rayleigh fading of the t+1th round and the tth round between the UAV and the i-th legitimate edge device respectively.

[0090] The signal-to-noise ratio difference v of the legal edge device is used as a random variable for binary hypothesis testing. The probability density function (PDF) of the binary hypothesis test of the qualified edge device can be derived as follows:

[0091]

[0092] The proof process is as follows: According to Rayleigh fading and We can re-express the distribution of v = g1-g0, where and Substituting it into the probability density function (PDF) of the binary hypothesis test, we can get the probability density functions of the two parts:

[0093]

[0094] Therefore, we can derive the following proof formula for the probability density function of qualified edge devices:

[0095]

[0096] In the above proof formulas, represents the Rayleigh fading between the drone and the i-th legitimate edge device, f G represents the probability density function of the two parts, f V (v) represents the overall qualified probability density function, λ i represents the rate of Rayleigh fading, v is the difference in signal-to-noise ratio of qualified edge devices, are the channel deterministic variables of the t+1th round and the tth round, respectively, and g1, g0 are the re-expressions of the signal-to-noise ratio differences of qualified edge devices.

[0097] By proving the probability density function of qualified edge devices, the probability density function of the qualified edge devices can be obtained, and the proof is complete.

[0098] The derivation process of the probability density function of unqualified edge devices is:

[0099] Considering that if the model parameters are uploaded by a malicious node, the difference in the signal-to-noise ratio of the malicious edge device between adjacent transmissions is expressed as:

[0100]

[0101] Where w is the signal-to-noise ratio difference of the malicious edge device, are the channel deterministic variables of the t+1th round and the tth round, They represent the Rayleigh fading of the t+1th round and the tth round between the UAV and the i-th legitimate edge device respectively.

[0102] The signal-to-noise ratio difference w of the malicious edge device is used as a random variable for binary hypothesis testing. The probability density function (PDF) of the binary hypothesis test of the malicious edge device can be derived as follows:

[0103]

[0104] The proof process is as follows: The distribution of The distribution of can be expressed as Since w=g2-g1, and Malicious edge device probability density function proof formula:

[0105]

[0106] In the above proof formulas, They represent the Rayleigh fading of the t+1th round and the tth round between the UAV and the i-th legitimate edge device respectively.

[0107] Malicious edge device signal-to-noise ratio difference w,λ i λ j denote the Rayleigh fading rates of qualified edge devices and malicious edge devices, respectively. are the channel deterministic variables in the t+1th round and the tth round, respectively, and g1, g0 are the re-expressions of the signal-to-noise ratio differences of the malicious edge devices.

[0108] By proving the probability density function formula of malicious edge devices, we can obtain the probability density function of the binary hypothesis test of malicious edge devices. The proof is complete.

[0109] Based on the received signal-to-noise ratio of two consecutive transmissions, the statistic for the binary hypothesis test is defined as follows:

[0110]

[0111] Among them, t is the optimal detection threshold, which is used to decide whether the malicious device participates in the transmission parameters in the t+1th time slot. D1 and D2 represent the decision thresholds supporting H0 and H1, respectively. ΔΓ represents the statistic, and Γ is the signal-to-noise ratio.

[0112] Through the probability density functions of qualified edge devices and malicious edge devices, we can further deduce the false alarm rate and missed alarm rate of drones during physical layer authentication and derive the optimal detection threshold.

[0113] The false alarm rate (FAR) refers to the probability that the local model uploaded by a qualified edge device fails to pass the drone aggregator certification, which can be derived by the following formula:

[0114]

[0115] In the formula, is the false alarm rate, Indicates that the drone is considered a malicious device. represents the null hypothesis, i.e., the drone is determined to be a qualified edge device, λ i represents the rate of Rayleigh fading, v is the difference in signal-to-noise ratio of qualified edge devices, and are the channel deterministic variables of the tth global round and the t+1th global round, Π t is the detection threshold.

[0116] In addition, the number of qualified edge devices affects the convergence and performance of federated learning, and some unselected qualified edge devices may be misidentified as malicious edge devices due to the high false positive rate. The number of misidentifications can be calculated by the following formula:

[0117]

[0118] In the formula, Indicates the number of qualified edge devices that are misidentified as malicious edge devices, is the false alarm rate, J represents all malicious edge devices, j represents a malicious edge device, where s t,i The edge device i is identified as a qualified device.

[0119] The false negative rate is determined by calculating the probability function based on the Rayleigh fading rate, the difference in signal-to-noise ratio of the malicious edge device, and the channel deterministic variable, specifically:

[0120] The Missing Report Rate (MDR) refers to the probability that an illegal malicious device disguises itself as a qualified device, passes authentication, and participates in federated learning. This is a more serious situation, and the corresponding probability of occurrence can be derived as follows:

[0121]

[0122] In the formula, represents the false negative rate, λ i λ j They represent the Rayleigh fading rates of qualified edge devices and malicious edge devices, w represents the signal-to-noise ratio difference of malicious edge devices, and are the channel deterministic variables of the tth global round and the t+1th global round, Π t is the detection threshold.

[0123] In reality, malicious edge devices will try to maximize their false negative rate. To increase the probability of intrusion, qualified edge devices hope that the maximum value of the potential false negative rate meets the following constraints:

[0124]

[0125] In the formula, ∈ represents the maximum possible intrinsic security threat level of system intrusion, Indicates the maximum potential false negative rate of qualified edge devices.

[0126] The process of determining the optimal detection threshold is as follows: In each global round t, the drone needs to determine the optimal detection threshold to minimize the false alarm rate and missed alarm rate. The formula for minimizing the false alarm rate and missed alarm rate is as follows:

[0127]

[0128] In the formula, represents the optimal detection threshold, Π t is the detection threshold, σ t represents the maximum potential false negative rate of qualified edge devices, and ∈ represents the intrinsic security threat level with the maximum possibility of system intrusion.

[0129] The proof process is: It can be deduced that the false negative rate is about the detection threshold Π t is a monotonically increasing function of the detection threshold Π t The derivation process of the monotone decreasing function is:

[0130]

[0131] and,

[0132]

[0133] In the formula, λ i λ j denote the Rayleigh fading rates of qualified edge devices and malicious edge devices, respectively. and are the channel deterministic variables of the t-th global round and the t+1-th global round respectively, and Π is the detection threshold.

[0134] Therefore, the reduction of the missed alarm rate may increase the false alarm rate. The optimal detection threshold can be obtained by the formula of minimizing the false alarm rate and the missed alarm rate, and the optimal solution can be obtained when the balance point is reached.

[0135] Furthermore, the drone scheduling resource model includes an objective function and constraints; the objective function is to minimize the security learning cost function based on the drone trajectory, federated learning participation variables, and delay energy consumption variables; the constraints include a comparison of the false alarm rate with the security threat threshold, a federated learning minimum participating device constraint based on the federated learning participation variables, and a drone motion constraint based on the drone trajectory. The delay energy consumption variables include bandwidth allocation, computing frequency, and transmission power.

[0136] Further, it is characterized in that the LSTM-SDRL algorithm includes an LSTM sub-algorithm and an SDRL sub-algorithm;

[0137] The LSTM sub-algorithm is used to strengthen the expected cumulative reward network and the expected degree of violation of the security constraint network to generate a strengthened expected cumulative reward network and a strengthened expected degree of violation of the security constraint network;

[0138] The SDRL sub-algorithm constructs a strategy function based on the enhanced expected cumulative reward network and the enhanced safety constraint violation expected degree network to solve it and generate the drone scheduling strategy.

[0139] Furthermore, the output layer hidden state of the LSTM sub-algorithm is connected to the SDRL sub-algorithm through a fully connected neural network.

[0140] Furthermore, a policy distribution function of the SDRL sub-algorithm is constructed according to the expected cumulative reward network, the expected degree of violation of the safety constraint network and the reward function.

[0141] Furthermore, the SDRL sub-algorithm constructs a reward function based on the negative penalty constraints of the security learning cost and the false negative rate.

[0142] In the above solution, the UAV can efficiently and reliably authenticate the edge devices and achieve the Pareto optimality of training accuracy, latency, and energy consumption, thus providing four-fold guarantees for improving the safety performance, communication performance, energy-saving performance, and model accuracy performance of UAVs in carrying out edge device data mining tasks.

[0143] Latency and energy consumption analysis: The latency and energy consumption of each edge device mainly involve the local computing stage and the transmission stage in the federated learning training process.

[0144] Local computing phase: At time slot t, the local training latency of the i-th edge device can be expressed as:

[0145]

[0146] Among them, C i,t represents the computational complexity of training a sample data using the back propagation algorithm on the i-th edge device, fi,t Indicates the CPU frequency of the i-th device, represents the local training delay, D i,t represents the dataset size of edge device i.

[0147] In the local computing stage, the energy consumption can be expressed as:

[0148]

[0149] in, represents energy consumption, f i,t represents the CPU frequency of the ith device, ζ i is the calculation coefficient related to the CPU of the i-th device, C i,t represents the computational complexity of training a sample data using the back propagation algorithm on the i-th edge device, Indicates the number of local training rounds.

[0150] Transmission phase: Given the fixed dimensions of the global model and the local model, the data size of the local model parameters that each edge device must send to the drone in each iteration is constant. Therefore, the latency of the transmission phase can be expressed as:

[0151]

[0152] In the formula, represents the delay of the local transmission stage, represents the communication data rate, and Ω represents the size of the model.

[0153] Accordingly, the energy consumption of the i-th device in the transmission phase can be expressed as:

[0154]

[0155] In the formula, Indicates energy consumption, represents the delay of the local transmission stage, p i,t represents the transmit power of edge device i.

[0156] The energy consumption of the drone includes motion and parameter aggregation. Since the latency is relatively low compared to transmission and local computation, the energy consumption during aggregation can be ignored. The motion power consumption can be calculated as follows:

[0157]

[0158] This formula can accurately reflect the complex change trend of UAV flight power by introducing key physical quantities in flight, such as flight distance, time, rotor speed, air resistance, and static power consumption of the UAV. represents the motion power consumption, P0 and P1 represent the blade profile power and thrust power of the drone in static hovering, respectively, r t , r t-1 denote the positions of the drones in the tth round and the t-1th round, τ t represents the total delay, U tip is the tip speed of the blade, v0 represents the average rotor induced speed in hover, d0, ρ0, s0, and A correspond to the airframe drag ratio, air density, rotor solidity, and rotor disk area, respectively.

[0159] In summary, the total delay of a global round can be expressed as follows:

[0160]

[0161] In the formula, τ t represents the total delay, They represent the delays of the local calculation stage and the transmission stage respectively.

[0162] In addition, the overall energy consumption of all devices in the tth global round can be expressed as:

[0163]

[0164] In the formula, E t represents the overall energy consumption, They represent the overall energy consumption of the local computation phase and the transmission phase, respectively.

[0165] In order to meet the various requirements of different federated learning, such as latency sensitivity or energy sensitivity, we use the training cost TC to strike a balance between latency and energy, namely:

[0166]

[0167] Where TC t is the training cost, τ t is the total delay, represents the weight coefficient, E t Indicates overall energy consumption.

[0168] At the same time, the system proposed in this embodiment needs to reduce the false alarm rate and maximize the number of security devices under the false alarm rate limit. To this end, a new metric is redefined, namely, the security learning cost (SLC):

[0169]

[0170] Among them, t represents the security learning cost, τ t is the total delay, Represents the weight coefficient, Et represents the total energy consumption, the denominator represents the participation degree, where i represents the qualified edge device, I represents the total number of qualified edge devices, and k t,i represents the federated learning participating variable, j represents the malicious edge device, J represents all malicious edge devices, represents the false negative rate. This metric takes into account the actual number of participants and the level of false positives.

[0171] Optimization problem of drones: After establishing the above models, from the perspective of drones, its mission goal is to ensure that the system intrusion is constrained within a certain range, while minimizing latency and energy consumption, and ensuring the participation of qualified devices. Our optimization problem and constraints are as follows:

[0172]

[0173] Where P1 is the objective function for minimizing the resources (computing resources and communication resources) of drone scheduling, and C1 to C7 are seven constraints. The objective function includes minimizing latency and energy consumption, and maximizing device participation, where: k t ,b t ,f t ,p t They are drone trajectory, federated learning participating variables, bandwidth allocation, computing frequency and transmission power. Represents the trajectory design of the drone and the variables involved in federated learning is a binary variable indicating whether the i-th qualified edge device is selected in the t-th round and the bandwidth is allocated Calculate frequency Transmission power

[0174] Constraint (C1) ensures the minimum detection rate (MDR), i.e., the intrusion threat to the system. Constraint (C2) is the minimum participating device constraint of federated learning based on the participating variables of the federated learning, which ensures the minimum participating devices. Constraint (C3) restricts the physical motion law of the drone, and constraints (C4) to (C6) are drone motion constraints based on the trajectory of the drone, which restrict wireless resources. Finally, constraint (C7) ensures that the energy consumption of the drone does not exceed the energy budget.

[0175] LSTM-SDRL algorithm: Optimizing the objective function P1 is a non-convex multi-step decision problem that requires online state information from the environment. In addition, due to the intrusion of potential malicious edge devices, the drone agent needs to avoid dangerous strategies that may interrupt the federated learning task while exploring the action space. Therefore, the problem is solved by the LSTM-SDRL method. The objective function P1 is transformed into a constrained Markov decision process (CMDP) and the secure deep reinforcement learning (SDRL) environment is clarified. Next, the state space, action space, reward, and cost in the model are described.

[0176] State space: In the tth round of global federated learning, the state s t ∈S contains the following information: the location of the drone in round t-1, the location of qualified edge devices and malicious edge devices r I,t-1 ,r J,t-1 , the power p transmitted by the malicious edge device in round t J,t , dataset size D t , computational complexity C t In addition, it also includes the remaining energy E of the drone in round t u,t , the state space can be summarized as follows:

[0177] s t = {r u,t-1 ,r I,t-1 ,r J,t-1 ,p J,t ,C t ,D t ,E u,t}.

[0178] In the formula, state s t represents the state space, r u,t-1 ,r I,t-1 ,r J,t-1 ,p J,t ,C t ,D t ,E u,t denote the locations of qualified edge devices and malicious edge devices, transmission power, constraints, dataset size, and the remaining energy E of the drone in round t. u,t .

[0179] Action Space: After observing its state, the drone will react, and its action space is defined as:

[0180] a t = {r u,t ,c t ,p t ,f t ,b t}.

[0181] In the formula, a t represents the action space, r u,t ,c t ,p t ,f t ,b t They represent drone trajectory, federated learning participating variables, transmission power, computing frequency, and bandwidth allocation respectively.

[0182] Reward function: According to the objective function and constraints of the optimization problem, the reward designed in this embodiment is to minimize the security learning cost of the entire federated learning (FL). The reward can be designed as:

[0183]

[0184] Among them, -Ψ is the negative value of the safety learning cost, and It represents the negative penalty constraint when the missed detection rate (MDR) exceeds the threshold, the UAV violates the motion law and the energy budget according to the constraints. All three penalties are scaled by multiplying the weights.

[0185] Cost: According to the definition of safety reinforcement learning, the cost designed in this embodiment is based on reducing the false detection rate to evaluate the risk value of the strategy to avoid dangerous exploration. Specifically, the safety constraints selected Quantified into M risk levels.

[0186]

[0187] Algorithm framework: The proposed LSTM-SDRL framework includes an expected cumulative reward network Q and an expected degree of safety constraint violation network E, hereinafter referred to as Q network and E network. Both Q network and E network consist of an online network and a target network. The Q network is used to estimate the state-action value function Q(s,a), which represents the expected cumulative reward under a given state s and action a. The E network is responsible for estimating the risk or penalty associated with potential safety constraint violations. Therefore, the function E(s,a) represents the expected degree of safety constraint violation. The target network is a copy of the online network and is used to improve training efficiency and stability. In addition, historical data is stored in a replay buffer. Subsequently, a subset of these data is randomly selected from the buffer to train the neural network, thereby significantly reducing the correlation between data points.

[0188] LSTM module: To capture the temporal patterns of environmental dynamics (e.g., time-varying federated learning tasks and observed malicious power transmitted by Byzantine nodes), this paper introduces an LSTM network, a recurrent neural network designed specifically for processing and predicting time series data. The input of the LSTM is the historical state sequence observed in time slot t and the previous seq-1 time slots. The output of the LSTM module is the hidden state h t , the state will be passed to the secure deep reinforcement learning part through the fully connected neural network (dense layer). Specifically, the LSTM module consists of a forget gate f t , input gate i t , output gate o t , unit door t and cell status C t Composition can be expressed as:

[0189] i t =σ(W i [s t ,h t-1 ]+b i ),

[0190] f t =σ(W f ·[s t ,h t-1 ]+b f ),

[0191] g t =tanh(W c ·[s t ,h t-1 ]+b c )

[0192] C t =f t *C t-1 +i t *g t ,

[0193] o t =σ(W o ·[s t ,h t-1 ]+b o )

[0194] h t =o t *tanh(C t ).

[0195] SDRL algorithm framework: The structure of the neural network of SDRL in this embodiment mainly includes Q network, target Q network, E network, and target E network. The output value of the Q network evaluates the long-term expected reward of each state-action pair, and the target Q value is calculated as follows:

[0196]

[0197] The learnable weight ω measures the importance of future rewards, which are evaluated by the target Q-network.

[0198] The E value describes the risk of each state-action pair, which is updated according to the reward function formula based on the future L risk values, as shown below:

[0199]

[0200] Among them, the discount rate χ represents the expected future risk value.

[0201] Different from the traditional Boltzmann distribution, the policy distribution of state-action pairs in secure deep reinforcement learning is based not only on Q-values ​​but also on E-values, which can be obtained by the following formula:

[0202]

[0203] Where Υ weighs the importance of long-term expected rewards, π(s t ,a t ) is the strategy distribution function.

[0204] The training process can be found in Figure 3 : First, the LSTM-enhanced Q network and E network are initialized, parameterized by parameters θ and φ respectively (line 1). Then, by copying the above online network, the target Q network and E network are initialized (line 2). In addition, the replay buffer D is initialized for storing historical samples (line 3).

[0205] In the exploration phase, the agent observes the state s t And build the state sequence (Line 6). Then, the Q network and E network output the evaluation of the Q value and E value of each state-action pair (Lines 7-8). Next, the agent selects an action according to formula (40) (Line 9). After executing the action, the agent enters the next state s i+1 and obtain a reward r from the environment i (Line 10). At the same time, in SDRL, the agent also evaluates the risk c t (Line 11). Afterwards, the sample data quintuple It will be saved in the playback buffer D for subsequent retrieval.

[0206] In the training phase (lines 14-21), a mini-batch of transformations of size H is randomly extracted from D to train the Q network and the E network. In addition, the target Q network and the E network are updated every set steps, and the update method is soft update.

[0207] In summary, a UAV-assisted physical layer authentication enhanced federated learning method is designed. This method achieves the Pareto optimality of training accuracy, latency, and energy consumption under the premise of ensuring efficient and reliable authentication of the server (UAV) to the client (edge ​​device), and improves the security performance, communication performance, energy saving performance, and model accuracy for UAVs to carry out edge device data mining tasks.

[0208] In this embodiment, malicious devices that fail identity authentication are identified through the physical layer to ensure the security, communication performance, energy-saving performance, model accuracy, high privacy and low communication overhead of federated learning. It satisfies the requirements of efficient and reliable authentication of edge devices by drones, achieves Pareto optimality of training accuracy, latency and energy consumption, and improves the security performance, communication performance, energy-saving performance and model accuracy performance of drones in edge device data mining tasks.

[0209] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0210] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A UAV federated learning method with enhanced physical layer authentication, characterized in that: The drone acts as a mobile server and performs federated learning with the edge devices on the ground, including: Determine whether the identity authentication of the edge device is qualified through a communication model based on physical layer authentication enhancement, if qualified, the drone receives the local model of federated learning updated by the edge device, if not qualified, the drone refuses to receive the local model, wherein the communication model is a binary false detection model using signal-to-noise ratio and false negative rate; A drone scheduling resource model is constructed based on the false negative rate, and the drone scheduling resource model is solved by the LSTM-SDRL algorithm to generate a drone scheduling strategy for determining safety performance, communication performance, energy saving performance and model accuracy performance.

2. The UAV federated learning method with enhanced physical layer authentication according to claim 1, characterized in that: The local optimization problem of the communication model enhanced by physical layer authentication aims to minimize the local loss function of the model change between the global model and the local model of the edge device in the time slot, and the local optimization problem includes a local change loss term, a local loss function gradient term and a global loss function gradient term; The local change loss term is a local loss function with the sum of the global model and the model change as variables; The local loss function gradient term includes the gradient of the local loss function in the global model; The global loss function gradient term includes the gradient of the global loss function in the global model.

3. The UAV federated learning method with enhanced physical layer authentication according to claim 1, characterized in that: The communication model performs a binary hypothesis test according to the difference in signal-to-noise ratios of adjacent global rounds of federated learning to determine whether the identity authentication of the edge device is qualified; The probability density function of the binary hypothesis test is calculated based on Rayleigh fading and channel deterministic variables.

4. The UAV federated learning method with enhanced physical layer authentication according to claim 3 is characterized in that: The channel deterministic variable is determined according to the reference channel power gain, the distance between the UAV and the edge device, the edge device transmit power and the noise power at the UAV; The probability density function is derived through the communication relationship between the signal-to-noise ratio, the Rayleigh fading and the channel deterministic variable.

5. The UAV federated learning method with enhanced physical layer authentication according to claim 1, characterized in that: The drone scheduling resource model includes an objective function and constraints; The objective function is to minimize the safety learning cost function based on the UAV trajectory, the federated learning participation variables, and the delay energy consumption variables; The constraint conditions include a comparison of a false negative rate and a security threat threshold, a federated learning minimum participating device constraint based on the federated learning participating variables, and a drone motion constraint based on the drone trajectory.

6. The UAV federated learning method with enhanced physical layer authentication according to claim 5, characterized in that: The false negative rate is determined by calculating a probability function based on the Rayleigh fading rate, the signal-to-noise ratio difference of the malicious edge device and the channel deterministic variable.

7. The UAV federated learning method with enhanced physical layer authentication according to any one of claims 1 to 6, characterized in that: The LSTM-SDRL algorithm includes an LSTM sub-algorithm and an SDRL sub-algorithm; The LSTM sub-algorithm is used to strengthen the expected cumulative reward network and the expected degree of violation of the security constraint network to generate a strengthened expected cumulative reward network and a strengthened expected degree of violation of the security constraint network; The SDRL sub-algorithm constructs a strategy function based on the enhanced expected cumulative reward network and the enhanced safety constraint violation expected degree network to solve it and generate the drone scheduling strategy.

8. The UAV federated learning method with enhanced physical layer authentication according to claim 7, characterized in that: The output layer hidden state of the LSTM sub-algorithm is connected to the SDRL sub-algorithm through a fully connected neural network.

9. The UAV federated learning method with enhanced physical layer authentication according to claim 7, characterized in that: The policy distribution function of the SDRL sub-algorithm is constructed according to the expected cumulative reward network, the expected degree of violation of the safety constraint network and the reward function.

10. The UAV federated learning method with enhanced physical layer authentication according to claim 7, characterized in that: The SDRL sub-algorithm constructs a reward function based on the negative penalty constraints of the security learning cost and the false negative rate.

Citation Information

Patent Citations

  • Federal learning-based reliability optimization method for digital twinning-assisted industrial Internet of Things

    CN115310360A

  • Full-stage credibility guarantee method for federated learning training of AIGC model

    CN118316623A

  • Differential privacy federated learning-based satellite-ground network edge computing task unloading method

    CN118660316A

  • Microgrid spatial-temporal perception energy management method based on safe deep reinforcement learning

    US20240330396A1

  • Artificial intelligence-assisted diagnosis model construction system for medical images

    WO2022222458A1

Cited By

  • Unmanned aerial vehicle flight authority management method and platform based on block chain technology

    CN120472720A

  • Unmanned aerial vehicle network intrusion detection method based on reinforcement learning

    CN120640294A