Low-overhead unmanned aerial vehicle identity authentication method based on deep reinforcement learning
Through deep reinforcement learning, optimized drone identity authentication strategies has been solved, and the problem of difficult to balance overhead and security in the existing technology has been achieved, and the identity authentication effect with low overhead and high security has been achieved.
Patent Information
- Application Number
- CN202411922893.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
In the face of harsh communication environments, existing drone identity authentication technology is difficult to balance overhead with security, which may lead to serious privacy leakage or excessive overhead.
Using a method based on deep reinforcement learning, the system state vector is constructed and the Q network optimization algorithm is combined with the UAV’s identity authentication strategy, including identity authentication method, session time and data encryption strategy, thereby optimizing authentication delay and energy consumption.
It realizes low overhead and high security in the drone identity authentication process, avoids high-risk strategies, and improves the efficiency and security of identity authentication.
Smart Images

Figure BDA0005208471380000048 
Figure BDA0005208471380000066 
Figure FDA0005208471370000016
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a low-overhead drone identity authentication method based on deep reinforcement learning. Background Art
[0002] Unmanned Aerial Vehicle (UAV) is widely used in military, industrial and civilian fields due to its advantages such as easy deployment and low maintenance cost. It brings great convenience to people in monitoring, detection, transportation, emergency rescue and other aspects. However, UAVs usually work in harsh natural environments, and their communication security faces various challenges. An untrusted communication environment can lead to risks such as valuable data leakage or loss of important goods carried by drones. Therefore, it is essential to provide a secure communication channel, in which identity authentication plays an important role.
[0003] There are a lot of related research on drone identity authentication at home and abroad. For example, the Chinese patent application publication number CN116506857A discloses a drone identity authentication method based on timely generation and updating of authentication information. Among them, the physical layer unclonable function (PUF) is used to generate unique identity information for each drone, which improves the reliability of drone identity authentication. In 2024, Chuang Tian et al. provided a unique identity for unmanned IoT devices through the physical unclonable function PUF, and used its challenge-response mechanism and the key generation function of the ground control station to generate session keys for the participants, ensuring the security of communication. In 2024, Raja Karmakar et al. used a fuzzy extractor to eliminate noise in the PUF output on the basis of verifying the identity of the drone using the uniqueness of PUF, further improving the reliability of authentication. At the same time, it is proposed to use the Thompson Sampling (TS) algorithm to dynamically adjust the duration of the authentication session. However, the defect of this solution is that the algorithm does not consider the risk of the session duration being too short or too long, and cannot avoid dangerous strategies. In 2024, Junfeng Miao et al. proposed a secure and effective drone-assisted IoT authentication protocol. The protocol uses elliptic curve cryptography (ECC) to complete identity authentication, which is not only secure but also resistant to known attacks. However, the disadvantage of this scheme is that the mathematical theory of ECC itself is relatively complex and difficult to implement in applications. Summary of the invention
[0004] The purpose of the present invention is to address the defects and shortcomings of the above-mentioned prior art and provide a low-overhead drone authentication method based on deep reinforcement learning, which is used to balance the overhead and security in the authentication process between drones and ground base stations. In order to ensure low overhead and high security in the drone authentication process, the present invention uses a deep reinforcement learning algorithm, takes authentication delay and energy consumption as security constraints, and dynamically selects drone authentication strategies, including authentication methods, session time, and data encryption strategies, so as to avoid high-risk strategies that may cause serious privacy leaks or excessive overhead.
[0005] The technical solution adopted by the present invention to solve the technical problem is: a low-overhead drone identity authentication method based on deep reinforcement learning, the method comprising the following steps:
[0006] Step 1: The drone identity authentication system consists of M drones and 1 ground base station, some of which have built-in PUF chips and some do not.
[0007] Step 2: Taking the i-th drone as an example, drone i sends a registration request with its real identity ID to the base station through a secure channel.
[0008] In step 2, each drone generates its own public key p i , and private key q i , and p i and ID form a message {p i ,ID} is sent to the base station as its own identity credential. The base station saves the received registration message and generates its own identity identifier BID, public key p0 and private key q0. The base station then sends {p0,BID} to the drone, and the drone verifies the legitimacy of the identity authentication with the return message. In addition, the drone with a PUF chip will also send a registration request to the base station using its remote ID. After receiving the registration request, the base station generates the temporary identity information of the drone and a series of challenges, and sends the challenges to the drone. After receiving the challenge, the drone uses the embedded PUF to generate a response and sends the response to the base station. After receiving the response, the base station stores the drone ID and challenge-response pair in the database.
[0009] Step 3: After the registration of M drones is completed, the identity authentication phase is carried out.
[0010] Step 4: Taking the i-th UAV as an example, the ground base station constructs a deep neural network, Q network, for UAV i, where 1≤i≤M. The weight parameters of the initial Q network are And initialize the learning rate to α and the discount factor to γ. And set the number of learning rounds The environment is reset in each round of learning. There are K moments in a round of learning. At each moment k, the base station selects an identity authentication strategy for drone i.
[0011] Step 5: Drone i sends an authentication message to the ground base station
[0012] In step 5, the drone generates a remote ID Including the drone's ID, altitude, speed, longitude and latitude information, and get the timestamp T according to the current time i (k) , generate random numbers Generate a pseudonym F through a hash function i (k) , an authentication message will be generated based on the above information Send to the base station.
[0013] Step 6: After receiving the message, the base station checks whether the drone has been registered. If not, drone i needs to register first. The base station will check the validity of the timestamp and whether the random number has been used to resist replay attacks.
[0014] Step 7: The ground base station constructs the current state vector for drone i
[0015] In step 7, the method for the base station to construct the system state vector for UAV i may be: recording the remote ID of the UAV at time k as PUF chip assembly status And the average authentication delay at the last moment and energy consumption And the number of authentications
[0016]
[0017] Step 8: The ground base station calculates the state vector Through strategy set A i Select the joint optimization strategy for identity authentication method, session time, and encryption strategy.
[0018] In step 8, the policy set is defined as in If it is a PUF-based authentication method, then it is an ECC-based authentication method. Represents the session time, represents the encryption strategy. The ground base station converts the state vector Input into the Q network. The Q network outputs the long-term discounted expected benefits of all participating node selection strategies in the current state And use the ò-greedy method to select an action based on this value. The action selected by the sensor device at this time is The specific method is: select Q with a probability of 1-ò i The action with the largest value randomly selects other actions with probability ò, where ò∈(0,1). ò determines the exploratory nature of the sensor device. The larger its value, the greater the randomness of the base station in selecting actions. This value is often set to a small positive number.
[0019] Step 9: The base station authenticates drone i according to the selected authentication strategy.
[0020] In step 9, when the base station selects the PUF-based authentication strategy, the base station extracts the corresponding challenge-response pair from the database based on the drone's ID, generates a timestamp and a random number, combines this information with the selected challenge, and uses the encryption strategy After receiving the message, the drone first decrypts and verifies the validity of the timestamp and whether the random number is reused. Then, it uses the built-in PUF hardware to generate a response based on the challenge received, and generates a new random number and timestamp. After encryption, it is transmitted back to the base station. After receiving it, the base station decrypts the authentication information, verifies the validity of the timestamp and the uniqueness of the random number, and verifies whether the received response matches the challenge-response pair stored in the database, thereby completing the identity authentication process. If the base station selects the ECC-based identity authentication strategy, the drone generates a random number and timestamp, which together with its own ID form the message n i , hash the message, using its own private key q i Sign and get signature i , the message n i With signature i Combining Encryption Strategies The base station decrypts the received authentication information and sends it to the base station according to the message n. i The identity ID in gets the corresponding public key p i Verify signature i The base station then generates a random number and a timestamp, and together with its own BID, forms a message n0, hashes the message, signs it with its own private key q0, obtains a signature ξ0, and combines the message n0 with the signature ξ0 using the encryption strategy The drone decrypts the received authentication information and obtains the corresponding public key p0 to verify the signature ξ0 according to the identity BID in the message n0. If the verification is successful, the two parties can have subsequent conversations.
[0021] Step 10: After the identity authentication is completed, the drone requests to create an authenticated session with the ground base station. The session duration is During the specified session time, the drone does not need to re-authenticate with the base station and can communicate continuously and securely. The base station recognizes the legitimacy of the drone. When the session time exceeds, the drone needs to re-authenticate with the base station.
[0022] Step 11: Ground Base Station Calculates Rewards
[0023] In step 11, according to the authentication delay Energy consumption The probability of the attacker successfully encrypting and decrypting Data protection level Calculate the current reward as follows:
[0024]
[0025] Among them, w T 、w Y 、w Q are weight coefficients, which measure the delay of executing the current task. Energy consumption And the probability of successful encryption and decryption by the attacker Importance in rewards.
[0026] Step 12: The ground base station stores the experience including status, action strategy, and reward into the experience pool.
[0027] In step 12, the ground base station changes the state Authentication strategy award Structured as an experience sequence And store the experience sequence into the experience pool middle.
[0028] Step 13: Randomly sample Z experiences from the experience pool to form a batch sample.
[0029] In step 13, the ground base station randomly samples Z experience {φ (z)} 1≤z≤k Form batch samples where z follows a uniform distribution from 1 to k.
[0030] Step 14: Update the weight parameters of the Q network
[0031] In step 14, the ground base station uses the Adam optimization algorithm to update the weight parameters of the Q network
[0032]
[0033] Among them, β is the discount factor for updating the weight parameters.
[0034] Step 15: Repeat steps 7-14 above until the ground base station learns a stable identity authentication method, session time, and data encryption strategy. converges to a stable value.
[0035] Beneficial effects:
[0036] 1. The present invention records the drone remote ID, PUF chip assembly status, average authentication delay and energy consumption at the previous moment, and the number of authentications to construct the system state, and combines the Q network optimization neural network selection action strategy based on the long-term discounted expected benefit value to balance the efficiency and security of the authentication process.
[0037] 2. The present invention can dynamically select drone identity authentication strategies, including identity authentication methods, session time, and data encryption strategies, to solve the problems of high overhead and low security in drone identity authentication. The method records the drone's remote ID (including drone ID, speed, altitude, longitude and latitude, etc.), the physical unclonable chip assembly, the average authentication delay and energy consumption at the last moment, and the number of authentications, thereby constructing the system state, and using deep reinforcement learning to select action strategies, thereby improving the efficiency and security of drone identity authentication. DETAILED DESCRIPTION
[0038] In order to more clearly understand the technical content of the present invention, the following embodiments are given in detail.
[0039] The present invention provides a low-overhead drone identity authentication method based on deep reinforcement learning, the method comprising the following steps:
[0040] Step 1: The drone identity authentication system includes two drones and one ground base station, one of which has a built-in PUF chip and the other does not have a PUF chip.
[0041] Step 2: Each drone generates its own public key p i , and private key q i , and p i and ID form a message {p i,ID} is sent to the base station as its own identity credential. The base station saves the received registration message and generates its own identity identifier BID, public key p0 and private key q0. The base station then sends {p0,BID} to the drone, and the drone verifies the legitimacy of the identity authentication with the return message. In addition, the drone with an embedded PUF chip will also send a registration request to the base station with its identity identifier. After receiving the registration request, the base station generates the temporary identity information of the drone and a series of challenges, and sends the challenges to the drone. After receiving the challenge, the drone uses the embedded PUF to generate a response and sends the response to the base station. After receiving the response, the base station stores the drone ID and challenge-response pair in the database.
[0042] Step 3: After the drone registration is completed, the identity authentication phase will begin.
[0043] Step 4: The ground base station constructs a deep neural network for the drone, the Q network, which consists of an input layer, a hidden layer, and an output layer. The input layer consists of 10 neurons, the hidden layer consists of 64 neurons, and the output layer consists of 18 neurons. The weight parameters of the initial Q network are The initial learning rate α is 0.001, the discount factor γ is 0.8, the number of randomly sampled experiences is 32, and the benefit function weight parameter w T 、w Y 、w Q They are 1, 0.5 and 1.5 respectively.
[0044] Step 5: Generate remote ID on drone Including the drone's ID, altitude, speed, longitude and latitude information, and get the timestamp T according to the current time i (k) , generate random numbers Generate a pseudonym F through a hash function i (k ), an authentication message will be generated based on the above information Send to the base station.
[0045] Step 6: After receiving the message, the base station checks whether the drone has been registered. If not, the drone needs to register first. The base station will check the validity of the timestamp and whether the random number has been used to resist replay attacks.
[0046] Step 7: According to the remote ID of the drone at time k PUF chip assembly status And the authentication delay of the previous moment Energy consumption and authentication times The base station builds the system state vector for the drone as:
[0047]
[0048] Step 8: The ground base station converts the state vector Input into the Q network. The Q network outputs the long-term discounted expected benefits of all participating node selection strategies in the current state And use the value to select the joint optimization strategy of identity authentication method, session time and encryption strategy using the ò-greedy method Note that the action selected by the sensor device at this time is The specific method is: select Q with a probability of 1-ò i The action with the largest value randomly selects other actions with probability ò, where ò∈(0,1). ò determines the exploratory nature of the sensor device. The larger its value, the greater the randomness of the base station in selecting actions. This value is often set to a small positive number.
[0049] Step 9: Perform authentication according to the selected authentication method.
[0050] Step 10: After the identity authentication is completed, the drone requests to create an authenticated session with the ground base station. The session duration is
[0051] Step 11: The ground base station determines the authentication delay Energy consumption The probability of the attacker successfully encrypting and decrypting Data protection level Calculate the current reward as follows:
[0052]
[0053] Step 12: The ground base station will be Authentication strategy award Structured as an experience sequence And store the experience sequence into the experience pool middle.
[0054] Step 13: The ground base station randomly samples 32 experience {φ( z )} 1≤z≤k Form batch samples.
[0055] Step 14: The ground base station uses the Adam optimization algorithm to update the weight parameters of the Q network as follows:
[0056]
[0057] Step 15: Repeat steps 7-14 until the ground base station learns a stable identity authentication method, session time, and data encryption strategy. converges to a stable value.
[0058] The above embodiments are only preferred embodiments of the present invention and cannot be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the present invention.
Claims
1. A low-overhead drone identity authentication method based on deep reinforcement learning, characterized in that: The method comprises the following steps: Step 1: The drone identity authentication system consists of M drones and 1 ground base station, some of which have built-in PUF chips and some do not; Step 2: Taking the i-th drone as an example, drone i sends a registration request with its real identity ID to the base station through a secure channel; Step 3: After the registration of M drones is completed, the identity authentication phase is carried out; Step 4: Taking the i-th UAV as an example, the ground base station constructs a deep neural network for UAV i, Q network, where 1≤i≤M, and initializes the weight parameters of Q network as And initialize the learning rate to α, the discount factor to γ, and set the number of learning rounds The environment is reset in each round of learning. There are K moments in a round of learning. At each moment k, the base station selects an identity authentication strategy for drone i. Step 5: Drone i sends an authentication message to the ground base station Step 6: After receiving the message, the base station checks whether the drone has been registered. If not, drone i needs to register first. The base station will check the validity of the timestamp and whether the random number has been used to resist replay attacks; Step 7: The ground base station constructs the current state vector for drone i Step 8: The ground base station calculates the state vector Through strategy set A i Select the joint optimization strategy of identity authentication method, session time, and encryption strategy; Step 9: The base station authenticates the drone i according to the selected authentication strategy; Step 10: After the identity authentication is completed, the drone requests to create an authenticated session with the ground base station. The session duration is During the specified session time, the drone does not need to re-authenticate with the base station and can communicate securely and continuously. The base station recognizes the legitimacy of the drone. When the session time is exceeded, the drone needs to re-authenticate with the base station. Step 11: Ground Base Station Calculates Rewards Step 12: The ground base station stores the experience including status, action strategy, and reward into the experience pool; Step 13: Randomly sample Z experiences from the experience pool to form a batch sample; Step 14: Update the weight parameters of the Q network Step 15: Repeat steps 7-14 above until the ground base station learns a stable identity authentication method, session time, and data encryption strategy. converges to a stable value.
2. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 2, each drone generates its own public key p i , and private key q i , and p i and ID form a message {p i ,ID} is sent to the base station as its own identity credential. The base station saves the received registration message and generates its own identity identifier BID, public key p0 and private key q0. The base station then sends {p0, BID} to the drone. The drone verifies the legitimacy of the identity authentication with the return message. In addition, the drone with PUF chip will also send a registration request to the base station with its remote ID. After receiving the registration request, the base station generates the temporary identity information of the drone and a series of challenges, and sends the challenges to the drone. After receiving the challenge, the drone uses the embedded PUF to generate a response and sends the response to the base station. After receiving the response, the base station stores the drone ID and challenge-response pair in the database.
3. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 7, the method by which the base station constructs the system state vector for drone i is: record the remote ID of the drone at time k as PUF chip assembly status And the average authentication delay at the last moment and energy consumption And the number of authentications 4. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 5, the drone generates a remote ID Includes the drone's ID, altitude, speed, longitude and latitude information, and obtains a timestamp based on the current time Generate random numbers Generate pseudonyms through hash functions An authentication message will be generated based on the above information Send to the base station.
5. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 8, the policy set is defined as in If it is a PUF-based authentication method, then it is an ECC-based authentication method. Represents the session time, Represents the encryption strategy, the ground base station converts the state vector Input into the Q network, the Q network outputs the long-term discounted expected benefits of all participating node selection strategies in the current state And use the ò-greedy method to select an action through this value. The action selected by the sensor device at this time is The specific method is: select Q with a probability of 1-ò i The action with the largest value randomly selects other actions with a probability of ò, where ò∈(0,1). ò determines the exploratory nature of the sensor device. The larger its value, the greater the randomness of the base station in selecting an action. This value is often set to a small positive number.
6. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 9, when the base station selects the PUF-based authentication strategy, the base station extracts the corresponding challenge-response pair from the database according to the drone's ID, generates a timestamp and a random number, and combines this information with the selected challenge and uses the encryption strategy. After receiving the message, the drone first decrypts and verifies the validity of the timestamp and whether the random number is reused. Then, it uses the built-in PUF hardware to generate a response based on the challenge received, and generates a new random number and timestamp. After encryption, it is transmitted back to the base station. After receiving it, the base station decrypts the authentication information, verifies the validity of the timestamp and the uniqueness of the random number, and verifies whether the received response matches the challenge-response pair stored in the database, thereby completing the identity authentication process. If the base station selects the ECC-based identity authentication strategy, the drone generates a random number and timestamp, which together with its own ID form the message n i , hash the message, using its own private key q i Sign and get signature i , the message n i With signature i Combining Encryption Strategies Encrypt and send them to the base station. The base station decrypts the received authentication information and sends it to the base station according to the message n. i The identity ID in gets the corresponding public key p i Verify signature i , then the base station generates a random number and a timestamp, and together with its own BID, forms a message n0, hashes the message, signs it with its own private key q0, obtains a signature ξ0, combines the message n0 with the signature ξ0, and uses the encryption strategy The two parties encrypt the information and send it to the drone. The drone decrypts the received authentication information and obtains the corresponding public key p0 to verify the signature ξ0 according to the identity BID in the message n0. If the verification is successful, the two parties can have subsequent conversations.
7. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 11, according to the authentication delay Energy consumption The probability of the attacker successfully encrypting and decrypting Data protection level Calculate the current reward as follows: Among them, w T 、w Y 、w Q are weight coefficients, which measure the delay of executing the current task. Energy consumption And the probability of successful encryption and decryption by the attacker Importance in rewards.
8. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 12, the ground base station sets the status Authentication strategy award Structured as an experience sequence And store the experience sequence into the experience pool middle.
9. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 13, the ground base station randomly samples Z experience {φ (z) } 1≤z≤k Form batch samples where z follows a uniform distribution from 1 to k.
10. A low-overhead drone identity authentication method based on deep reinforcement learning according to claim 1, characterized in that: In step 14, the ground base station uses the Adam optimization algorithm to update the weight parameters of the Q network Among them, β is the discount factor for updating the weight parameters.
Citation Information
Patent Citations
Elliptic curve encryption-based unmanned aerial vehicle and base station communication identity authentication method
CN112073964A
Unmanned aerial vehicle identity authentication method based on timely generation and updating of authentication information
CN116506857A
Unmanned aerial vehicle group access authentication method based on ECC algorithm and certificate
CN118317305A
Apparatus for compacting test piece
KR102282031B1
Artificial Intelligence-Based Generation of Anthropomorphic Signatures and use Thereof
US20210326433A1