Optimization method and device of unmanned aerial vehicle communication network, electronic equipment and storage medium
By optimizing the resource allocation of the UAV communication network and utilizing a reinforcement learning environment and the H-SAC algorithm, the problems of low transmission rate and high risk of eavesdropping in the UAV semantic communication system were solved, thereby improving security and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-10
AI Technical Summary
Existing drone semantic communication systems suffer from low semantic transmission rates for users and a high risk of eavesdropping, failing to provide secure and reliable drone semantic transmission.
By acquiring the resource configuration and average secure semantic transmission rate of the UAV communication network, and utilizing a reinforcement learning environment and the hybrid flexible action-evaluation H-SAC algorithm, the resource configuration of the UAV communication network is optimized, including UAV flight trajectory, transmission power, phase shift of intelligent reflector, and number of semantic symbols. A semantic communication simulation model is then constructed to improve transmission rate and security.
Significant optimization of the drone communication network has been achieved, improving the timeliness and effectiveness of semantic transmission rate, reducing the risk of eavesdropping, and enhancing the security and reliability of semantic transmission for users' drones.
Smart Images

Figure CN121645267A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to an optimization method, apparatus, electronic device, and computer-readable storage medium for unmanned aerial vehicle (UAV) communication networks. Background Technology
[0002] Unmanned aerial vehicle (UAV) communication systems are crucial for enabling information exchange between UAVs and ground stations, other aircraft, and other communication nodes. A typical UAV communication system includes communication links, spectrum management, and security components. Communication links encompass data links, control links, and feedback links, while spectrum management is essential for the efficient use of limited spectrum resources by UAVs.
[0003] Semantic communication is a communication method based on semantic understanding, aiming to make communication more intelligent by parsing and understanding the semantics of the communication content. In the field of unmanned aerial vehicles (UAVs), semantic communication can be used to achieve more accurate information delivery and intelligent decision support, thereby improving the overall efficiency of UAV systems.
[0004] IRS (Intelligent Reflective Surface) technology utilizes programmable elements on a smart surface to control the propagation direction and path of electromagnetic waves through passive reflection. This makes the transmission path of communication signals more complex and diverse, making it difficult for eavesdroppers to accurately intercept user information. Therefore, intelligent reflective surface technology can effectively combat eavesdropping. IRS can also enhance security by setting timed triggering changes for the programmable elements on the smart surface within a specific Wi-Fi (wireless fidelity) environment.
[0005] Existing drone semantic communication systems suffer from low semantic transmission rates and a significant risk of eavesdropping, leading to semantic leakage issues and failing to provide users with secure and reliable drone semantic transmission. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to address the above-mentioned shortcomings of the prior art by providing an optimization method, apparatus, electronic device and computer-readable storage medium for UAV communication networks. This method can achieve significant and effective optimization of UAV communication networks, reduce the situation of low optimization efficiency and latency, improve the timeliness and effectiveness of semantic transmission rate optimization, reduce the risk of eavesdropping, reduce semantic leakage problems, and improve the security and reliability of semantic transmission of user UAVs.
[0007] In a first aspect, the present invention provides an optimization method for an unmanned aerial vehicle (UAV) communication network, comprising: obtaining the resource configuration and average security semantic transmission rate of the UAV communication network at the current time T; determining whether the average security semantic transmission rate at the current time T is less than a required rate threshold; in response to the average security semantic transmission rate at the current time T being less than the required rate threshold, determining a target optimization strategy for the UAV communication network based on a reinforcement learning environment, a hybrid flexible action-evaluation (H-SAC) algorithm, and the resource configuration at the current time T; and optimizing the resource configuration of the UAV communication network based on the target optimization strategy and resource configuration constraints.
[0008] Preferably, after determining whether the average secure semantic transmission rate at the current time T is less than the required rate threshold, and before determining the target optimization strategy of the UAV communication network based on the reinforcement learning environment, the hybrid flexible action-evaluation H-SAC algorithm, and the resource allocation at the current time T, the optimization method for the UAV communication network further includes: constructing a reinforcement learning environment for the UAV communication network based on a Markov decision process, wherein the reinforcement learning environment includes an agent, a state space, an action space, a reward function, and state transition probabilities.
[0009] The intelligent agent serves as the control center for the unmanned aerial vehicle (UAV) communication network.
[0010] state space
[0011] Action space a t ={{q t},Θ μ(i) ,P i,t ,k i,t},
[0012] Reward function
[0013] State transition probability
[0014] in, This represents the channel factor for all users and eavesdroppers in the drone communication network at time t. a represents the average secure semantic transmission rate of the UAV communication network at time t. t-1 The r represents the target action selected by the UAV communication network at time t-1. t Let {q} represent the reward value of the drone communication network at time t. t} indicates that the drone's flight trajectory is adjusted at time t, Θ μ(i) This indicates adjusting the phase shift of the smart reflector corresponding to user i, P i,t This indicates adjusting the transmit power from the UAV to user i at time t, k i,t This indicates the adjustment of the number of semantic symbols for user i at time t. This represents the secure semantic transmission rate of user i in the UAV communication network. Γ represents the average user semantic transmission rate of the UAV communication network. i Γ represents the semantic transmission rate of user i in the UAV communication network. EAV Let s represent the semantic transmission rate of user i corresponding to the eavesdropper in the UAV communication network, D represent the total number of users in the UAV communication network, and s represent the total number of users in the UAV communication network. t+1 s represents the state of the UAV communication network at time t+1. t a represents the state of the UAV communication network at time t. t Let i represent the target action selected by the UAV communication network at time t, where i = 1, 2, ..., D, and D represents the total number of users in the UAV communication network.
[0015] Preferably, the step of determining the target optimization strategy for the UAV communication network based on the reinforcement learning environment, the Hybrid Flexible Action-Evaluation (H-SAC) algorithm, and the resource allocation at the current time T specifically includes: generating the initial state of the UAV communication network based on the resource allocation at the current time T; selecting a target action and calculating the state, reward, and state transition probability of the UAV communication network under the target action; calculating the state action value of the target action based on the Q-value function, the reward, and the state transition probability under the target action, and optimizing the Q-value function based on the state action value of the target action; and selecting the target action corresponding to the maximum state action value as the target optimization strategy for the UAV communication network.
[0016] Preferably, the resource configuration includes the UAV flight trajectory, UAV transmission power, the number of semantic symbols for each user, and the phase shift of the intelligent reflector corresponding to each user. The step of generating the initial state of the UAV communication network based on the resource configuration at the current time T specifically includes: calculating the channel factor of user i and the eavesdropper corresponding to user i at the current time T based on the UAV flight trajectory at the current time T and the phase shift of the intelligent reflector corresponding to user i; calculating the signal-to-noise ratio (SNR) of user i and the eavesdropper corresponding to user i at the current time T based on the channel factor and the UAV transmission power at the current time T; matching the semantic transmission rate of user i and the eavesdropper corresponding to user i at the current time T based on the mapping relationship between SNR and semantic transmission rate and the SNR at the current time T; calculating the average secure semantic transmission rate and the average user semantic transmission rate of the UAV communication network at the current time T based on the semantic transmission rate at the current time T; assuming that the target action of the UAV communication network at time T-1 is empty, and the report of the UAV communication network at the current time T is empty.
[0017] Preferably, after calculating the signal-to-noise ratio (SNR) of user i and the corresponding eavesdropper at current time T based on the channel factor and UAV transmit power at current time T, and before matching the semantic transmission rate of user i and the corresponding eavesdropper at current time T based on the mapping relationship between SNR and semantic transmission rate and the SNR at current time T, the optimization method for the UAV communication network further includes: acquiring the first text transmission data of user i, the channel bandwidth and text transmission requirements, and the true SNR of user i and the corresponding eavesdropper; constructing a semantic communication simulation model based on the DeepSC spectral clustering algorithm of deep learning, and predicting the second text transmission data after semantic communication between user i and the corresponding eavesdropper based on the semantic communication simulation model and the first text transmission data; calculating the semantic similarity of semantic communication between user i and the corresponding eavesdropper based on the first text transmission data and the second text transmission data; calculating the semantic transmission rate of user i and the corresponding eavesdropper based on the semantic similarity, the channel bandwidth and the number of semantic symbols of user i; and obtaining the mapping relationship between SNR and semantic transmission rate by associating the semantic transmission rate and the true SNR.
[0018] Preferably, the text transmission requirements include the expected amount of semantic information and the expected number of characters in the text. The calculation of the semantic transmission rate of user i and the corresponding eavesdropper based on the semantic similarity, the channel bandwidth, and the number of semantic symbols for user i specifically includes: calculating the semantic transmission rate of user i and the corresponding eavesdropper according to formulas (1) and (2).
[0019]
[0020] Where W is the channel bandwidth, I is the expected amount of semantic information, and k i Let L represent the number of semantic symbols for user i, L represent the expected number of words in the text, and φ represent the number of semantic symbols for user i. i φ represents the semantic similarity between user i and user i. EAV This represents the semantic similarity between user i and the eavesdropper.
[0021] Preferably, the drone trajectory includes the drone's initial position, collision avoidance position, and speed, and the resource configuration constraints include a first constraint, a second constraint, a third constraint, a fourth constraint, a fifth constraint, and a sixth constraint.
[0022] The first constraint is the starting point position of the UAV, q[0] = q initial ,
[0023] The second constraint is speed.
[0024] The third constraint is the collision avoidance position.
[0025] The fourth constraint is the UAV's transmit power.
[0026] The fifth constraint is the phase shift θ of the smart reflector. m ∈{0,2π},
[0027] The sixth constraint is the number of semantic symbols: 0 ≤ k ≤ k max Where q[0] represents the starting position of the drone, q initial V represents the preset starting point position, q[n+1] represents the position of the drone at time n+1, q[n] represents the position of the drone at time n, and V min V represents the minimum speed of the drone. max Let δ represent the maximum speed of the drone, and let R = {1, ..., R}. Indicates the first The location of the intelligent reflective surface, P u Indicates the drone's transmission power. θ represents the maximum transmit power of the drone. m Indicates intelligent reflective surface m The phase shift, M represents the set of intelligent reflective surfaces, k represents the number of semantic symbols, k max Indicates the maximum number of semantic symbols.
[0028] Secondly, the present invention also provides an optimization device for an unmanned aerial vehicle (UAV) communication network, comprising: an acquisition module, a judgment module, a determination module, and an optimization module. The acquisition module is used to acquire the resource configuration and average security semantic transmission rate of the UAV communication network at the current time T. The judgment module, connected to the acquisition module, is used to determine whether the average security semantic transmission rate at the current time T is less than the required rate threshold. The determination module, connected to the judgment module, is used to determine a target optimization strategy for the UAV communication network based on a reinforcement learning environment, the H-SAC algorithm, and the resource configuration at the current time T in response to the average security semantic transmission rate at the current time T being less than the required rate threshold. The optimization module, connected to the determination module, is used to optimize the resource configuration of the UAV communication network based on the target optimization strategy and resource configuration constraints.
[0029] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the optimization method for the UAV communication network provided in the first aspect above.
[0030] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the optimization method for the UAV communication network provided in the first aspect.
[0031] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for optimizing unmanned aerial vehicle (UAV) communication networks. By acquiring the resource configuration and average secure semantic transmission rate (SSE) of the UAV communication network in real time, and based on the average SSE and demand rate thresholds, it determines whether the resource configuration of the UAV communication network needs optimization. In response to the need for optimization, it determines a target optimization strategy for the UAV communication network based on a reinforcement learning environment, the H-SAC algorithm, and the resource configuration. Based on the target optimization strategy and resource configuration constraints, it optimizes the resource configuration of the UAV communication network. Therefore, this invention can achieve significant and effective optimization of UAV communication networks, reducing low optimization efficiency and latency, improving the timeliness and effectiveness of semantic transmission rate optimization, reducing the risk of eavesdropping, reducing semantic leakage problems, and improving the security and reliability of semantic transmission for user UAVs. Attached Figure Description
[0032] Figure 1 This is a flowchart of an optimization method for a UAV communication network according to Embodiment 1 of the present invention;
[0033] Figure 2 This is an example diagram of a drone communication network according to Embodiment 1 of the present invention;
[0034] Figure 3 This is a schematic diagram of the structure of an optimization device for a drone communication network in Embodiment 2 of the present invention. Detailed Implementation
[0035] To enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0036] It is understood that the specific embodiments and accompanying drawings described herein are merely for explaining the invention and are not intended to limit the invention.
[0037] It is understood that, without conflict, the various embodiments and features in the embodiments of the present invention can be combined with each other.
[0038] It is understood that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, while the parts unrelated to the present invention are not shown in the drawings.
[0039] It is understood that each unit or module involved in the embodiments of the present invention may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units or modules may be integrated into one entity structure.
[0040] It is understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of this invention may occur in a different order than that marked in the accompanying drawings.
[0041] It is understood that the flowcharts and block diagrams of this invention illustrate the possible architecture, functions, and operations of systems, apparatuses, devices, and methods according to various embodiments of this invention. Each block in the flowchart or block diagram may represent a unit, module, program segment, or code, containing executable instructions for implementing the specified function. Furthermore, each block or combination of blocks in the block diagram and flowchart can be implemented using a hardware-based system to achieve the specified function, or using a combination of hardware and computer instructions.
[0042] It is understood that the units and modules involved in the embodiments of the present invention can be implemented by software or by hardware. For example, the units and modules can be located in a processor.
[0043] Example 1:
[0044] like Figure 1 As shown, this embodiment provides a method for optimizing a drone communication network. The method for optimizing a drone communication network includes:
[0045] S101, obtain the resource configuration and average secure semantic transmission rate of the UAV communication network at the current time T.
[0046] Specifically, resource allocation includes the drone's flight trajectory, drone's transmission power, the number of semantic symbols for each user, and the phase shift of the intelligent reflector corresponding to each user.
[0047] In this embodiment, as Figure 2 As shown, the communication nodes of the UAV communication network include a smart reflector, the UAV, the user, and the eavesdropper. The UAV provides semantic communication services to the user and is equipped with a single antenna; each main user is equipped with one antenna. A three-dimensional Cartesian coordinate system is established for all communication nodes, and the position of communication node A in the UAV communication network can be represented as w. A =[x A ,y A ,z A ] T , where x A y A z ALet x, y, and z represent the corresponding x, y, and z axis coordinates, respectively. The user set can be represented as D = {1, ..., D}, the eavesdropper set as K = {1, ..., K}, the intelligent reflector set as R = {1, ..., R}, the time set as T = {1, ..., T}, and the total bandwidth as B. During the flight time, the UAV trajectory set can be represented as Q = {Q...} t}, where t∈T, Q t This represents the drone's flight trajectory from the start time to time t during the flight time. Assume the i-th user is matched with the i-th... Each intelligent reflective surface is denoted by μ(i), where 1 < μ(i) < R. According to the formula... Calculate the phase shift matrix of the intelligent reflector μ(i), where φ m,μ(i) and α m,μ(i) Let k represent the phase and switching state of the m-th reflecting element in the intelligent reflecting surface μ(i), respectively. i Let be the number of semantic symbols for the i-th user.
[0048] S102, determine whether the average secure semantic transmission rate at the current time T is less than the required rate threshold.
[0049] In this embodiment, the actual semantic transmission rates of users and eavesdroppers in the UAV communication network are collected. The difference between the actual semantic transmission rate of each user in the UAV communication network and the actual semantic transmission rate of the corresponding eavesdropper is calculated to obtain the actual secure semantic transmission rate of each user in the UAV communication network. The average actual secure semantic transmission rate of each user in the UAV communication network is calculated, and the average actual secure semantic transmission rate is used as the average secure semantic transmission rate at the current time T. Semantic transmission rate is a comprehensive indicator that considers both the efficiency of information transmission and emphasizes the quality and effectiveness of information. It is an important reference indicator for measuring the user experience and usability of an information system. This embodiment, by evaluating the semantic transmission rate of the UAV communication network, can determine the efficiency, quality, and effectiveness of information transmission in the UAV communication network.
[0050] It should be noted that, generally, the average user semantic transmission rate of a UAV communication network is greater than the average security semantic transmission rate. Therefore, this embodiment takes determining whether the UAV communication network needs resource configuration optimization based on the average security semantic transmission rate and the demand rate threshold as an example. This embodiment can also determine whether the average security semantic transmission rate and the average user semantic transmission rate are less than their respective demand rate thresholds. If both are less than their respective demand rate thresholds, or if either the average security semantic transmission rate or the average user semantic transmission rate is less than its respective demand rate threshold, then the resource configuration of the UAV communication network is optimized. Additionally, this embodiment can also determine whether the actual semantic transmission rates of the collected users and eavesdroppers are less than the demand rate threshold to determine whether the resource configuration of the UAV communication network needs optimization.
[0051] S103, in response to the fact that the average secure semantic transmission rate at the current time T is less than the required rate threshold, the target optimization strategy of the UAV communication network is determined based on the reinforcement learning environment, the hybrid flexible action-evaluation H-SAC algorithm and the resource allocation at the current time T.
[0052] In this embodiment, although existing drones are equipped with high-performance processors, their computing power is still limited compared to ground computing platforms (such as servers or cloud computing), making it difficult to handle highly complex computing tasks. Furthermore, under high load, processor heat dissipation issues can lead to performance degradation, affecting the drone's real-time processing capabilities and flight stability. Additionally, despite advancements in battery technology, current battery life remains limited, with typical drone flight times ranging from 20 minutes to several hours, restricting their application in long-duration missions. Real-time analysis and decision-making can be time-consuming, especially when processing highly uncertain data. Therefore, this embodiment optimizes resource allocation based on a reinforcement learning environment and the H-SAC algorithm when the average secure semantic transmission rate is below the required rate threshold. It achieves a reasonable trade-off and optimization between algorithm complexity, application scenarios, and drone hardware capabilities to ensure the effectiveness and reliability of drone communication while improving the average secure semantic transmission rate and average user semantic transmission rate of the drone communication network.
[0053] It should be noted that, given the superior computing and power resources of the UAV communication network and the UAV's real-time response capabilities, this embodiment can further divide the UAV flight time into multiple time periods, collect resource configuration data for each time period, and directly determine the target optimization strategy for each time period based on the resource configuration of each time period. This target optimization strategy is then used to optimize the resource configuration for the next time period. By optimizing the resource configuration for each time period, the UAV's payload, energy, and other resources can be utilized to the maximum extent, reducing resource waste. Furthermore, by dynamically adjusting resource configuration and optimization strategies according to real-time environmental changes, the UAV's adaptability can be improved. Ultimately, this better achieves the predetermined goals of increasing the average secure semantic transmission rate and average user semantic transmission rate of the UAV communication network.
[0054] Optionally, after S102: determining whether the average secure semantic transmission rate at the current time T is less than the required rate threshold, and before determining the target optimization strategy of the UAV communication network based on the reinforcement learning environment, the hybrid flexible action-evaluation H-SAC algorithm, and the resource allocation at the current time T, the optimization method for the UAV communication network further includes:
[0055] S105, based on Markov decision processes, constructs a reinforcement learning environment for UAV communication networks.
[0056] Specifically, the reinforcement learning environment includes the agent, state space, action space, reward function, and state transition probabilities.
[0057] The intelligent agent serves as the control center for the unmanned aerial vehicle (UAV) communication network.
[0058] state space
[0059] Action space a t ={{q t},Θ μ(i) ,P i,t ,k i,t},
[0060] Reward function
[0061] State transition probability
[0062] in, This represents the channel factor for all users and eavesdroppers in the drone communication network at time t. a represents the average secure semantic transmission rate of the UAV communication network at time t. t-1 The r represents the target action selected by the UAV communication network at time t-1. t Let {q} represent the reward value of the drone communication network at time t. t} indicates that the drone's flight trajectory is adjusted at time t, Θ μ(i) This indicates adjusting the phase shift of the smart reflector corresponding to user i, P i,t This indicates adjusting the transmit power from the UAV to user i at time t, k i,t This indicates the adjustment of the number of semantic symbols for user i at time t. This represents the secure semantic transmission rate of user i in the UAV communication network. Γ represents the average user semantic transmission rate of the UAV communication network. i Γ represents the semantic transmission rate of user i in the UAV communication network. EAV Let s represent the semantic transmission rate of user i corresponding to the eavesdropper in the UAV communication network, D represent the total number of users in the UAV communication network, and s represent the total number of users in the UAV communication network. t+1 s represents the state of the UAV communication network at time t+1. t a represents the state of the UAV communication network at time t. t Let i represent the target action selected by the UAV communication network at time t, where i = 1, 2, ..., D, and D represents the total number of users in the UAV communication network.
[0063] In this embodiment, the problem of maximizing the average security semantic transmission rate and the average user semantic transmission rate of the UAV communication network is transformed into a Markov decision process problem. Key elements for optimizing the semantic transmission rate of the UAV communication network are extracted, and a model-free deep reinforcement learning environment is established. This embodiment considers both the average security semantic transmission rate and the average user semantic transmission rate in the reward function design, ensuring both a high security semantic transmission rate and a high user semantic transmission rate, thereby improving the user experience.
[0064] Specifically, based on the reinforcement learning environment, the hybrid flexible action-evaluation H-SAC algorithm, and the resource allocation at the current time T, the target optimization strategy of the UAV communication network is determined, including steps S1031-S1034:
[0065] S1031, Based on the resource configuration at the current time T, generate the initial state of the UAV communication network.
[0066] Specifically, S1031: Based on the resource configuration at the current time T, generate the initial state of the UAV communication network, including: calculating the channel factor of user i and the corresponding eavesdropper at the current time T based on the UAV flight trajectory at the current time T and the phase shift of the intelligent reflector corresponding to user i; calculating the signal-to-noise ratio (SNR) of user i and the corresponding eavesdropper at the current time T based on the channel factor and the UAV transmission power; matching the semantic transmission rate of user i and the corresponding eavesdropper at the current time T based on the mapping relationship between SNR and semantic transmission rate and the SNR at the current time T; calculating the average secure semantic transmission rate and the average user semantic transmission rate of the UAV communication network at the current time T based on the semantic transmission rate at the current time T; assuming that the target action of the UAV communication network at time T-1 is empty, and the report of the UAV communication network at the current time T is empty.
[0067] In this embodiment, it is assumed that the channel from the UAV to the l-th user on the ground follows Rayleigh fading, and the corresponding channel factor is . Where ρ represents the channel power gain at a reference distance of 1m, α represents the path loss factor, and d u,l [n] represents the distance from the drone to the l-th user on the ground. q[n] represents the position of the drone at any given time. l This represents the location of the l-th user. Similarly, the channel factor from the drone to the ground eavesdropper can be obtained as follows:
[0068] Assuming the drone reaches the... The channels of each intelligent reflector follow a line-of-sight distribution, and the corresponding channel factors are: in, Indicates the drone has reached the The distance between each intelligent reflective surface
[0069] Assuming the first The channel from the intelligent reflector to the i-th user on the ground follows a Ricean distribution, with the corresponding channel factor being...
[0070] Where β represents the path loss exponent and K represents the Rice factor. Similarly, the corresponding path loss exponent can be obtained. The channel factor from the intelligent reflector to the ground eavesdropper is:
[0071]
[0072] Suppose the i-th user is matched with the i-th user. Each intelligent reflector is μ(i). The i-th user includes the following two types of channel paths: Meanwhile, since the eavesdropper is trying to eavesdrop on all existing messages as much as possible, the corresponding total channel factor is:
[0073]
[0074] In summary, the signal-to-noise ratios at the i-th user and the corresponding eavesdropper are as follows: in, Let be the noise variance at the i-th user. Let V be the noise variance at the location of the eavesdropper.
[0075] Optionally, after calculating the signal-to-noise ratio (SNR) of user i and the corresponding eavesdropper at current time T based on the channel factor and UAV transmit power at current time T, and before matching the semantic transmission rate of user i and the corresponding eavesdropper at current time T based on the mapping relationship between SNR and semantic transmission rate and the SNR at current time T, the optimization method for the UAV communication network further includes: obtaining the first text transmission data of user i, the channel bandwidth and text transmission requirements, and the true SNR of user i and the corresponding eavesdropper; constructing a semantic communication simulation model based on the DeepSC spectral clustering algorithm of deep learning, and predicting the second text transmission data after semantic communication between user i and the corresponding eavesdropper based on the semantic communication simulation model and the first text transmission data; calculating the semantic similarity of semantic communication between user i and the corresponding eavesdropper based on the first text transmission data and the second text transmission data; calculating the semantic transmission rate of user i and the corresponding eavesdropper based on the semantic similarity, channel bandwidth and the number of semantic symbols of user i; and obtaining the mapping relationship between SNR and semantic transmission rate by associating the semantic transmission rate and the true SNR.
[0076] In this embodiment, in the UAV communication network, the semantic encoder and channel encoder are deployed at the sender, and the semantic decoder and channel decoder are deployed at the receiver. At the UAV sender, the transmitted text data can be represented as s = [w1, w2, ..., w...]. i ,…,w l ], where l represents the length of the currently transmitted sentence, w i This represents the i-th word in the sentence. At both the user's and the eavesdropper's ends, the received signal is decoded via channel decoding and semantic decoding to recover the transmitted text data.
[0077] This embodiment establishes a semantic communication simulation model based on DeepSC. Given a range of the number of sent semantic symbols, the first text data s transmitted by the UAV is taken as the first text data transmitted by user i. The first text data s is processed by a semantic encoder and a channel encoder at the sending end and input into the simulation channel. At the receiving end, it is processed by a channel decoder and a semantic encoder to recover the first text data at the semantic level, thus obtaining the user's second text data s. -1According to the formula The semantic similarity between the first and second transmitted text data is calculated, where ζ(·) represents the BERT pre-trained model. The semantic similarity φ is a continuous value between 0 and 1; the more similar the predicted sentence and the target sentence are, the higher the semantic similarity. This embodiment uses DeepSC to establish a semantic communication simulation model to predict the second text transmission data between the user and the eavesdropper, improving the prediction accuracy of the text transmission data and enhancing the accuracy of the average secure semantic transmission rate and the average user semantic transmission rate of the UAV communication network.
[0078] Specifically, text transmission requirements include the expected amount of semantic information and the expected number of words in the text.
[0079] Specifically, based on semantic similarity, channel bandwidth, and the number of semantic symbols for user i, the semantic transmission rate of user i and the corresponding eavesdropper is calculated, including: calculating the semantic transmission rate of user i and the corresponding eavesdropper according to formulas (1) and (2):
[0080]
[0081] Where W is the channel bandwidth, I is the expected amount of semantic information, and k i Let L represent the number of semantic symbols for user i, L represent the expected number of words in the text, and φ represent the number of semantic symbols for user i. i φ represents the semantic similarity between user i and user i. EAV This represents the semantic similarity between user i and the eavesdropper.
[0082] In this embodiment, after calculating the semantic similarity of semantic communication between user i and the corresponding eavesdropper, the semantic transmission rate of user i and the corresponding eavesdropper is calculated based on the semantic similarity of semantic communication between user i and the corresponding eavesdropper, formula (1), and formula (2). The true signal-to-noise ratio (SNR) of user i and the corresponding eavesdropper is collected, and the semantic similarity of semantic communication between user i and the corresponding eavesdropper is associated with the true SNR to obtain the mapping relationship between SNR and semantic transmission rate. The mapping relationship between SNR and semantic transmission rate is imported into the reinforcement learning environment. This embodiment improves the interaction rate between the agent and the environment and increases the efficiency of determining the target optimization strategy of the UAV communication network by constructing the mapping relationship between SNR and semantic transmission rate and importing the mapping relationship between SNR and semantic transmission rate into the reinforcement learning environment.
[0083] After calculating the signal-to-noise ratio (SNR) of user i and the corresponding eavesdropper at the current time T, the semantic transmission rate of user i and the corresponding eavesdropper at the current time T can be quickly matched based on the mapping relationship between SNR and semantic transmission rate and the SNR at the current time T.
[0084] According to the formula formula Given the semantic transmission rate at the current time T, calculate the average secure semantic transmission rate and average user semantic transmission rate of the UAV communication network at the current time T. Furthermore, assuming the target action of the UAV communication network at time T-1 is empty and the reward of the UAV communication network at the current time T is empty, quickly determine the initial state of the UAV communication network based on the calculated channel factor and average secure semantic transmission rate.
[0085] S1032, Select the target action and calculate the state, reward, and state transition probability of the UAV communication network under the target action.
[0086] In this embodiment, the agent iteratively selects and optimizes the UAV trajectory, UAV transmission power, number of semantic symbols, and phase shift of the intelligent reflector through the H-SAC algorithm, and inputs them into the UAV communication network, that is, selects the target action, enters the state under the target action, and obtains the current reward value and state transition probability through the reward function.
[0087] After selecting a target action, similar to the process of quickly determining the initial state of the UAV communication network, based on the resource configuration after selecting the target action, the channel factor and signal-to-noise ratio under the target action are calculated. Based on the signal-to-noise ratio under the target action, the semantic transmission rates of users and responders under the target action are matched to calculate the average security semantic transmission rate and the average user semantic transmission rate under the target action, and to determine the state, reward, and transfer probability under the target action.
[0088] S1033 calculates the state-action value of the target action in the UAV communication network based on the Q-value function, the reward under the target action, and the state transition probability, and optimizes the Q-value function based on the state-action value of the target action.
[0089] In this embodiment, after selecting a target action and calculating the state, reward, and state transition probability of the UAV communication network under the target action, the cumulative reward under the target action is obtained. Where γ∈(0,1] is the reward coefficient. Given a policy π, which is the mapping between the target action and the state under the target action, the state-action value function Q of policy π is calculated based on the Q-value function. π (s t ,a t ) = E π [R t |s t =s,a t =a].
[0090] In each iteration, the agent stores the state before selecting the target action, the state under the target action, the selected target action, and the reward value obtained from interacting with the environment into an experience pool, providing training samples for network training. After the agent interacts with the environment a set number of times, it learns by replaying the experiences stored in the experience pool, aiming to minimize the loss function. It updates the parameters of the policy network and value network of reinforcement learning using backpropagation gradients, where the value network includes a Q-value function. After the agent interacts with the environment a set number of times, it randomly draws a tuple of size N from the experience pool (s...). i ,a i ,r i+1 ,s i+1 ) i∈N With the goal of minimizing the loss function, the back gradient of the loss function is calculated, and the parameters of the continuous-SAC and discrete-SAC networks are updated. If the training iterations do not reach the set number, the process returns to the agent to iteratively select and optimize the UAV trajectory, UAV transmit power, number of semantic symbols, and phase shift of the intelligent reflector using the H-SAC algorithm, and then inputs these parameters into the UAV communication network.
[0091] This embodiment continuously explores the environment—namely, the state space, action space, reward function, and state transition probabilities—using the Bellman equation and leveraging existing knowledge in the experience pool. It gradually adjusts the Q-value, iteratively updating and refining the Q-value function to enable it to more effectively learn the optimal policy, as shown in the formula.
[0092] As shown. Accordingly, the formula
[0093] It can be mathematically equivalent to the Bellman optimality equation. This embodiment utilizes the H-SAC algorithm to simultaneously process continuous high-dimensional action space and accurately estimate Q-values to determine the target strategy of the UAV communication network. Based on the target strategy, it optimizes UAV trajectory, UAV transmission power, number of semantic symbols, and phase shift of intelligent reflectors to maximize the average secure semantic transmission rate and the average user semantic transmission rate.
[0094] S1034, Select the target action corresponding to the maximum state action value as the target optimization strategy for the UAV communication network.
[0095] S104 optimizes the resource allocation of the UAV communication network based on the target optimization strategy and resource allocation constraints.
[0096] Specifically, the drone trajectory includes the drone's initial position, collision avoidance position, and speed.
[0097] Specifically, resource allocation constraints include the first constraint, the second constraint, the third constraint, the fourth constraint, the fifth constraint, and the sixth constraint.
[0098] The first constraint is the starting point position of the UAV, q[0] = q initial ,
[0099] The second constraint is speed.
[0100] The third constraint is the collision avoidance position.
[0101] The fourth constraint is the UAV's transmit power.
[0102] The fifth constraint is the phase shift θ of the smart reflector. m ∈{0,2π},
[0103] The sixth constraint is the number of semantic symbols: 0 ≤ k ≤ k max Where q[0] represents the starting position of the drone, q initial V represents the preset starting point position, q[n+1] represents the position of the drone at time n+1, q[n] represents the position of the drone at time n, and V min V represents the minimum speed of the drone. max Let δ represent the maximum speed of the drone, and let R = {1, ..., R}. Indicates the first The location of the intelligent reflective surface, P u Indicates the drone's transmission power. θ represents the maximum transmit power of the drone. m Indicates intelligent reflective surface m The phase shift, M represents the set of intelligent reflective surfaces, k represents the number of semantic symbols, k max Indicates the maximum number of semantic symbols.
[0104] In this embodiment, if the target optimization strategy is to adjust the drone trajectory, then according to q[0]=q initial , Given the drone's current position at time T, determine the drone's position after time T. If the objective optimization strategy is to adjust the drone's transmission power, then based on... Determine the UAV's transmit power after the current time T. If the target optimization strategy is to adjust the intelligent reflector, then based on θ... m ∈{0,2π}, Determine the phase shift of the UAV's intelligent reflector after the current time T. If the objective optimization strategy is to adjust the number of semantic symbols, then according to 0≤k≤k... max This embodiment determines the number of semantic symbols for each user after the current time T in the UAV communication network. By adjusting the number of semantic symbols for each user in the UAV communication network, this embodiment can effectively improve communication efficiency, enhance the accuracy of information transmission, improve network capacity management, strengthen the system's anti-interference capability, and promote intelligent decision-making and multi-user collaboration.
[0105] This embodiment provides an optimization method for UAV communication networks. By acquiring the resource configuration and average secure semantic transmission rate of the UAV communication network in real time, and based on the average user semantic rate and the demand rate threshold, it determines whether the resource configuration of the UAV communication network needs optimization. If optimization is required, a target optimization strategy for the UAV communication network is determined based on a reinforcement learning environment, the H-SAC algorithm, and the resource configuration. Based on the target optimization strategy and resource configuration constraints, the resource configuration of the UAV communication network is optimized. Therefore, this invention can achieve significant and effective optimization of UAV communication networks, reducing low optimization efficiency and latency, improving the timeliness and effectiveness of semantic transmission rate optimization, reducing the risk of eavesdropping, reducing semantic leakage problems, and improving the security and reliability of user UAV semantic transmission. Furthermore, based on a reinforcement learning environment and the H-SAC algorithm, resource configurations where the average secure semantic transmission rate is lower than the demand rate threshold are optimized. A reasonable trade-off and optimization is made between algorithm complexity, application scenarios, and UAV hardware capabilities to ensure the effectiveness and reliability of UAV communication while improving the average secure semantic transmission rate and average user semantic transmission rate of the UAV communication network. The reward function design considers both the average secure semantic transmission rate and the average user semantic transmission rate, ensuring a high secure semantic transmission rate while maximizing the user semantic transmission rate to improve user experience. A semantic communication simulation model is established using DeepSC to predict the second text transmission data between users and eavesdroppers, improving the prediction accuracy of text transmission data and enhancing the accuracy of the average secure semantic transmission rate and average user semantic transmission rate of the UAV communication network. By constructing a mapping relationship between signal-to-noise ratio (SNR) and semantic transmission rate and importing this relationship into a reinforcement learning environment, the interaction rate between the agent and the environment is improved, increasing the efficiency of determining the target optimization strategy for the UAV communication network. By utilizing the H-SAC algorithm's ability to simultaneously process continuous high-dimensional action spaces and accurately estimate Q-values, the target strategy for the UAV communication network is determined. Based on this target strategy, the UAV trajectory, UAV transmission power, number of semantic symbols, and phase shift of the intelligent reflector are optimized to maximize the average secure semantic transmission rate and average user semantic transmission rate. By adjusting the number of semantic symbols for each user in an unmanned aerial vehicle (UAV) communication network, communication efficiency can be effectively improved, the accuracy of information transmission can be enhanced, network capacity management can be improved, the system's anti-interference capability can be strengthened, and intelligent decision-making and multi-user collaboration can be promoted.
[0106] Example 2:
[0107] like Figure 3As shown, this embodiment provides an optimization device for an unmanned aerial vehicle (UAV) communication network, including: an acquisition module 21, a judgment module 22, a determination module 23, and an optimization module 24. The acquisition module 21 is used to acquire the resource configuration and average security semantic transmission rate of the UAV communication network at the current time T. The judgment module 22, connected to the acquisition module 21, is used to determine whether the average security semantic transmission rate at the current time T is less than the required rate threshold. The determination module 23, connected to the judgment module 22, is used to determine the target optimization strategy of the UAV communication network based on the reinforcement learning environment, the H-SAC algorithm, and the resource configuration at the current time T in response to the average security semantic transmission rate at the current time T being less than the required rate threshold. The optimization module 24, connected to the determination module 23, is used to optimize the resource configuration of the UAV communication network based on the target optimization strategy and resource configuration constraints.
[0108] Optionally, the optimization device for the UAV communication network further includes: a construction module 25 for constructing a reinforcement learning environment for the UAV communication network based on a Markov decision process.
[0109] Specifically, the determining module 23 includes: a generation unit 231, a selection unit 232, a calculation unit 233, and a selection unit 234. The generation unit 231 is used to generate the initial state of the UAV communication network based on the resource configuration at the current time T. The selection unit 232 is used to select a target action and calculate the state, reward, and state transition probability of the UAV communication network under the target action. The calculation unit 233 is used to calculate the state action value of the target action of the UAV communication network based on the Q-value function, the reward, and the state transition probability under the target action, and optimize the Q-value function based on the state action value of the target action. The selection unit 234 is used to select the target action corresponding to the maximum state action value as the target optimization strategy of the UAV communication network.
[0110] Specifically, the generation unit 231 further includes: a first calculation subunit, a second calculation subunit, a matching subunit, a third calculation subunit, and a generation subunit. The first calculation subunit is used to calculate the channel factor of user i and the eavesdropper corresponding to user i at the current time T based on the UAV flight trajectory at the current time T and the phase shift of the smart reflector corresponding to user i. The second calculation subunit is used to calculate the signal-to-noise ratio of user i and the eavesdropper corresponding to user i at the current time T based on the channel factor at the current time T and the UAV transmission power. The matching subunit is used to match the semantic transmission rate of user i and the eavesdropper corresponding to user i at the current time T based on the mapping relationship between the signal-to-noise ratio and the semantic transmission rate and the signal-to-noise ratio at the current time T. The third calculation subunit is used to calculate the average secure semantic transmission rate and the average user semantic transmission rate of the UAV communication network at the current time T based on the semantic transmission rate at the current time T. The generation subunit is used to set the target action of the UAV communication network at time T-1 as empty and the report of the UAV communication network at the current time T as empty.
[0111] Optionally, the generation unit 231 further includes: an acquisition subunit, a construction subunit, a fourth calculation subunit, a fifth calculation subunit, and an association subunit. The acquisition subunit is used to acquire the first text transmission data of user i, the channel bandwidth and text transmission requirements, and the true signal-to-noise ratio of user i and the eavesdropper corresponding to user i. The construction subunit constructs a semantic communication simulation model based on the DeepSC algorithm of deep learning spectral clustering, and predicts the second text transmission data after semantic communication between user i and the eavesdropper corresponding to user i based on the semantic communication simulation model and the first text transmission data. The fourth calculation subunit calculates the semantic similarity of semantic communication between user i and the eavesdropper corresponding to user i based on the first text transmission data and the second text transmission data. The fifth calculation subunit calculates the semantic transmission rate of user i and the eavesdropper corresponding to user i based on the semantic similarity, the channel bandwidth, and the number of semantic symbols of user i. The association subunit associates the semantic transmission rate and the true signal-to-noise ratio to obtain the mapping relationship between the signal-to-noise ratio and the semantic transmission rate.
[0112] Specifically, the fifth calculation subunit includes: a minimum calculation unit, used to calculate the semantic transmission rate of user i and the corresponding eavesdropper of user i according to formulas (1) and (2):
[0113]
[0114] Where W is the channel bandwidth, I is the expected amount of semantic information, and k i Let L represent the number of semantic symbols for user i, L represent the expected number of words in the text, and φ represent the number of semantic symbols for user i. i φ represents the semantic similarity between user i and user i. EAV This represents the semantic similarity between user i and the eavesdropper.
[0115] Understandably, the above-described optimization device for a UAV communication network executes the optimization method for the UAV communication network corresponding to Embodiment 1 provided above. Therefore, the beneficial effects it can achieve can be referred to the beneficial effects of the scheme corresponding to the optimization method for the UAV communication network in Embodiment 1 above, which will not be repeated here.
[0116] Example 3:
[0117] This embodiment also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to implement the optimization method for the UAV communication network in Embodiment 1 above.
[0118] Example 4:
[0119] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the optimization method for the UAV communication network in Embodiment 1 above.
[0120] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of the present invention, and the present invention is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A method for optimizing a UAV communication network, the method comprising: The method comprises: acquiring resource configuration and average safe semantic transmission rate of the UAV communication network at the current time T; judging whether the average safe semantic transmission rate at the current time T is less than a demand rate threshold; in response to the average safe semantic transmission rate at the current time T being less than the demand rate threshold, determining a target optimization strategy of the UAV communication network based on a reinforcement learning environment, a hybrid flexible action-evaluation H-SAC algorithm and the resource configuration at the current time T; optimizing the resource configuration of the UAV communication network based on the target optimization strategy and resource configuration constraints.
2. The method of claim 1, wherein, After the judgment of whether the average safe semantic transmission rate at the current time T is less than the demand rate threshold, and before the determination of the target optimization strategy of the UAV communication network based on the reinforcement learning environment, the hybrid flexible action-evaluation H-SAC algorithm and the resource configuration at the current time T, the method further comprises: constructing a reinforcement learning environment of the UAV communication network based on a Markov decision process, the reinforcement learning environment comprises an agent, a state space, an action space, a reward function and a state transition probability, the agent is a control center of the UAV communication network, state space Action space a t = {{q t}, Θ μ(i) , P i,t , k i,t}, Return function State transition probabilities wherein, denotes the channel factors of all users and eavesdroppers in the UAV communication network at time t, denotes the average secure semantic transmission rate of the UAV communication network at time t, a t-1 denotes the selected target action of the UAV communication network at time t-1, r t denotes the reward value of the UAV communication network at time t, {q t} denotes adjusting the UAV flight trajectory at time t, Θ μ(i) denotes adjusting the phase shift of the smart reflector corresponding to user i, P i,t denotes adjusting the transmission power of the UAV to user i at time t, k i,t denotes adjusting the number of semantic symbols of user i at time t, denotes the secure semantic transmission rate of user i in the UAV communication network, a denotes the average user semantic transmission rate of the UAV communication network, Γ i denotes the semantic transmission rate of user i in the UAV communication network, Γ EAV denotes the semantic transmission rate of the eavesdropper corresponding to user i in the UAV communication network, D denotes the total number of users in the UAV communication network, s t+1 denotes the state of the UAV communication network at time t+1, s t denotes the state of the UAV communication network at time t, a t denotes the selected target action of the UAV communication network at time t, i = 1, 2, …, D, D denotes the total number of users in the UAV communication network.
3. The method of claim 2, wherein, the determination of the target optimization strategy of the UAV communication network based on the reinforcement learning environment, the hybrid flexible action-evaluation H-SAC algorithm and the resource configuration at the current time T specifically comprises: generating an initial state of the UAV communication network based on the resource configuration at the current time T; selecting a target action and calculating the state, the reward and the state transition probability of the UAV communication network under the target action; calculating the state-action value of the target action of the UAV communication network based on the Q value function, the reward under the target action and the state transition probability, and optimizing the Q value function based on the state-action value of the target action; selecting the target action corresponding to the maximum state-action value as the target optimization strategy of the UAV communication network.
4. The method of claim 3, wherein, the resource configuration comprises a UAV flight trajectory, a UAV transmission power, a semantic symbol quantity of each user and a phase shift of the smart reflecting surface corresponding to each user, the generation of the initial state of the UAV communication network based on the resource configuration at the current time T specifically comprises: calculating the channel factor of user i and the eavesdropper corresponding to user i at the current time T based on the UAV flight trajectory at the current time T and the phase shift of the smart reflecting surface corresponding to user i; calculating the signal-to-noise ratio of user i and the eavesdropper corresponding to user i at the current time T based on the channel factor at the current time T and the UAV transmission power; matching the semantic transmission rate of user i and the eavesdropper corresponding to user i at the current time T based on the mapping relationship between the signal-to-noise ratio and the semantic transmission rate and the signal-to-noise ratio at the current time T; calculating the average safe semantic transmission rate and the average user semantic transmission rate of the UAV communication network at the current time T based on the semantic transmission rate at the current time T; the target action of the UAV communication network at T-1 time is empty, and the reward of the UAV communication network at the current time T is empty.
5. The method of claim 4, wherein, Before the calculating the signal-to-noise ratio of the user i and the eavesdropper corresponding to the user i at the current time T according to the channel factor at the current time T and the unmanned aerial vehicle transmission power, and before the matching the semantic transmission rate of the user i and the eavesdropper corresponding to the user i at the current time T according to the mapping relationship between the signal-to-noise ratio and the semantic transmission rate and the signal-to-noise ratio at the current time T, the method further comprises: obtaining the first text transmission data of the user i, the channel bandwidth, the text transmission demand of the user i, and the real signal-to-noise ratio of the user i and the eavesdropper corresponding to the user i; constructing a semantic communication simulation model based on a deep learning-based spectral clustering DeepSC algorithm, and predicting the second text transmission data of the user i and the eavesdropper corresponding to the user i after semantic communication based on the semantic communication simulation model and the first text transmission data; calculating the semantic similarity of the semantic communication of the user i and the eavesdropper corresponding to the user i according to the first text transmission data and the second text transmission data; calculating the semantic transmission rate of the user i and the eavesdropper corresponding to the user i based on the semantic similarity, the channel bandwidth, and the number of semantic symbols of the user i; obtaining the mapping relationship between the signal-to-noise ratio and the semantic transmission rate based on the semantic transmission rate and the real signal-to-noise ratio.
6. The method of claim 5, wherein, The text transmission demand comprises an expected number of semantic information and an expected number of words, The calculating the semantic transmission rate of the user i and the eavesdropper corresponding to the user i based on the semantic similarity, the channel bandwidth, and the number of semantic symbols of the user i specifically comprises: calculating the semantic transmission rate of the user i and the eavesdropper corresponding to the user i according to formulas (1) and (2): where W is the channel bandwidth, I is the expected number of semantic information, k i is the number of semantic symbols of user i, L represents the expected number of words of text, i i represents the semantic similarity of user i, i EAV represents the semantic similarity of user i to the eavesdropper.
7. The method of claim 1, wherein, The unmanned aerial vehicle trajectory comprises an initial position of the unmanned aerial vehicle, a collision avoidance position, and a speed, The resource configuration constraints comprise a first constraint, a second constraint, a third constraint, a fourth constraint, a fifth constraint, and a sixth constraint, The first constraint is that the starting point position of the UAV q[0] = q initial , The second constraint is speed The third constraint is a collision avoidance position The fourth constraint is the UAV launch power The fifth constraint is the phase shift of the smart reflector The sixth constraint is the semantic symbol number 0≤k≤k max , wherein q[0] represents a starting point position of the UAV, q initial represents a preset starting point position, q[n+1] represents a position of the UAV at n+1 time, q[n] represents a position of the UAV at n time, V min represents a minimum speed of the UAV, V max represents a maximum speed of the UAV, δ represents a time consumed by the UAV from the position at n+1 time to the position at n time, R={1,…,R}, represents a position of the i th intelligent reflecting surface, P u represents a transmission power of the UAV, represents a maximum transmission power of the UAV, θ m represents a phase shift of the intelligent reflecting surface m , M represents a set of the intelligent reflecting surfaces, k represents a semantic symbol number, and k max represents a maximum semantic symbol number.
8. An optimization device for a drone communication network, characterized in that, comprise: an obtaining module, a judging module, a determining module, and an optimizing module, The obtaining module is configured to obtain the resource configuration and the average safe semantic transmission rate of the unmanned aerial vehicle communication network at the current time T. The judging module is connected with the obtaining module and is configured to judge whether the average safe semantic transmission rate at the current time T is less than a demand rate threshold. The determining module is connected with the judging module and is configured to, in response to the average safe semantic transmission rate at the current time T being less than the demand rate threshold, determine a target optimization strategy of the unmanned aerial vehicle communication network based on a reinforcement learning environment, an H-SAC algorithm, and the resource configuration at the current time T.
9. An electronic device, comprising: The optimizing module is connected with the determining module and is configured to optimize the resource configuration of the unmanned aerial vehicle communication network based on the target optimization strategy and the resource configuration constraints.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the optimization method of the unmanned aerial vehicle communication network according to any one of claims 1 to 7. The computer program is executed by the processor to implement the optimization method of the unmanned aerial vehicle communication network according to any one of claims 1 to 7.