Unmanned aerial vehicle emergency communication spectrum scheduling method and system based on dynamic entropy optimization

Through the dynamic entropy optimization of the UAV emergency communication spectrum scheduling method, the DSA and ME-MAAC algorithms are used to optimize spectrum perception and aggregation, which solves the problems of scarce spectrum resources and multi-user competition in disaster relief, and achieves efficient and reliable spectrum utilization and communication guarantee.

CN120769244APending Publication Date: 2025-10-10SHANDONG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511062534.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In disaster relief scenarios, drone emergency communication networks face the problems of scarce spectrum resources, limited energy, and dynamic competition among multiple users, resulting in inefficient spectrum utilization and poor communication reliability.

Method used

A spectrum scheduling method for UAV emergency communications based on dynamic entropy optimization is adopted. Through an intelligent distributed spectrum management mechanism, DSA and ME-MAAC algorithms are used for spectrum perception and aggregation. The spectrum access strategy is optimized in combination with a deep reinforcement learning model to ensure the priority of primary users and efficient access of secondary users.

Benefits of technology

It significantly improves spectrum utilization efficiency, reduces conflict rates, ensures high quality and reliability of rescue communications, meets the differentiated bandwidth requirements of multiple users, and adapts to dynamic spectrum environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769244A_ABST
    Figure CN120769244A_ABST
Patent Text Reader

Abstract

The invention relates to an unmanned aerial vehicle emergency communication spectrum scheduling method and system based on dynamic entropy optimization, and the method comprises the steps: 1, constructing an emergency communication system, and calculating the gain of a received signal; 2, constructing a spectrum sensing and user access model; step 3, evaluating the channel quality in the emergency communication system; 4, establishing a state, an action and a reward of deep reinforcement learning according to spectrum sensing, a user access model and channel quality evaluation; 5, constructing a deep reinforcement learning model, and performing training to obtain a trained deep reinforcement learning model; and step 6, obtaining an optimal strategy of unmanned aerial vehicle emergency communication spectrum scheduling in a spectrum environment by using the trained deep reinforcement learning model, and realizing multi-user aggregation spectrum access. According to the invention, through intelligent distributed spectrum management, the problems of limited spectrum resources and energy, multi-user dynamic competition and the like in a disaster rescue scene are solved, and the spectrum utilization efficiency and the communication reliability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to an unmanned aerial vehicle emergency communication spectrum scheduling method and system based on dynamic entropy optimization, and belongs to the technical field of wireless communication. BACKGROUND

[0002] In recent years, natural disasters have occurred frequently, leading to the vulnerability of traditional communication networks, and the emergency communication system is facing severe challenges. Unmanned aerial vehicles (UAVs) are widely used in the construction of rescue cognitive networks (RCN) due to their flexible deployment capabilities, but they have three major core bottlenecks: the limited battery capacity (generally less than 2 hours), which leads to insufficient spectrum access service duration; the available spectrum resources are concentrated in narrowband licensed frequency bands such as 2.4 / 5.8 GHz, which cannot meet the multi-service needs of heterogeneous users (rescue personnel and disaster-stricken people); and the competitive access of primary users (PUs) and secondary users (SUs) in dynamic spectrum environments leads to a conflict rate of more than 15%, severely affecting transmission reliability. Traditional cognitive radio technology alleviates resource shortages through static spectrum allocation and preemptive access strategies, but the preset fixed priority model has poor real-time dynamic adaptability in disaster scenarios, and the spectrum aggregation efficiency is limited by hardware heterogeneity (the difference in sensing accuracy is up to 40%).

[0003] In terms of dynamic spectrum access mechanisms, domestic and foreign researches mainly focus on single-user reinforcement learning frameworks. The DQN channel selection algorithm proposed by a team from the University of Texas in the United States optimizes single-SU throughput through value function iteration, but its strategy convergence rate is lower than 0.8 episode / s, and the conflict probability increases to more than 35% in multi-user scenarios. The DDQN architecture developed by the Swiss Federal Institute of Technology in Lausanne reduces estimation errors by decoupling action selection and evaluation, but the training time increases by 170% in a 32-channel system. To address the multi-user conflict problem, the Technion-Israel Institute of Technology first proposed the distributed MA-DQN scheme, which uses priority experience replay to enable agent collaboration, but its fixed-length observation window cannot adapt to time-dependent channels (decision error rate increases by 22% when Markov transition probability is greater than 0.7).

[0004] In the field of spectrum aggregation technology, the D-OFDM aggregation model built by a university team improves throughput by 18.7% through non-contiguous subcarrier binding, but it relies on ideal channel state information, and the bit error rate significantly deteriorates in 3D mobile scenarios. The DISCO project funded by the FP7 program developed an aliasing channel splicing algorithm, but it did not consider the nonlinear characteristics of hardware transceivers (phase noise tolerance -125 dBc / Hz), resulting in a high actual deployment failure rate of 41%. The winning solution of the SC2 competition of DARPA uses game theory to optimize multi-user aggregation strategies, but the Nash equilibrium convergence time exceeds 5 minutes in a 30-node network, which cannot meet the real-time needs of emergency rescue.

[0005] The current technical pain points are that the strategy exploration efficiency defect leads to high local optimal probability, the dynamic cooperation capability defect causes the estimation deviation amplification, the spectrum-energy joint optimization defect leads to low energy efficiency, and the hardware adaptability limitation causes the pseudo idle channel access rate to be too high. Research shows that the full spectrum utilization rate, cooperation efficiency and energy consumption control comprehensive score of the existing system are only at a suboptimal level, and it is urgent to improve the heterogeneous hardware dynamic aggregation mechanism, the robust search strategy based on environmental entropy and the cross-layer joint optimization technology, so as to realize the efficient and reliable operation of the rescue network. SUMMARY

[0006] In view of the shortcomings of the prior art, the present application provides a kind of unmanned aerial vehicle emergency communication spectrum scheduling method and system based on dynamic entropy optimization, specially designed for unmanned aerial vehicle rescue cognitive network;This method solves the key technical problems such as spectrum resource scarcity, energy limitation and multi-user dynamic competition in disaster rescue scene through intelligent distributed spectrum management mechanism, significantly improves the system spectrum utilization efficiency and communication reliability.

[0007] The present application proposes an emergency communication network model in extreme conditions, in which the communication equipment in the disaster area is severely damaged, resulting in complete interruption of the communication network, and the disaster area needs to be configured with mobile base stations to build a temporary communication network. Therefore, a UAV-assisted rescue cognitive network model is designed, in which the UAV carrying the mobile base station can provide spectrum access for disaster area users through satellite link and ground control station. All users in the disaster area can be generally divided into two categories: disaster relief personnel (i.e. primary users) and disaster victims (i.e. secondary users). Assuming that there are M primary users and N secondary users in the disaster area, the primary users as the main force of rescue need to communicate with the ground control center in the form of voice, image, video, etc., i.e. the primary users have higher bandwidth requirements. The secondary users with lower communication requirements need to access the dynamically changing spectrum without conflict and conduct emergency communication through text, voice, etc. Due to the destruction of infrastructure in the disaster area, the spectrum carried by the UAV is in a completely idle state. In order to ensure the high-quality communication of the rescue team, the primary users have higher access priority.

[0008] Term explanation: 1、UAV-RCNs: UAV-assisted Rescue Cognitive Networks, UAV-assisted Rescue Cognitive Networks is the core communication node architecture proposed in the present application, which specifically refers to a mobile emergency communication base station formed by a multi-rotor unmanned aerial vehicle carrying communication equipment.

[0009] 2、DSA: Dynamic Spectrum Aggregation, DSA mechanism perceives idle channels through energy detection, supports intelligent aggregation of up to 4 non-contiguous subcarriers, and adopts Q-learning to optimize access strategy. In a test of 33 subchannels, a spectrum efficiency of 40 Kbps / MHz is achieved, the conflict rate is reduced to 0.13%, and the 50ms low-latency priority access of rescue communication (PU) is ensured.

[0010] 3、Primary User (PU) and Secondary User (SU): In the spectrum sharing scenario, the primary user refers to the user who has priority access to the spectrum, usually referring to critical rescue personnel or important communication subjects such as command centers. The secondary user refers to the ordinary user who needs to dynamically perceive idle channels and access them. In this invention, the distribution of PU and SU is modeled according to the actual application scenario, and the SU needs to dynamically select available spectrum resources for data transmission without interfering with PU communication.

[0011] 4、ME-MAAC algorithm: Maximum Entropy Multi-Agent Actor-Critic algorithm, which enhances the exploration ability of the strategy by introducing an entropy regularization term, avoiding premature convergence of the model to a local optimum. In this invention, each SU is assigned an agent, which selects actions (spectrum segments) through the Actor network and evaluates the value of the action through the Critic network. In this way, the ME-MAAC algorithm can find the optimal spectrum selection strategy for the SU, thereby maximizing the long-term cumulative reward of the SU while ensuring PU communication.

[0012] 5、Entropy regularization: Entropy regularization is a technique used in reinforcement learning to enhance the exploration of the strategy. By introducing an entropy reward term in the policy gradient update, the randomness of the strategy is increased, thereby avoiding premature convergence of the model to a local optimum. In this invention, the ME-MAAC algorithm introduces an entropy regularization term to adjust the weight of the entropy ( =0.2), ensuring a balance between exploration and exploitation.

[0013] 6、OFDM: Orthogonal Frequency Division Multiplexing (OFDM) is a highly efficient digital modulation technique that divides the communication spectrum into multiple independent subchannels, each carrying different data symbols.

[0014] 7、Experience replay buffer: Experience replay buffer is a data structure used to store training samples for offline training in reinforcement learning. In the present invention, the ME-MAAC algorithm stores experience unit groups such as observation state, action selection, reward value and new observation state of SU through the experience replay buffer. During the training process, the system randomly samples from the experience replay buffer to break data correlation, improve learning efficiency and stability.

[0015] 8、Line of Sight (LoS): In wireless communication, it refers to the propagation path from the transmitter to the receiver without any obstacles blocking the signal, which can directly reach the receiver, usually providing better transmission quality and higher data rate.

[0016] 9、Non-Line of Sight (NLoS): In wireless communication, it refers to the propagation path from the transmitter to the receiver with obstacles, causing the signal to reach the receiver indirectly through reflection, scattering, etc. In this case, the signal quality may be significantly affected, such as attenuation and delay increase.

[0017] The technical solutions of the present invention are as follows: The first aspect of the present invention provides a dynamic entropy optimization-based unmanned aerial vehicle emergency communication spectrum scheduling method, comprising: Step 1: Construct an emergency communication system for user and mobile base station communication and calculate the received signal gain; Step 2: Construct a spectrum sensing and user access model to identify available channels and develop user access strategies; Step 3: Evaluate the channel quality in the emergency communication system; Step 4: Establish the state, action and reward of deep reinforcement learning based on the spectrum sensing and user access model and the evaluation of channel quality; Step 5: Construct a deep reinforcement learning model and train it to obtain a trained deep reinforcement learning model; Step 6: Use the trained deep reinforcement learning model to obtain the optimal strategy for unmanned aerial vehicle emergency communication spectrum scheduling in the spectrum environment, realizing multi-user aggregated spectrum access.

[0018] According to the present invention, the emergency communication system is constructed for user and mobile base station communication, and the received signal gain is calculated; comprising: The emergency communication system includes unmanned aerial vehicles and mobile base stations carried by unmanned aerial vehicles, and the unmanned aerial vehicles communicate with users through air-to-ground channels; the users include primary users (disaster relief personnel) and secondary users (disaster victims); The UAV in the emergency communication system hovers at a fixed height h and flies along a fixed circular trajectory at a constant speed v; the three-dimensional coordinates of the UAV, primary user and secondary user are set as (x, y, h), ( , ,0) and (x n , y n ,0); then the distances between the UAV and the mth primary user and the nth secondary user are as follows: (1); (2); Among them, d m represents the distance between the UAV and the mth primary user, d n represents the distance between the drone and the nth user, x represents the horizontal coordinate of the drone, y represents the vertical coordinate of the drone, and x m Indicates the horizontal coordinate of the primary user, y m Indicates the vertical coordinate of the primary user, x n Indicates the horizontal coordinate of the secondary user, y n represents the ordinate of the secondary user; compared with ground base stations, UAVs have a higher line of sight (LoS) link; the present invention calculates the path loss by adopting the air-to-ground (A2G) channel model and ignoring the influence of small-scale fading; it is assumed that there are LoS links, i.e., line-of-sight links, and NLoS links, i.e., non-line-of-sight (NLoS) links, with a certain probability between the UAV and all users; the path losses of the LoS link and the NLoS link are defined as shown in equations (3) and (4), respectively: (3); (4); in, represents the path loss of the line-of-sight link, represents the path loss of the non-line-of-sight link, represents the unique carrier frequency of each channel, i.e., the air-to-ground channel, and d is the distance between the UAV and the user. To simplify the calculation, it is assumed that the distance d is fixed when the user accesses the spectrum (the radio frequency resources provided by the mobile base station carried by the UAV); represents the speed of light, and The additional losses of LoS and NLoS links respectively, and their magnitude depends on the environmental parameters of the urban environment; the spectrum configured by the drone includes multiple independent channels, each with a unique carrier frequency ; Assume that the probability of the LoS link and NLoS link between users is expressed as: (5); (6); in, represents the probability of the line-of-sight link appearing, represents the probability of non-line-of-sight links; a and b are parameters determined by the urban environment; the total path loss formula between the drone and each user is obtained from the above formulas (5) and (6), and its value depends on the distance d and the carrier frequency , as shown in formula (7): (7); in, represents the total path loss between the UAV and each user; In order to enhance the robustness of channel estimation, the Rice channel model is introduced to describe small-scale fading, and the received signal gain is calculated using the Rice channel model as shown below: (8); Where g represents the received signal gain, Indicates the remaining energy level of the signal after path loss. ,Right now Determined by path loss; is the Rice factor, which represents the energy ratio between the LoS link component and the scattered component (the signal component caused by all reflections and diffraction except the direct path); the phase corresponding to the LoS link component follows a uniform distribution, and represents a complex Gaussian random variable corresponding to the NLoS link component. The Complex Normal Distribution is used to describe random variables with real and imaginary parts. represents the complex exponential function, where e is the base of the natural logarithm (approximately equal to 2.71828), represents the imaginary unit, Represents the phase angle.

[0019] Preferably, according to the present invention, a spectrum sensing and user access model is constructed to identify available channels and formulate user access strategies, including: In emergency communications scenarios, spectrum resources are often extremely scarce. To maximize spectrum utilization efficiency, a dynamic spectrum aggregation and access control system has been designed. This system uses a slotted frame structure and divides each operating cycle into four phases: First, there is the random access phase, led by primary users (PUs). This phase ensures absolute priority for communications for key users such as disaster relief commanders and emergency doctors. This is followed by the spectrum sensing phase for secondary users (SUs), where ordinary disaster victims can scan for available channels using efficient energy detection technology. Next comes the critical spectrum aggregation phase, where the system automatically merges discrete idle sub-channels into continuous, available frequency bands. Finally, there is the intelligent network update phase, where the system automatically adjusts parameters based on the current operating status to optimize resource allocation strategies for the next operating cycle. The spectrum sensing and user access model includes the random access phase, spectrum sensing phase, spectrum aggregation phase, and network update phase; In the random access phase, M primary users start at the beginning of each time slot based on their own bandwidth requirements. Randomly access multiple channels, namely air-to-ground channels (users access the mobile base station carried by drones through air-to-ground channels), and the spectrum is no longer completely idle due to the occupation of primary users; In the spectrum sensing stage, N secondary users observe idle channels through their own spectrum sensing capabilities. The accuracy of spectrum sensing is the basis for reliable operation of the system. Spectrum sensing capabilities include detection probability and false alarm probability; detection probability Indicates the probability of correctly identifying an occupied channel and the false alarm probability The probability of misjudging an idle channel as an occupied channel; the spectrum sensing capability of the secondary user is expressed in spectrum sensing accuracy. express, The higher it is, the more accurate the secondary user's judgment of the channel status is. The formula is as follows: (9); in, represents the spectrum sensing accuracy of the nth secondary user; In the spectrum aggregation stage, secondary users The observation state, or the estimated state of each channel occupancy obtained through spectrum sensing, is obtained. Based on the spectrum aggregation capability, the limited frequency bands are aggregated into different aggregated bands. However, due to hardware limitations such as transceivers, spectrum aggregation capabilities dictate that secondary users can only aggregate continuous spectrum holes, or spectrum resources not currently occupied by primary users and available to secondary users, with a fixed aggregation length, L. Secondary users select the appropriate aggregated band for access, and access success is determined by base station interaction information. This includes two scenarios: successful access and failed access. Successful access by a secondary user: Based on the secondary user's correct observation status, when the number of channels in the aggregated frequency band selected by the secondary user is greater than or equal to the secondary user's bandwidth requirement B, and the selected channels do not conflict with any other users, the secondary user will select multiple high-quality channels in the aggregated segment (signal-to-noise ratio greater than 20dB, path loss less than 60dB, gain greater than 10dBi, unoccupied channels, and channels with good historical performance) for access. The mobile base station carried by the drone will then provide feedback to the secondary user, including rewards such as transmission rate, data packet loss rate, signal-to-noise ratio, and current channel load. The secondary user will then send multimedia data (video images, location information of disaster victims, etc.) to the mobile base station carried by the drone. If there are still idle channels in the aggregated frequency band, other users can occupy them for data transmission. Failed access by a secondary user: If the secondary user experiences erroneous perception, the aggregated band does not meet its needs, or the selected channel conflicts, the secondary user will abandon the aggregated band and reselect the spectrum in the next time slot based on the previous experience and update the strategy to retransmit the unsent data. (After a failed access attempt, the secondary user uses a reinforcement learning algorithm to update its behavior strategy based on the experience of the previous time slot, such as perception errors, improper aggregation, and conflicts, to make a better access decision in the next time slot.) In the final stage of the time slot, the access strategy update stage, the user agent (an independent intelligent decision-making agent corresponding to each secondary user) will use environmental information and feedback information from the base station (such as the current channel status and spectrum sensing results) to update the access strategy, continuously optimizing the action strategy so that the strategy is gradually transformed into the optimal strategy.

[0020] Preferably, according to the present invention, evaluating the channel quality in the emergency communication system comprises: The spectrum of the mobile base station carried by the drone is limited. Assuming that the bandwidth of each channel is are the same, the entire spectrum is evenly divided into K channels; at the same time, the transmission power P of each channel is also the same, but the carrier frequency of each channel is unique; therefore, the channel gain of the nth secondary user on the K channels is defined as: (10); in, represents the channel gain of the nth user, represents the channel gain of the nth secondary user on the kth channel; to evaluate the quality performance of the channel, the signal-to-noise ratio and transmission rate are set as the channel quality standards; to make full use of the idle channels in the spectrum, the aggregated frequency band selected by a secondary user contains not only the channels occupied by the primary user, but also the channels accessed by other users; therefore, when there is no conflict between secondary users, the signal-to-noise ratio obtained by the nth secondary user is is defined as: (11); in, is the noise spectral density, is the gain of the ith channel in the segment selected by the nth secondary user, is the gain of the j-th channel of the m-th primary user, Indicates the nth user selection channels, Indicates the mth primary user's selection channels, Indicates the bandwidth requirement of the nth secondary user, N is the number of secondary users, and M is the number of primary users; If there are K channels in the aggregated frequency band selected by the nth secondary user that conflict with other secondary users, the signal-to-noise ratio is represented as: (12); in, represents the gain of the vth conflicting channel between the nth secondary user and the remaining u secondary users, represents the transmission power of the kth user, represents the number of channels in the aggregated frequency band selected by the u-th secondary user that have access conflicts with other secondary users; According to the above SNR formula, the transmission rate of the nth secondary user is Expressed as: (13); Or expressed as: (14); in, is the loss coefficient of the signal-to-noise ratio, which represents the loss between the theoretical channel capacity and a specific modulation and coding scheme; If the signal-to-noise ratio is greater than the set threshold (20dB) and the transmission rate is greater than the set threshold (1 megabit), the channel quality is determined to be high.

[0021] Preferably, according to the present invention, the state, action and reward of deep reinforcement learning are established based on spectrum sensing and user access model and channel quality evaluation; including: Secondary users act as independent intelligent agents. Their state reflects spectrum channel occupancy, including spectrum status and user observation status. Their actions involve selecting a specific aggregated frequency band for access. Rewards are a combination of transmission rate and conflict penalties to encourage users to optimize channel selection. The goal is to maximize long-term discounted rewards. Through policy iteration, secondary users are able to adaptively learn the optimal access strategy in complex spectrum environments. In a complex spectrum environment, each user, acting as an independent intelligent agent, comprehensively considers its own capabilities and bandwidth requirements, monitors the entire spectrum in real time, and selects channels that meet its needs and provide good quality to ensure data transmission for different users. Based on the principles of reinforcement learning algorithms, this paper describes in detail the state, action, and reward-related elements to prepare for the detailed explanation of the algorithm. (1) Spectrum state and user observation state: The spectrum is evenly divided into K independent channels, and the state of the spectrum at time slot t is expressed as ,in Indicates the state of the kth channel in time slot t. Each channel has two states: occupied state (S=1) or idle state ( =0); however, at the initial moment, the spectrum is actually completely idle, and changes in the channel state are caused by the access of the primary user. The secondary user has no idea about the primary user's spectrum usage pattern, so the secondary user needs to learn the primary user's occupancy pattern through spectrum sensing. The observed states of N secondary users are defined as Since the spectrum carried by the drone is very limited, the secondary user no longer performs partial spectrum perception, but directly perceives the entire spectrum; therefore, the observation state of the nth secondary user for the entire spectrum at time slot t is expressed as ,in, represents the state of the kth channel perceived by the nth secondary user in time slot t; (2) Action: The action consists of the set of actions of all secondary users, i.e. , Represents the action of the nth sub-user in time slot t. Each sub-user is regarded as an intelligent agent, according to its own observation state Select actions independently ; There are two situations for the selected action: when the selected frequency band can meet the bandwidth requirement B, and the selected channel does not conflict with the channel occupied by the primary user, the secondary user will select the channel with higher channel gain in the aggregated frequency band to meet its own bandwidth requirement B; otherwise, when the aggregated frequency band cannot meet the requirement, the secondary user will give up accessing the channel; according to the aggregation length of the secondary user ( ), the entire spectrum is divided into K- +1 aggregated frequency band; the secondary user will select the most appropriate aggregated frequency band for access based on the action strategy; (3) Reward: The reward set of N users is expressed as ,in represents the reward of the nth secondary user in time slot t; to encourage the secondary user to choose an appropriate action, a penalty term f (f = -10) is set. When the secondary user chooses an action that does not meet expectations, the mobile base station will feedback the penalty term to remind the secondary user to update the action strategy; the reward function combines the secondary user's transmission rate and the penalty term f, and is divided into four cases. The reward of the nth secondary user in time slot t is expressed as: (15); Since the access scheme cancels the priority setting, the relationship between secondary users is no longer cooperative, but they compete to select high-quality channels; The goal of each sub-user is to find an optimal action strategy that maximizes the long-term cumulative discounted benefit, as shown below: (16); in, ∈[0,1] is the discount rate that measures the impact of future rewards on cumulative rewards, and E[] represents the discount rate in the strategy The expected value of the weighted sum of all future rewards under Indicates that in the strategy The long-term cumulative discount benefit under the The total revenue that can be obtained.

[0022] Preferably, according to the present invention, a deep reinforcement learning model is constructed and trained to obtain a trained deep reinforcement learning model; comprising: Build a deep reinforcement learning model, namely the ME-MAAC algorithm reinforcement learning model, which includes the Actor network (generating action strategies), the Critic network (evaluating Q values), the target network, and the experience replay area; The Actor network outputs discrete actions (Softmax scores) based on the observed state, while the Critic network calculates the action value and introduces an entropy term to enhance exploration. Policy gradient updates combine Q-values ​​and entropy regularization to balance exploration and exploitation. The target network improves stability through soft updates, and experience replay reduces data correlation to accelerate convergence. The network parameters are The Actor network uses the exploration strategy To choose actions, strategies To observe the status For input value, take action is the output value; the action selection function is described as ; Since there is no correlation between channels, the action space is discrete, and the score of each aggregated frequency band is evaluated by using a softmax function in the network; actions with higher scores are more likely to be selected than actions with lower scores; The parameters of the critic network are , the user agent obtains the action and observation status of the current time slot from the Actor network , and input to the Critic network, then the user agent will get the output Q value , to evaluate the action of the current time slot; both the Actor network and the Critic network use the Adam optimizer to optimize the parameters; The ME-MAAC algorithm reinforcement learning model uses a policy gradient method to update the actor network and a loss function to update the critic network. Because all secondary users need to perceive the entire spectrum to obtain the observed states of all channels and select an appropriate spectrum segment, their state and action spaces have the same size, so all user agents have the same network structure. The policy gradient exists in the Actor network and is used to update the Actor network. When training the network, each secondary user, as an intelligent agent, randomly samples a small batch of observation state and action data from the experience replay area, and then transmits this data to the Actor network and the Critic network to calculate the Q value respectively. and action-value function ; In order to add some randomness when exploring actions and avoid converging to non-optimal deterministic policies, a maximum entropy method is introduced in the ME-MAAC algorithm, that is, an entropy term that supports random policies is added to the policy gradient and loss function. , then the policy gradient of the nth user is expressed as: (17); in, represents the policy gradient of the nth user, Indicates network parameters The derivative operation of is the control factor that controls the randomness of the optimal strategy, D represents the experience replay area, Indicates the state sampled from the experience playback area D and actions The expected value of Indicates that in a given state Take action The Q value is estimated by the Critic network with parameter θ; The nth user samples a certain number of experience units , respectively, through the Critic network and its target network to obtain the Q value and target Q value ,in Represents the next observed state, and then adds the target Q value and the entropy term obtained from the target Actor network Combined to calculate the output value of the target network , the secondary user updates the Critic network by minimizing the regression loss function, which is defined as: (18); (19); in, represents the loss function of the Critic network of the nth user, Represents the experience unit sampled from the experience playback area D The expected value of Indicates that given the next observation state Next, the action distribution generated by the target Actor network Action expected value; The network structure of the target network is the same as that of the source network, namely the Actor network and the Critic network, and the parameters of the target Actor network and the target Critic network in the target network are and Frequently updating the Q value may lead to unstable training. Therefore, the target network soft-updates its own parameters by copying a small part of the source network parameters to enhance the update stability of the network. The update formula of the target network parameters is expressed as: (20); (twenty one); Where τ∈[0,1] is the soft update coefficient; In each time slot, the user integrates all the data obtained from interacting with the environment, namely the observed state, action and reward, into a set of experience units. , and stored in the experience replay area D; during each network training process, the secondary user randomly extracts experience units to update the network parameters; this is because the continuous data obtained by the user's interaction with the environment has a certain correlation, but the data used by the deep neural network must be independent and identically distributed. Experience replay helps to reduce the correlation of data samples and improve learning efficiency and convergence speed; through continuous updating of training by the target network and the source network, a trained deep reinforcement learning model is obtained.

[0023] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a method for scheduling spectrum of unmanned aerial vehicle emergency communications based on dynamic entropy optimization are implemented.

[0024] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for scheduling spectrum of unmanned aerial vehicle emergency communications based on dynamic entropy optimization.

[0025] A second aspect of the present invention provides a spectrum scheduling system for UAV emergency communications based on dynamic entropy optimization, comprising: The emergency communication system building module is configured to: build an emergency communication system, perform communication between users and mobile base stations, and calculate received signal gain; The spectrum sensing and user access model building module is configured to: build spectrum sensing and user access models, identify available channels, and formulate user access policies; The channel quality assessment module is configured to: assess the channel quality in the emergency communication system; A reinforcement learning parameter establishment module is configured to establish states, actions, and rewards for deep reinforcement learning based on spectrum sensing and user access models and channel quality assessments; The reinforcement learning model building module is configured to: build a deep reinforcement learning model and perform training to obtain a trained deep reinforcement learning model; The spectrum aggregation module is configured to use the trained deep reinforcement learning model to obtain the optimal strategy for spectrum scheduling of drone emergency communications in the spectrum environment and achieve aggregated spectrum access for multiple users.

[0026] The beneficial effects of the present invention are: 1. Efficient emergency communication deployment: Drones equipped with mobile base stations quickly established a temporary communication network via satellite links, resolving communication disruptions caused by damaged infrastructure in disaster areas. Fixed-altitude hovering combined with an optimized path loss model ensured line-of-sight link stability and maximized coverage.

[0027] 2. Precise Spectrum Resource Utilization: Dynamic spectrum aggregation technology, based on sensing precision μ, enables secondary users to accurately identify and integrate idle channels to meet bandwidth requirements. A continuous spectrum hole aggregation strategy, combined with a conflict detection mechanism, significantly reduces the probability of channel collisions between users and improves spectrum utilization.

[0028] 3. Differentiated Service Quality Assurance: A priority mechanism is used to reserve high-quality channels for disaster relief personnel, ensuring critical voice and video transmission needs. Secondary users utilize a competitive reinforcement learning access strategy to ensure basic text and voice emergency communications while avoiding interference from primary users.

[0029] 4. Intelligent Dynamic Spectrum Management: A multi-agent reinforcement learning framework enables each user to autonomously perceive the full spectrum status and make real-time decisions, adapting to dynamic changes caused by random access by primary users. A dual-dimensional quality assessment mechanism based on signal-to-noise ratio and transmission rate guides the agents to optimize high-gain channel combinations.

[0030] 5. Optimizing anti-interference communication performance: The Rice channel model, combined with path loss calculation, enhances the robustness of channel estimation in complex environments. A dynamic correction formula for the signal-to-noise ratio in conflict scenarios accurately reflects the impact of multi-user interference on transmission quality, providing reliable data support for strategy optimization.

[0031] 6. Efficient Policy Learning Mechanism: The maximum entropy method adds an entropy regularization term to the policy gradient, effectively preventing agents from prematurely converging to suboptimal policies. The combination of soft updates of the target network and experience replay technology significantly improves the stability and convergence speed of the multi-agent collaborative learning process.

[0032] 7. Adaptability to Multiple Service Needs: A hierarchical bandwidth allocation mechanism meets the differentiated needs of primary users (high-bandwidth video transmission) and secondary users (basic text communication). A flexible design with adjustable aggregation length (L) supports on-demand spectrum access for devices with varying hardware capabilities.

[0033] The combined effect of these excellent characteristics enables the UAV-RCNs system provided by the present invention to demonstrate significant technical advantages and application value in emergency communication scenarios such as disaster relief, providing reliable technical support for emergency rescue communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a schematic diagram of the UAV emergency communication system architecture of the present invention; Figure 2 Schematic diagram of the UAV-RCNs system model of the present invention; Figure 3 Schematic diagram of the spectrum access timing model of the present invention; Figure 4 Schematic diagram of the ME-MAAC algorithm framework of the present invention; Figure 5 Schematic diagram of key indicators comparing the algorithm proposed in this invention with the DQN algorithm; Figure 5 (a) is a schematic diagram showing the changing trend of the long-term cumulative discounted rewards of different algorithms during the training process; Figure 5 (b) is a schematic diagram of the change in collision probability between secondary users due to spectrum conflicts; Figure 5 (c) is a schematic diagram showing the probability of spectrum conflict between secondary users and primary users; Figure 5 (d) is a schematic diagram showing the probability of users giving up access due to unmet bandwidth requirements; Figure 6Schematic diagram comparing the multi-SU load performance of the algorithm proposed in this invention and the DQN algorithm; Figure 6 (a) is a diagram showing the long-term reward changes of the algorithm when the number of users increases; Figure 6 (b) is a performance diagram of the user's actual data transmission rate; Figure 6 (c) is a schematic diagram showing the impact of the number of users on the probability of conflict; Figure 6 (d) is a schematic diagram of energy consumption of different algorithms; Figure 7 Schematic diagram of the impact of the algorithm proposed in this invention and the DQN algorithm on perception accuracy; Figure 7 (a) is a schematic diagram of the cumulative discounted rewards for each secondary user under different algorithms; Figure 7 (b) is a schematic diagram of the average collision rate between secondary users and primary users for different algorithms; Figure 7 (c) is a schematic diagram of the interruption probability of different algorithms; Figure 8 This is a multi-PU load performance curve diagram of the algorithm proposed in this invention and the DQN algorithm; Figure 8 (a) is a diagram showing the average successful access rate of users under different aggregation lengths and bandwidth requirements; Figure 8 (b) is a schematic diagram showing the impact of user aggregation capability and bandwidth requirements on the interruption probability; Figure 8 (c) is a schematic diagram of the average collision rate of major users; Figure 8 (d) is a schematic diagram of the average collision rate of secondary users. DETAILED DESCRIPTION

[0035] The present invention will be further described below with reference to embodiments and accompanying drawings, but is not limited thereto.

[0036] Example 1 A spectrum scheduling method for UAV emergency communication based on dynamic entropy optimization, such as Figure 1 Shown, including: Step 1: Build an emergency communication system to communicate between users and mobile base stations and calculate the received signal gain; Step 2: Build a spectrum sensing and user access model to identify available channels and formulate user access strategies. Step 3: Evaluate the channel quality in the emergency communication system; Step 4: Establish the state, action, and reward of deep reinforcement learning based on spectrum sensing and user access models and channel quality assessment; Step 5: Build a deep reinforcement learning model and train it to obtain a trained deep reinforcement learning model; Step 6: Use the trained deep reinforcement learning model to obtain the optimal strategy for spectrum scheduling of drone emergency communications in the spectrum environment and achieve aggregated spectrum access for multiple users.

[0037] Example 2 The difference between the method for scheduling spectrum for UAV emergency communication based on dynamic entropy optimization described in Example 1 is that: Construct an emergency communication system to communicate between users and mobile base stations and calculate the received signal gain; Figure 2 Shown, including: The emergency communication system includes drones and mobile base stations carried by drones. The drones communicate with users via air-to-ground channels. Users include primary users (rescue personnel) and secondary users (disaster victims). Air-to-ground channel communications consider both line-of-sight (LoS) and non-line-of-sight (NLoS) path losses. The total path loss is calculated using a probabilistic model, and the Rice channel model is introduced to describe small-scale fading. The distance between the drone and the user, the carrier frequency, and environmental parameters all influence the channel quality, laying the foundation for subsequent spectrum allocation. The UAV in the emergency communication system hovers at a fixed height h and flies along a fixed circular trajectory at a constant speed v; the three-dimensional coordinates of the UAV, primary user and secondary user are set as (x, y, h), ( , ,0) and (x n , y n ,0); then the distances between the UAV and the mth primary user and the nth secondary user are as follows: (1); (2); Among them, d m represents the distance between the UAV and the mth primary user, d n represents the distance between the drone and the nth user, x represents the horizontal coordinate of the drone, y represents the vertical coordinate of the drone, and x m Indicates the horizontal coordinate of the primary user, y m Indicates the vertical coordinate of the primary user, x n Indicates the horizontal coordinate of the secondary user, y nrepresents the ordinate of the secondary user; compared with the ground base station, the unmanned aerial vehicle has a higher line of sight (LoS) link; the application calculates the path loss by adopting an air to ground (A2G) channel model and ignoring the influence of small-scale fading; it is assumed that there is a certain probability of LoS link and NLoS link (Non Line of Sight, NLoS) between the unmanned aerial vehicle and all users; the path losses of the LoS link and the NLoS link are defined as shown in formula (3) and (4) respectively: (3); (4); wherein, represents the path loss of the LoS link, represents the path loss of the NLoS link, represents a carrier frequency unique to each channel, i.e. air to ground channel, and d is the distance between the unmanned aerial vehicle and the user; for the purpose of simplifying calculation, it is assumed that the distance d is fixed when the user accesses the spectrum (the radio frequency resource provided by the mobile base station carried by the unmanned aerial vehicle); represents the speed of light, and are the additional losses of the LoS link and the NLoS link respectively, and the sizes thereof depend on the environmental parameters of the urban environment; the spectrum configured by the unmanned aerial vehicle includes multiple independent channels, and each channel has a unique carrier frequency ; It is assumed that the probabilities of the LoS link and the NLoS link between the users are represented as: (5); (6); wherein, represents the probability of the LoS link, represents the probability of the NLoS link; a and b are parameters determined by the urban environment; the total path loss formula between the unmanned aerial vehicle and each user is obtained from the above formula (5) and (6), and the value thereof depends on the distance d and the carrier frequency , as shown in formula (7): (7); wherein, represents the total path loss between the unmanned aerial vehicle and each user; In order to enhance the robustness of channel estimation, a Rician channel model is introduced to describe small-scale fading, and a Rician channel model is adopted to calculate the received signal gain, as shown below: (8); where g denotes the received signal gain, represents the energy level of the signal after passing through the path loss, , i.e. is determined by the path loss; is the Rician factor, which represents the energy ratio between the LoS link component and the scattering component (all signal components caused by reflection, diffraction, etc. in addition to the direct path); the phase corresponding to the LoS link component follows a uniform distribution, and represents that the NLoS link component corresponds to a complex Gaussian random variable, represents a complex normal distribution (Complex Normal Distribution) for describing a random variable with real and imaginary parts, represents a complex exponential function, where e is the base of natural logarithm (approximately equal to 2.71828), represents the imaginary unit, represents the phase angle.

[0038] A spectrum sensing and user access model is constructed to identify available channels and formulate user access strategies; as shown in the following table, the model includes: Figure 3 In the emergency communication scenario, the spectrum resource is often extremely scarce. In order to maximize the spectrum utilization efficiency, a dynamic spectrum aggregation and access control system is designed. The system adopts a time slot frame structure design, which divides each working period into four stages: first, a random access stage dominated by primary users (PUs), which ensures the absolute priority of communication for key users such as disaster relief commanders and emergency doctors; then, a spectrum sensing stage for secondary users (SUs), in which ordinary disaster victims can scan available channels through efficient energy detection technology; then, a critical spectrum aggregation stage, in which the system automatically combines discrete idle sub-channels into continuous available frequency bands; finally, an intelligent network updating stage, in which the system automatically adjusts parameters according to the current running state to optimize the resource allocation strategy for the next working period. The spectrum sensing and user access model includes a random access stage, a spectrum sensing stage, a spectrum aggregation stage, and a network updating stage. In the random access stage, M primary users randomly access multiple channels according to their bandwidth requirements at the beginning of each time slot, the spectrum is not completely idle due to the occupation of primary users; In the spectrum sensing stage, N secondary users observe the idle channels through their spectrum sensing capabilities. The accuracy of spectrum sensing is the basis for reliable operation of the system, and the spectrum sensing capability includes detection probability and false alarm probability; the detection probability ​Indicates the probability of correctly identifying an occupied channel and the false alarm probability The probability of misjudging an idle channel as an occupied channel; the spectrum sensing capability of the secondary user is expressed in spectrum sensing accuracy. express, The higher it is, the more accurate the secondary user's judgment of the channel status is. The formula is as follows: (9); in, represents the spectrum sensing accuracy of the nth secondary user; In the spectrum aggregation stage, secondary users The observation state, or the estimated state of each channel occupancy obtained through spectrum sensing, is obtained. Based on the spectrum aggregation capability, the limited frequency bands are aggregated into different aggregated bands. However, due to hardware limitations such as transceivers, spectrum aggregation capabilities dictate that secondary users can only aggregate continuous spectrum holes, or spectrum resources not currently occupied by primary users and available to secondary users, with a fixed aggregation length, L. Secondary users select the appropriate aggregated band for access, and access success is determined by base station interaction information. This includes two scenarios: successful access and failed access. Successful access by a secondary user: Based on the secondary user's correct observation status, when the number of channels in the aggregated frequency band selected by the secondary user is greater than or equal to the secondary user's bandwidth requirement B, and the selected channels do not conflict with any other users, the secondary user will select multiple high-quality channels in the aggregated segment (signal-to-noise ratio greater than 20dB, path loss less than 60dB, gain greater than 10dBi, unoccupied channels, and channels with good historical performance) for access. The mobile base station carried by the drone will then provide feedback to the secondary user, including rewards such as transmission rate, data packet loss rate, signal-to-noise ratio, and current channel load. The secondary user will then send multimedia data (video images, location information of disaster victims, etc.) to the mobile base station carried by the drone. If there are still idle channels in the aggregated frequency band, other users can occupy them for data transmission. Failed access by a secondary user: If the secondary user experiences erroneous perception, the aggregated band does not meet its needs, or the selected channel conflicts, the secondary user will abandon the aggregated band and reselect the spectrum in the next time slot based on the previous experience and update the strategy to retransmit the unsent data. (After a failed access attempt, the secondary user uses a reinforcement learning algorithm to update its behavior strategy based on the experience of the previous time slot, such as perception errors, improper aggregation, and conflicts, to make a better access decision in the next time slot.) In the final stage of the time slot, the access strategy update stage, the user agent (an independent intelligent decision-making agent corresponding to each secondary user) will use environmental information and feedback information from the base station (such as the current channel status and spectrum sensing results) to update the access strategy, continuously optimizing the action strategy so that the strategy is gradually transformed into the optimal strategy.

[0039] Evaluate the channel quality in emergency communication systems; including: The spectrum of the mobile base station carried by the drone is limited. Assuming that the bandwidth of each channel is are the same, the entire spectrum is evenly divided into K channels; at the same time, the transmission power P of each channel is also the same, but the carrier frequency of each channel is unique; therefore, the channel gain of the nth secondary user on the K channels is defined as: (10); in, represents the channel gain of the nth user, represents the channel gain of the nth secondary user on the kth channel; to evaluate the quality performance of the channel, the signal-to-noise ratio and transmission rate are set as the channel quality standards; to make full use of the idle channels in the spectrum, the aggregated frequency band selected by a secondary user contains not only the channels occupied by the primary user, but also the channels accessed by other users; therefore, when there is no conflict between secondary users, the signal-to-noise ratio obtained by the nth secondary user is is defined as: (11); in, is the noise spectral density, is the gain of the ith channel in the segment selected by the nth secondary user, is the gain of the j-th channel of the m-th primary user, Indicates the nth user selection channels, Indicates the mth primary user's selection channels, Indicates the bandwidth requirement of the nth secondary user, N is the number of secondary users, and M is the number of primary users; If there are K channels in the aggregated frequency band selected by the nth secondary user that conflict with other secondary users, the signal-to-noise ratio is represented as: (12); in, represents the gain of the vth conflicting channel between the nth secondary user and the remaining u secondary users, represents the transmission power of the kth user, represents the number of channels in the aggregated frequency band selected by the u-th secondary user that have access conflicts with other secondary users; According to the above SNR formula, the transmission rate of the nth secondary user is Expressed as: (13); Or expressed as: (14); in, It is the loss coefficient of the signal-to-noise ratio, which indicates the loss between the theoretical channel capacity and the specific modulation and coding scheme. If the signal-to-noise ratio is greater than the set threshold (20dB) and the transmission rate is greater than the set threshold (1M), the channel quality is judged to be high.

[0040] Based on spectrum sensing and user access models and channel quality assessment, establish the states, actions, and rewards of deep reinforcement learning; including: Secondary users act as independent intelligent agents. Their state reflects spectrum channel occupancy, including spectrum status and user observation status. Their actions involve selecting a specific aggregated frequency band for access. Rewards are a combination of transmission rate and conflict penalties to encourage users to optimize channel selection. The goal is to maximize long-term discounted rewards. Through policy iteration, secondary users are able to adaptively learn the optimal access strategy in complex spectrum environments. In a complex spectrum environment, each user, acting as an independent intelligent agent, comprehensively considers its own capabilities and bandwidth requirements, monitors the entire spectrum in real time, and selects channels that meet its needs and provide good quality to ensure data transmission for different users. Based on the principles of reinforcement learning algorithms, this paper describes in detail the state, action, and reward-related elements to prepare for the detailed explanation of the algorithm. (1) Spectrum state and user observation state: The spectrum is evenly divided into K independent channels, and the state of the spectrum at time slot t is expressed as ,in Indicates the state of the kth channel in time slot t. Each channel has two states: occupied state (S=1) or idle state ( =0); however, at the initial moment, the spectrum is actually completely idle, and changes in the channel state are caused by the access of the primary user. The secondary user has no idea about the primary user's spectrum usage pattern, so the secondary user needs to learn the primary user's occupancy pattern through spectrum sensing. The observed states of N secondary users are defined as Since the spectrum carried by the drone is very limited, the secondary user no longer performs partial spectrum perception, but directly perceives the entire spectrum; therefore, the observation state of the nth secondary user for the entire spectrum at time slot t is expressed as ,in, represents the state of the kth channel perceived by the nth secondary user at time slot t; (2) Action: The action consists of the set of actions of all secondary users, i.e. , Represents the action of the nth sub-user in time slot t. Each sub-user is regarded as an intelligent agent, according to its own observation state Select actions independently ; There are two situations for the selected action: when the selected frequency band can meet the bandwidth requirement B, and the selected channel does not conflict with the channel occupied by the primary user, the secondary user will select the channel with higher channel gain in the aggregated frequency band to meet its own bandwidth requirement B; otherwise, when the aggregated frequency band cannot meet the requirement, the secondary user will give up accessing the channel; according to the aggregation length of the secondary user ( ), the entire spectrum is divided into K- +1 aggregated frequency band; the secondary user will select the most appropriate aggregated frequency band for access based on the action strategy; (3) Reward: The reward set of N users is expressed as ,in represents the reward of the nth secondary user in time slot t; to encourage the secondary user to choose an appropriate action, a penalty term f (f = -10) is set. When the secondary user chooses an action that does not meet expectations, the mobile base station will feedback the penalty term to remind the secondary user to update the action strategy; the reward function combines the secondary user's transmission rate and the penalty term f, and is divided into four cases. The reward of the nth secondary user in time slot t is expressed as: (15); Since the access scheme cancels the priority setting, the relationship between secondary users is no longer cooperative, but they compete to select high-quality channels; The goal of each sub-user is to find an optimal action strategy that maximizes the long-term cumulative discounted benefit, as shown below: (16); in, ∈[0,1] is the discount rate that measures the impact of future rewards on cumulative rewards, and E[] represents the discount rate of the strategy The expected value of the weighted sum of all future rewards under Indicates that in the strategy The long-term cumulative discount benefit under the The total revenue that can be obtained.

[0041] Construct a deep reinforcement learning model and train it to obtain a trained deep reinforcement learning model; Figure 4 Shown, including: Build a deep reinforcement learning model, namely the ME-MAAC algorithm reinforcement learning model, which includes the Actor network (generating action strategies), the Critic network (evaluating Q values), the target network, and the experience replay area; The Actor network outputs discrete actions (Softmax scores) based on the observed state, while the Critic network calculates the action value and introduces an entropy term to enhance exploration. Policy gradient updates combine Q-values ​​and entropy regularization to balance exploration and exploitation. The target network improves stability through soft updates, and experience replay reduces data correlation to accelerate convergence. The network parameters are The Actor network uses the exploration strategy To choose actions, strategies To observe the status For input value, take action is the output value; the action selection function is described as ; Since there is no correlation between channels, the action space is discrete, and the score of each aggregated frequency band is evaluated by using a softmax function in the network; actions with higher scores are more likely to be selected than actions with lower scores; The parameters of the critic network are , the user agent obtains the action and observation status of the current time slot from the Actor network , and input to the Critic network, then the user agent will get the output Q value , to evaluate the action of the current time slot; both the Actor network and the Critic network use the Adam optimizer to optimize the parameters; The ME-MAAC algorithm reinforcement learning model uses a policy gradient method to update the actor network and a loss function to update the critic network. Because all secondary users need to perceive the entire spectrum to obtain the observed states of all channels and select an appropriate spectrum segment, their state and action spaces have the same size, so all user agents have the same network structure. The policy gradient exists in the Actor network and is used to update the Actor network. When training the network, each secondary user, as an intelligent agent, randomly samples a small batch of observation state and action data from the experience replay area, and then transmits this data to the Actor network and the Critic network to calculate the Q value respectively. and action-value function ; In order to add some randomness when exploring actions and avoid converging to non-optimal deterministic policies, a maximum entropy method is introduced in the ME-MAAC algorithm, that is, an entropy term that supports random policies is added to the policy gradient and loss function. , then the policy gradient of the nth user is expressed as: (17); in, represents the policy gradient of the nth user, Indicates network parameters The derivative operation of is the control factor that controls the randomness of the optimal strategy, D represents the experience replay area, Indicates the state sampled from the experience playback area D and actions The expected value of Indicates that in a given state Take action The Q value is estimated by the Critic network with parameter θ; The nth user samples a certain number of experience units , respectively, through the Critic network and its target network to obtain the Q value and target Q value ,in Represents the next observed state, and then adds the target Q value and the entropy term obtained from the target Actor network Combined to calculate the output value of the target network , the secondary user updates the Critic network by minimizing the regression loss function, which is defined as: (18); (19); in, represents the loss function of the Critic network of the nth user, Represents the experience unit sampled from the experience playback area D The expected value of Indicates that in a given next observation state Next, the action distribution generated by the target Actor network Action expected value; The network structure of the target network is the same as that of the source network, namely the Actor network and the Critic network, and the parameters of the target Actor network and the target Critic network in the target network are and Frequently updating the Q value may lead to unstable training. Therefore, the target network soft-updates its own parameters by copying a small part of the source network parameters to enhance the update stability of the network. The update formula of the target network parameters is expressed as: (20); (twenty one); Where τ∈[0,1] is the soft update coefficient; In each time slot, the user integrates all the data obtained from interacting with the environment, namely the observed state, action and reward, into a set of experience units. , and stored in the experience replay area D; during each network training process, the secondary user randomly extracts experience units to update the network parameters; this is because the continuous data obtained by the user's interaction with the environment has a certain correlation, but the data used by the deep neural network must be independent and identically distributed. Experience replay helps to reduce the correlation of data samples and improve learning efficiency and convergence speed; through continuous updating of training by the target network and the source network, a trained deep reinforcement learning model is obtained.

[0042] exist Figure 1 In the perception layer, the system first performs multi-spectral detection to collect spectrum information in the environment. This perception data is passed to the channel modeling module, where LoS / NLoS identification is performed to assess the signal quality of different paths. Simultaneously, this perception data is sent to the decision layer. The decision layer is the core of the system and consists of two submodules: the ME-MAAC algorithm and dynamic spectrum aggregation. The ME-MAAC algorithm uses maximum entropy optimization and multi-agent collaboration to achieve efficient spectrum resource allocation and management. Dynamic spectrum aggregation flexibly aggregates available spectrum resources based on current network status and needs. Based on these analysis results, the decision layer generates control instructions and passes them to the execution layer. The execution layer is responsible for specific implementation operations, including power control and priority access. Power control ensures stable and efficient signal transmission, while the priority access mechanism ensures that primary users receive priority access to high-quality channels to meet their high bandwidth requirements. Furthermore, status feedback from the execution layer is fed back to the decision layer, forming a closed-loop control loop to continuously optimize system performance.

[0043] Figure 5 Two simple scenarios in UAV-RCNs are considered, namely single-user N=1 scenario and multi-user N=2 scenario, where there are M=3 bandwidth requirements in the network scenario. =[5,4,3], and there are N primary users with the same perception accuracy. =0.99, bandwidth requirements are =2 and the aggregation length is =4 secondary users, that is, the secondary users have the same spectrum sensing and aggregation capabilities.

[0044] from Figure 5As can be seen in (a), the cumulative discounted rewards of the proposed algorithm and the DQN algorithm increase with the increase in the number of training times, and the reward of the proposed algorithm is significantly higher than that of the DQN algorithm. However, in the single-user scenario, the reward curve of the traditional Actor-Critic algorithm rises very slowly. In the multi-user scenario, the reward curve is flat, and even lower than the reward value of the DQN algorithm in the single-user scenario in the later stage of training. The reason for the low reward value may be too many collisions between users or the failure of users to find a suitable frequency band. The average collision rate is respectively given by Figure 5 (b) and Figure 5 (c) show; Figure 5 (b) shows how the average collision rate between users changes with the number of training times. Both the proposed algorithm and the DQN algorithm can reduce the average collision rate. The proposed algorithm reduces the average collision rate from 3.09% to 0.13% after 200,000 training times, while the DQN algorithm has a more obvious downward trend in the collision rate in the first 50,000 training times, but the average collision rate shows an upward trend in the subsequent longer training time. Figure 5 (c) shows the variation of the average collision rate between the secondary user and the primary user over time. Figure 5 Panel (d) shows the outage probability of the three algorithms in single-user and multi-user scenarios, that is, the probability that a secondary user will be unable to access the channel. In both single-user and multi-user scenarios, the Actor-Critic algorithm fails to reduce the outage probability. This is because the secondary user's action strategy is unable to find an aggregated frequency band with multiple idle channels, and the strategy has not been trained by the algorithm, resulting in a flat out outage probability curve. The DQN algorithm, however, can reduce the outage probability to a certain extent through training and learning. Compared with the aforementioned two algorithms, the proposed algorithm significantly reduces the outage probability, even dropping it from 3.97% to 0.83% in the multi-user scenario.

[0045] In summary, the multi-agent algorithm using a stochastic strategy offers significant advantages over the deterministic DQN and Actor-Critic algorithms. The proposed algorithm explores a wider range of dynamically changing environments, rapidly and significantly reduces collision and outage probabilities, and encourages secondary users to access higher-quality channels whenever possible, thereby increasing transmission rates. However, in single-user scenarios, the Actor-Critic algorithm can only reach a local optimum after a short training period. Consequently, the strategy is unable to explore alternative actions, resulting in a stagnation in the cumulative discounted reward. In multi-user scenarios, the Actor-Critic algorithm is unable to handle multi-user conflicts and even performs worse than the DQN algorithm.

[0046] Figure 6 The impact of different numbers of secondary users on the performance of the emergency communication network is shown, where the bandwidth requirements of M=3 primary users in the network are =[5,4,3], these primary users randomly access the channel at the beginning of each time slot, there are also N e {2,3,4,5} secondary users with the same =0.99, =2 and =4 secondary users share the spectrum with the primary users.

[0047] Figure 6 In Fig. 4(a), the cumulative discounted rewards obtained by the two algorithms are compared, which is the sum of the long-term cumulative rewards that all secondary users can obtain. The cumulative discounted rewards of the proposed algorithm and the DQN algorithm both increase with the increase of the number of secondary users. The reward of the proposed algorithm increases more significantly and stably. However, compared with the proposed algorithm, the reward of the DQN algorithm performs poorly, that is, when the number of users increases too much, the reward will decrease, and when N = 5, the total reward of the secondary users is significantly lower than that when N = 4, which shows that the DQN algorithm can only handle a certain number of secondary users at the same time, and even a slight increase in the number may lead to a decline in network performance, which cannot meet the requirements of emergency communication networks to handle large-scale traffic peaks.

[0048] Figure 6 In Fig. 4(b), the average achievable transmission rate of the secondary users is shown. With the increase of the number of secondary users, the average achievable rate of the two algorithms shows a downward trend. It can be further observed that the proposed algorithm can achieve a higher average achievable rate than the DQN algorithm, even when N = 5, the achievable rate of the proposed algorithm is higher than that of the DQN algorithm when N = 2, which shows that the UAV-RCNs using the proposed algorithm can serve more secondary users than the DQN algorithm, and try to provide each user with a good quality channel as much as possible.

[0049] Figure 6 In Fig. 4(c), the average collision rate between secondary users under different algorithms is mainly illustrated. The average collision rate of the proposed algorithm and the DQN algorithm both increases with the increase of the number of secondary users, which shows that the more users, the more collisions may occur. The collision rate of the DQN algorithm shows an upward trend in the later training period under any given number of secondary users, and if the training time is too long, the average collision rate may actually increase, which is undesirable. Therefore, in dealing with the conflict problem between secondary users, the proposed algorithm can more effectively reduce the average collision rate than the DQN algorithm, so that users can find an idle channel to communicate at each access.

[0050] Figure 6Panel (d) compares the total power consumption of all secondary users in the two algorithms under different load conditions. As the number of users increases, the power consumption of both algorithms increases, as multiple users inevitably lead to increased power consumption. However, when the number of users remains the same, the power consumption of the proposed algorithm is significantly lower than that of the DQN algorithm. In fact, the power consumption of N secondary users in the proposed algorithm is approximately equal to that of N-1 users in the DQN algorithm.

[0051] In summary, the proposed algorithm can handle multi-user networks more effectively than the DQN algorithm, and reduces channel conflicts and power consumption between secondary users through action strategies, thereby increasing the average achievable rate of each user and thus improving the efficiency of the entire network.

[0052] Figure 7 Shows the performance of secondary users with different spectrum sensing accuracy in the same network, where each user has the same bandwidth requirement =2 and aggregation length =4, and there are M=3 bandwidth requirements in the network. =[5,4,3] primary users. Assume that each secondary user in the same network has independent spectrum sensing accuracy , by the detection probability ∈[0.85,0.99] and false alarm probability ∈[0.01,0.25], that is, the accuracy of all secondary users is μ=[0.6375,0.7425,0.8415,0.9801], among which secondary user 1 has the lowest perception accuracy =0.6375, user 4 has the highest accuracy =0.9801.

[0053] Figure 7 (a) compares the cumulative discounted reward for each sub-user under the two algorithms. The cumulative discounted reward is a key indicator for evaluating the optimal benefits of spectrum aggregation and access schemes. Although the reward values ​​of both algorithms increase with the increase of perception accuracy, the proposed algorithm can obtain higher and more stable rewards than the DQN algorithm. When there are a lot of perception errors in the observation state, it is more difficult for the sub-user to find a frequency band that meets the bandwidth requirements, resulting in a decrease in reward. Sub-user 1 has the worst perception accuracy. =0.6375, both the proposed algorithm and the DQN algorithm can only obtain negative rewards, and even the DQN algorithm reward curve is declining. At this time, the DQN strategy has completely failed. The proposed algorithm evaluates all frequency bands through a strategy to enable users to select the frequency band in the correct state to increase the reward. Therefore, as the perception accuracy increases, the proposed algorithm can obtain a positive reward value. However, the rewards for sub-users 2 and 3 using the DQN algorithm are still negative. In fact, only sub-user 4 with the highest accuracy obtains a positive reward. This result shows that compared with the proposed algorithm, the DQN algorithm's strategy is less efficient in finding channels that meet user needs. This requires that sub-users using the DQN algorithm must perceive the spectrum very accurately to obtain better benefits. Combined with Figure 6 (b) and Figure 7 Furthermore, when there are users with different perception accuracy in the same network, users with poor perception accuracy will affect the performance of other users with higher perception accuracy, which leads to a significant decrease in the overall network performance.

[0054] because As the channel power decreases, the probability of users misperceiving the channel increases, thereby increasing the probability of collision between secondary users and primary users, which causes negative interference to the network. Figure 7 (b) shows the average collision rate between secondary users and primary users for different μ,. The results show that secondary users with higher accuracy have lower and more stable average collision rates compared to secondary users with lower accuracy. When M=0.6375, the rising probability curve of the proposed algorithm indicates that the secondary user has limited ability to handle conflicts, resulting in negative rewards for the user. Figure 7 In (b), we can further observe that the average collision rate of the DQN algorithm is lower than that of the proposed algorithm, but only by Figure 7 (b) doesn't indicate that the proposed algorithm performs worse than the DQN algorithm in handling collisions. This requires analysis in conjunction with the performance of the outage probability. This phenomenon occurs because the DQN algorithm significantly reduces the total number of successful accesses, resulting in fewer collisions. In reality, collisions between secondary and primary users are inevitable due to perception errors. This requires the algorithm to minimize the number of failed accesses and find suitable frequency bands to maximize the cumulative reward.

[0055] Figure 7The impact of spectrum sensing accuracy on the outage probability of secondary users is investigated in (c). The DQN algorithm has a higher outage probability than the proposed algorithm, i.e., secondary users using the DQN algorithm have more failed access attempts. When the spectrum sensing accuracy is high, the DQN algorithm can reduce the outage probability, but when the accuracy is the worst, the DQN algorithm cannot find a suitable spectrum, resulting in a high outage probability. Notably, when secondary users use the proposed algorithm, the outage probability decreases rapidly with the number of training iterations, even for the secondary user 1 with the lowest sensing accuracy, the outage probability is significantly reduced. Even in general, the outage probability of the secondary user 4 with a sensing accuracy of = 0.9801 using the DQN algorithm is significantly higher than that of the secondary user 1 with the lowest accuracy M = 0.6375 using the proposed algorithm, which indicates that the proposed algorithm has stronger robustness in dealing with spectrum sensing errors than the DQN algorithm, which is beneficial for users to adapt to extremely harsh communication environments stably.

[0056] Figure 8 The impact of different user aggregation lengths ∈ {4, 5, 6, 7, 8} and bandwidth requirements ∈ {2, 3, 4} on the performance of emergency communication networks is investigated. The primary users in the network are the same as described above, but to reduce the interference of collision and sensing errors between secondary users, it is assumed that there are only N = 2 = 0.99 secondary users in the UAV-RCNs.

[0057] Figure 8 The average successful access rate of users under different aggregation lengths and bandwidth requirements is compared in (a). It can be seen from the figure that the proposed algorithm is superior to the DQN algorithm under all conditions. When the aggregation length is fixed, the successful access rate decreases as the bandwidth requirement increases. This is because, under the condition of limited aggregation capacity, the increase in user demand means that the agent needs to find a frequency band with more free channels in the given spectrum for data transmission, but such frequency bands are very rare, and there is also competition between users, resulting in a decrease in the probability of successful access. When users use the proposed algorithm, for high aggregation length conditions, the access performance difference between different user bandwidth requirements becomes less significant, for example, for = 8, the successful access probability of users with a demand of = 2 is only 2.45% higher than that of = 3, = 3 is only 2.15% higher than that of = 4. It is further observed that when the bandwidth requirement is fixed, the successful access probability increases with the aggregation length, because the longer the selected frequency band, the greater the likelihood that it contains a large number of free channels available for secondary users to access.

[0058] Figure 8 (b) shows the impact of user aggregation capability and bandwidth demand on the interruption probability. When the user spectrum aggregation capability is equal to the bandwidth demand, that is, = =4, the outage probability of the DQN algorithm is 68.20%, and the outage probability of the proposed algorithm is 50.5%. Although the proposed algorithm outperforms the DQN algorithm in this case, the successful access rate of both algorithms is very low because there are very few completely idle aggregated frequency bands in the spectrum, which is not conducive to the access of secondary users. As the aggregation length increases and the bandwidth demand decreases, the outage probability of both algorithms decreases sharply, and the achievable outage probability of the proposed algorithm is significantly lower than that of the DQN algorithm. For example, when L=4 and =3, the interruption probability of the DQN algorithm is 25.37% higher than that of the proposed algorithm. = 2, the aggregated frequency bands selected by the user are easy to meet the bandwidth requirements, so the outage probability of both algorithms is small and stable, especially when the proposed algorithm is used.

[0059] Figure 8 Section (c) studies the average collision rate between secondary users and primary users. Although secondary users all have high spectrum sensing accuracy, the possibility of collisions still exists. Both algorithms have a certain average collision rate, but compared to the interruption probability, the impact of this collision rate is small and can be ignored. However, when the aggregation capacity is equal to the bandwidth demand, the average collision rate is actually lower. This is not because the algorithm avoids conflicts with the primary user, but because the algorithm first determines whether the frequency band meets the requirements when evaluating it. If the frequency band does not meet the requirements, the user abandons access, which means an interruption occurs. Therefore, when the interruption probability is too high, the collision probability is reduced. This is also reflected in the collision situation between secondary users.

[0060] Figure 8 (d) shows the average collision rate between users. = = 4 when the collision rate is lower than > =4 average collision rate, this is because in this case, a large number of aggregated frequency bands cannot be accessed. = = 4, only a small number of aggregated frequency bands can be selected, which also reduces the number of frequency bands that may cause conflicts. In other cases, the collision rate curve of the DQN algorithm shows an upward trend. When the user has more frequency bands to choose from, the poor strategy of DQN makes the conflict more intense, while the collision rate of the proposed algorithm shows a downward trend, especially =4, so this shows that the proposed algorithm can handle channel conflicts between secondary users more effectively than the DQN algorithm.

[0061] In summary, when secondary users have different spectrum aggregation capabilities and bandwidth requirements, the proposed algorithm significantly reduces the outage probability and average collision rate, achieving a higher successful access rate for each user. The algorithm's flexible adaptability facilitates smooth data transmission for users with diverse service requirements in emergency communication networks, especially those with high-quality communications requirements.

[0062] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for scheduling spectrum of unmanned aerial vehicle emergency communications based on dynamic entropy optimization described in Example 1 or 2.

[0063] Example 4 A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for scheduling spectrum for unmanned aerial vehicle emergency communications based on dynamic entropy optimization described in Example 1 or 2 are implemented.

[0064] Example 5 A spectrum scheduling system for UAV emergency communication based on dynamic entropy optimization, including: The emergency communication system building module is configured to: build an emergency communication system, perform communication between users and mobile base stations, and calculate received signal gain; The spectrum sensing and user access model building module is configured to: build spectrum sensing and user access models, identify available channels, and formulate user access policies; The channel quality assessment module is configured to: assess the channel quality in the emergency communication system; A reinforcement learning parameter establishment module is configured to establish states, actions, and rewards for deep reinforcement learning based on spectrum sensing and user access models and channel quality assessments; The reinforcement learning model building module is configured to: build a deep reinforcement learning model and perform training to obtain a trained deep reinforcement learning model; The spectrum aggregation module is configured to use the trained deep reinforcement learning model to obtain the optimal strategy for spectrum scheduling of drone emergency communications in the spectrum environment and achieve aggregated spectrum access for multiple users.

Claims

1. A spectrum scheduling method for UAV emergency communication based on dynamic entropy optimization, characterized in that: include: Step 1: Build an emergency communication system to communicate between users and mobile base stations and calculate the received signal gain; Step 2: Build a spectrum sensing and user access model to identify available channels and formulate user access strategies. Step 3: Evaluate the channel quality in the emergency communication system; Step 4: Establish the state, action, and reward of deep reinforcement learning based on spectrum sensing and user access models and channel quality assessment; Step 5: Build a deep reinforcement learning model and train it to obtain a trained deep reinforcement learning model; Step 6: Use the trained deep reinforcement learning model to obtain the optimal strategy for spectrum scheduling of drone emergency communications in the spectrum environment and achieve aggregated spectrum access for multiple users.

2. The spectrum scheduling method for UAV emergency communication based on dynamic entropy optimization according to claim 1 is characterized in that: Build an emergency communication system to communicate between users and mobile base stations and calculate the received signal gain; including: The emergency communication system includes a drone and a mobile base station carried by the drone. The drone communicates with users through an air-to-ground channel; users include primary users and secondary users. The UAV in the emergency communication system hovers at a fixed height h and flies along a fixed circular trajectory at a constant speed v; the three-dimensional coordinates of the UAV, primary user and secondary user are set as (x, y, h), ( , ,0) and (x n , y n ,0); then the distances between the UAV and the mth primary user and the nth secondary user are as follows: (1); (2); Among them, d m represents the distance between the UAV and the mth primary user, d n represents the distance between the drone and the nth user, x represents the horizontal coordinate of the drone, y represents the vertical coordinate of the drone, and x m Indicates the horizontal coordinate of the primary user, y m Indicates the vertical coordinate of the primary user, x n Indicates the horizontal coordinate of the secondary user, y n represents the ordinate of the secondary user; it is assumed that there are LoS links (i.e., line-of-sight links) and NLoS links (i.e., non-line-of-sight links) with a certain probability between the UAV and all users; the path losses of the LoS link and the NLoS link are defined as shown in equations (3) and (4), respectively: (3); (4); in, represents the path loss of the line-of-sight link, represents the path loss of the non-line-of-sight link, represents the unique carrier frequency of each channel, i.e., the air-to-ground channel, d is the distance between the drone and the user, represents the speed of light, and The additional losses of LoS and NLoS links respectively, and their magnitude depends on the environmental parameters of the urban environment; the spectrum configured by the drone includes multiple independent channels, each with a unique carrier frequency ; Assume that the probability of the LoS link and NLoS link between users is expressed as: (5); (6); in, represents the probability of the line-of-sight link appearing, represents the probability of non-line-of-sight links; a and b are parameters determined by the urban environment; the total path loss formula between the drone and each user is obtained from the above formulas (5) and (6), and its value depends on the distance d and the carrier frequency , as shown in formula (7): (7); in, represents the total path loss between the UAV and each user; The received signal gain is calculated using the Rice channel model as follows: (8); Where g represents the received signal gain, Indicates the remaining energy level of the signal after path loss. ,Right now Determined by path loss; is the Ricean factor, which represents the energy ratio between the LoS link component and the scattering component; represents the complex exponential function, where e is the base of the natural logarithm, represents the imaginary unit, Represents the phase angle.

3. The spectrum scheduling method for UAV emergency communication based on dynamic entropy optimization according to claim 2 is characterized in that: Build spectrum sensing and user access models, identify available channels, and formulate user access strategies; including: The spectrum sensing and user access model includes the random access phase, spectrum sensing phase, spectrum aggregation phase, and network update phase; In the random access phase, M primary users start at the beginning of each time slot based on their own bandwidth requirements. Randomly access multiple channels, namely air-to-ground channels; In the spectrum sensing phase, N secondary users observe idle channels through their own spectrum sensing capabilities. Spectrum sensing capabilities include detection probability and false alarm probability. Indicates the probability of correctly identifying an occupied channel and the false alarm probability The probability of misjudging an idle channel as an occupied channel; the spectrum sensing capability of the secondary user is expressed as spectrum sensing accuracy. express, The higher it is, the more accurate the secondary user's judgment of the channel status is. The formula is as follows: (9); in, represents the spectrum sensing accuracy of the nth secondary user; In the spectrum aggregation stage, secondary users The observation state, i.e., the estimated state of each channel occupancy obtained through spectrum sensing, is obtained. Limited frequency bands are aggregated into different aggregated frequency bands based on spectrum aggregation capabilities. Secondary users select appropriate aggregated frequency bands for access, and access success is determined by base station interaction information. There are two scenarios: successful secondary user access and failed secondary user access. Successful access by a secondary user: Based on the secondary user's correct observation status, when the number of channels in the aggregated frequency band selected by the secondary user is greater than or equal to the secondary user's bandwidth requirement B, and the selected channels do not conflict with any other users, the secondary user will select multiple high-quality channels in the aggregated segment for access. The mobile base station carried by the drone will then provide feedback to the secondary user, including rewards such as transmission rate, data packet loss rate, signal-to-noise ratio, and current channel load. The secondary user will then send multimedia data to the mobile base station carried by the drone. Secondary user access failure: If the secondary user misperceives, the aggregated frequency band does not meet the user's needs, or the selected channel conflicts, the secondary user will abandon the aggregated frequency band and reselect the spectrum in the next time slot based on the previous experience update strategy to retransmit the unsent data. In the last stage of the time slot, namely the access strategy update stage, the user agent will use environmental information and feedback information from the base station to update the access strategy, and continuously optimize the action strategy so that the strategy is gradually transformed into the optimal strategy.

4. The spectrum scheduling method for UAV emergency communication based on dynamic entropy optimization according to claim 3 is characterized in that: Evaluate the channel quality in emergency communication systems; including: The spectrum of the mobile base station carried by the drone is limited. Assuming that the bandwidth of each channel is are the same, the entire spectrum is evenly divided into K channels; at the same time, the transmission power P of each channel is also the same, but the carrier frequency of each channel is unique; therefore, the channel gain of the nth secondary user on the K channels is defined as: (10); in, represents the channel gain of the nth user, Indicates the channel gain of the nth secondary user on the kth channel; sets the signal-to-noise ratio and transmission rate as the standards of channel quality; when there is no conflict between secondary users, the signal-to-noise ratio obtained by the nth secondary user is is defined as: (11); in, is the noise spectral density, is the gain of the ith channel in the segment selected by the nth secondary user, is the gain of the j-th channel of the m-th primary user, Indicates the nth user selection channels, Indicates the mth primary user's selection channels, Indicates the bandwidth requirement of the nth secondary user, N is the number of secondary users, and M is the number of primary users; If there are K channels in the aggregated frequency band selected by the nth secondary user that conflict with other secondary users, the signal-to-noise ratio is represented as: (12); in, represents the gain of the vth conflicting channel between the nth secondary user and the remaining u secondary users, represents the transmission power of the kth user, represents the number of channels in the aggregated frequency band selected by the u-th secondary user that have access conflicts with other secondary users; According to the above SNR formula, the transmission rate of the nth secondary user is Expressed as: (13); Or expressed as: (14); in, It is the loss coefficient of the signal-to-noise ratio, which indicates some loss between the theoretical channel capacity and a specific modulation and coding scheme. If the signal-to-noise ratio is greater than the set threshold and the transmission rate is greater than the set threshold, the channel quality is judged to be high.

5. The method for scheduling spectrum of emergency communication of unmanned aerial vehicles based on dynamic entropy optimization according to claim 4 is characterized in that: Based on spectrum sensing and user access models and channel quality assessment, establish the states, actions, and rewards of deep reinforcement learning; including: Secondary users act as independent intelligent agents. Their state reflects spectrum channel occupancy, including spectrum status and user observation status. Their actions involve selecting a specific aggregated frequency band for access. Rewards are a combination of transmission rate and conflict penalties to encourage users to optimize channel selection. The goal is to maximize long-term discounted rewards. Through policy iteration, secondary users are able to adaptively learn the optimal access strategy in complex spectrum environments. (1) Spectrum state and user observation state: The spectrum is evenly divided into K independent channels, and the state of the spectrum at time slot t is expressed as ,in Indicates the state of the kth channel in time slot t. Each channel has two states: occupied or idle; The observed states of N secondary users are defined as The observation state of the nth secondary user for the entire spectrum at time slot t is expressed as ,in, represents the state of the kth channel perceived by the nth secondary user in time slot t; (2) Action: The action consists of the set of actions of all secondary users, i.e. , Represents the action of the nth sub-user in time slot t. Each sub-user is regarded as an intelligent agent, according to its own observation state Select actions independently ; There are two situations for the selected action: when the selected frequency band can meet the bandwidth requirement B, and the selected channel does not conflict with the channel occupied by the primary user, the secondary user will select the channel with higher channel gain in the aggregated frequency band to meet its own bandwidth requirement B; otherwise, when the aggregated frequency band cannot meet the requirement, the secondary user will give up accessing the channel; according to the aggregation length of the secondary user ( ), the entire spectrum is divided into K- +1 aggregated frequency band; the secondary user will select the most appropriate aggregated frequency band for access based on the action strategy; (3) Reward: The reward set of N users is expressed as ,in represents the reward of the nth secondary user in time slot t; a penalty term f is set. When the secondary user chooses an action that does not meet expectations, the mobile base station will feedback the penalty term to remind the secondary user to update the action strategy; the reward function combines the secondary user's transmission rate and the penalty term f, and is divided into four cases. The reward of the nth secondary user in time slot t is expressed as: (15); The goal of each sub-user is to find an optimal action strategy that maximizes the long-term cumulative discounted benefit, as shown below: (16); in, ∈[0,1] is the discount rate that measures the impact of future rewards on cumulative rewards, and E[] represents the discount rate of the strategy The expected value of the weighted sum of all future rewards under Indicates that in the strategy The long-term cumulative discount benefit under the The total revenue that can be obtained.

6. The spectrum scheduling method for UAV emergency communication based on dynamic entropy optimization according to claim 5 is characterized in that: Build a deep reinforcement learning model and train it to obtain a trained deep reinforcement learning model; including: Build a deep reinforcement learning model, namely the ME-MAAC algorithm reinforcement learning model, including the Actor network, Critic network, target network and experience replay area; The network parameters are The Actor network uses the exploration strategy To choose actions, strategies To observe the status For input value, take action is the output value; the action selection function is described as ; The parameters of the critic network are , the user agent obtains the action and observation status of the current time slot from the Actor network , and input to the Critic network, then the user agent will get the output Q value , to evaluate the action of the current time slot; both the Actor network and the Critic network use the Adam optimizer to optimize the parameters; The ME-MAAC algorithm reinforcement learning model updates the Actor network through the policy gradient method and updates the Critic network by calculating the loss function; When training the network as an intelligent agent, each secondary user randomly samples a small batch of observation state and action data from the experience replay area, and then transmits this data to the Actor network and the Critic network to calculate the Q value respectively. and action-value function ; A maximum entropy method is introduced in the ME-MAAC algorithm, which adds an entropy term that supports random strategies to the policy gradient and loss function. , then the policy gradient of the nth user is expressed as: (17); in, represents the policy gradient of the nth user, Indicates network parameters The derivative operation of is the control factor that controls the randomness of the optimal strategy, D represents the experience replay area, Indicates the state sampled from the experience playback area D and actions The expected value of Indicates that in a given state Take action The Q value is estimated by the Critic network with parameter θ; The nth user samples a certain number of experience units , respectively, through the Critic network and its target network to obtain the Q value and target Q value ,in Represents the next observed state, and then adds the target Q value and the entropy term obtained from the target Actor network Combined to calculate the output value of the target network , the secondary user updates the Critic network by minimizing the regression loss function, which is defined as: (18); (19); in, represents the loss function of the Critic network of the nth user, Represents the experience unit sampled from the experience playback area D The expected value of Indicates that given the next observation state Next, the action distribution generated by the target Actor network Action expected value; The network structure of the target network is the same as that of the source network, namely the Actor network and the Critic network, and the parameters of the target Actor network and the target Critic network in the target network are and ; The update formula of the target network parameters is expressed as: (20); (21); Where τ∈[0,1] is the soft update coefficient; In each time slot, the user integrates all the data obtained from interacting with the environment, namely the observed state, action and reward, into a set of experience units. , and stored in the experience replay area D; in each network training process, the secondary user randomly extracts experience units to update the network parameters; through the target network and the source network, the training is continuously updated to obtain a trained deep reinforcement learning model.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for scheduling spectrum of unmanned aerial vehicle emergency communications based on dynamic entropy optimization as described in any of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for scheduling spectrum of unmanned aerial vehicle emergency communications based on dynamic entropy optimization according to any of claims 1 to 6 are implemented.

9. A spectrum scheduling system for UAV emergency communications based on dynamic entropy optimization, comprising: The emergency communication system building module is configured to: build an emergency communication system, perform communication between users and mobile base stations, and calculate received signal gain; The spectrum sensing and user access model building module is configured to: build spectrum sensing and user access models, identify available channels, and formulate user access policies; The channel quality assessment module is configured to: assess the channel quality in the emergency communication system; A reinforcement learning parameter establishment module is configured to establish states, actions, and rewards for deep reinforcement learning based on spectrum sensing and user access models and channel quality assessments; The reinforcement learning model building module is configured to: build a deep reinforcement learning model and perform training to obtain a trained deep reinforcement learning model; The spectrum aggregation module is configured to use the trained deep reinforcement learning model to obtain the optimal strategy for spectrum scheduling of drone emergency communications in the spectrum environment and achieve aggregated spectrum access for multiple users.

Citation Information

Cited By

  • Entropy driving step length self-adaption-based diffusion reinforcement learning channel access method

    CN121463143A

  • An entropy-driven step length self-adaptive diffusion reinforcement learning channel access method

    CN121463143B

  • Joint optimization method for framing time and bandwidth allocation of emergency communication network

    CN122120947A

  • A Joint Optimization Method for Frame Time and Bandwidth Allocation in Emergency Communication Networks

    CN122120947B