Low-altitude Internet of Things dynamic spectrum allocation and access method and system based on deep reinforcement learning
By adopting deep reinforcement learning MAEAC algorithm in low-altitude intelligent networked drone networks, the problem of communication quality assurance of drone networks under limited spectrum resources is solved, efficient spectrum resource allocation and access is achieved, and stable communication of drone networks is ensured.
Patent Information
- Application Number
- CN202510564543.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-20
AI Technical Summary
Low-altitude intelligent networked drone networks face the problem of communication quality assurance under limited spectrum resources, especially in dynamic multi-channel spectrum environments. How to effectively allocate spectrum resources to ensure stable communications of drone networks has become a key issue that needs to be solved urgently.
Using the multi-agent actor-critic algorithm (MAEAC) based on deep reinforcement learning, each drone is equipped with an actor-critic network, and a maximum entropy method is introduced into the network for optimization to form a MAEAC reinforcement learning model. Through dynamic spectrum perception and aggregation, this model optimizes the spectrum access strategy of the drone in each time slot to ensure efficient utilization of spectrum resources.
It realizes stable communication of low-altitude intelligent networked drone network under limited spectrum resources, improves system performance and cooperation efficiency between drones, avoids local optimal solutions, and significantly improves the effects of spectrum allocation and access.
Smart Images

Figure CN120186773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for dynamic spectrum allocation and access in a low-altitude intelligent network based on deep reinforcement learning, belonging to the fields of wireless communication technology, cognitive radio technology, and low-altitude network. Background Art
[0002] With the rapid development and improvement of 5G technology and the emergence and development of 6G technology, the development of society and technology will enter a new level. 5G has the characteristics of "ultra-high data rate", "ultra-low latency", "massive connection ability", and "high reliability and stability", enabling the rapid development of the Low-Altitude Intelligent Network (LAIN). However, due to the huge scale of the drone network in the low-altitude intelligent network, there has always been a problem of insufficient spectrum resources.
[0003] The rise of the low-altitude economy has promoted the expansion of the traditional transportation network and the ground Internet into the three-dimensional space, realizing the transformation of the digital economy layout from "plane" to "three-dimensional", forming a new digital economic form, and giving birth to a new space for trillion-dollar industries. As the digital foundation for the development of the low-altitude economy, the low-altitude intelligent network serves the "communication, monitoring, guidance, meteorology, computing" needs of low-altitude applications and is the basic technical guarantee for applications such as low-altitude logistics, urban governance, and air traffic.
[0004] However, with the development of the low-altitude intelligent network and the increasing scale of the drone network, the difficulty of ensuring the communication quality of the drone network will be further increased, and a huge contradiction will be formed between the limited spectrum resources and the huge drone network. How to allocate the limited spectrum resources to the drone network in a dynamic spectrum environment to ensure its communication has become a key problem that urgently needs to be solved. To ensure the development of the low-altitude intelligent network, it is necessary to face the problem directly and actively seek solutions.
[0005] As a tool for intelligent decision-making and control, reinforcement learning has received extensive attention in the field of communication resource allocation. The communication resource allocation problem is essentially a complex dynamic optimization problem, whose goal is to maximize the system performance under limited communication resources (such as spectrum, power, and time), while ensuring the fairness and quality of service of multiple users. Traditional resource allocation methods usually rely on preset rules or static optimization models. However, in modern communication networks, due to the dynamic changes in user requirements, the complexity of the network environment, and the high-speed movement of devices, these traditional methods are often difficult to cope with. Reinforcement learning can achieve real-time optimization of resource allocation in a complex dynamic environment by interacting with the environment to learn the optimal strategy, especially showing great potential in multi-user and multi-task scenarios. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technologies, the present invention provides a method and system for dynamic spectrum allocation and access in a low-altitude intelligent network based on deep reinforcement learning. The communication system of the low-altitude intelligent network drone network includes a data center, a communication satellite, and a drone network. The data center is used to manage and analyze the data of the drone network. The communication satellite is a data communication transfer station between the data center and the drone network. The drone network is the part directly participating in the management of the low-altitude area in the low-altitude intelligent network. The drone network is extremely large. The communication satellite needs to communicate with the entire drone network under limited spectrum resources and face a dynamic multi-channel spectrum environment.
[0007] Term Explanation: 1. Reinforcement Learning (RL): Also known as re-inforcement learning, evaluation learning, or enhancement learning, it is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent achieves maximum reward or realizes a specific goal by learning strategies during the interaction with the environment.
[0008] 2. Deep Reinforcement Learning (DRL): It is a technology that combines reinforcement learning (RL) and deep learning (DL). Reinforcement learning is a method of machine learning, aiming to guide an agent to make decisions in the environment through a reward signal to achieve the goal of maximizing long-term rewards. Deep learning uses neural networks for feature extraction and decision-making modeling. In deep reinforcement learning, the agent interacts with the environment and learns strategies based on the rewards or punishments generated by its actions. Neural networks are usually used to approximate the value function, policy function, or directly predict future actions.
[0009] 3. Agent: It refers to an agent that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system, with autonomy, adaptability, and interaction capabilities. The agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or achieve a predetermined goal.
[0010] 4. AC Network: actor-critic network, which is a reinforcement learning method that combines policy gradient and value function. This algorithm consists of two parts: the actor and the critic network. The actor is responsible for selecting actions according to the current policy, while the critic evaluates the value of these actions and provides feedback to optimize the actor's policy.
[0011] 5. MAEAC (Multi-Agent Maximum Entropy Actor-Critic): A multi-agent maximum entropy actor-critic network reinforcement learning algorithm.
[0012] 6. Maximum Entropy is an important concept in statistics and information theory, widely used in fields such as machine learning, natural language processing, and signal processing. Its core idea is to select the distribution with the maximum entropy as the model among all probability distributions that satisfy the known constraints. This can ensure that the selected model makes the fewest assumptions about unknown information, thereby achieving the goal of maximizing information uncertainty.
[0013] 7. The learning rate, usually denoted as, is an important hyperparameter in optimization algorithms for machine learning and deep learning. It determines the step size of the model during each parameter update, thus affecting the training speed of the model and the stability of the optimization process.
[0014] 8. The discount factor is a key parameter in reinforcement learning used to measure the importance weight of future rewards in the current value.
[0015] 9. The exploration rate is one of the important parameters in reinforcement learning for balancing exploration and exploitation. It is a value between 0 and 1 used to control the probability of randomly selecting an action.
[0016] The technical solution of the present invention is as follows: The first aspect of the present invention provides a method for dynamic spectrum allocation and access in a low-altitude intelligent network based on deep reinforcement learning, including: S1. Construct a drone network for the low-altitude intelligent network; S2. Each drone is equipped with an actor-critic network, and the maximum entropy method is introduced into the actor-critic network for optimization to form a MAEAC reinforcement learning model; S3. Train the MAEAC reinforcement learning model to obtain a trained MAEAC reinforcement learning model and acquire the optimal strategy for spectrum allocation and access.
[0017] Preferably according to the present invention, constructing a drone network for the low-altitude intelligent network includes: The drone network includes M drones, where M drones access the communication satellite dynamically through N channels; Each drone has a bandwidth W m and spectrum sensing and spectrum aggregation capabilities, where the spectrum sensing ability determines the detection probability of the m th drone , the detection probability is used to judge the channel state; the spectrum aggregation ability represents the length of the spectrum segment aggregated by the m th drone L m ; Drones randomly access multiple channels in each time slot, and in each time slot, the drones are sorted according to the bandwidth from large to small to avoid conflicts; Each drone sequentially relies on spectrum sensing to sense the channel, and through spectrum aggregation, discrete idle channels are aggregated into a spectrum segment. Finally, according to the bandwidth W m selects a suitable aggregated spectrum segment for access; when the m th drone correctly senses the channel state and the number of idle channels in the aggregated spectrum segment selected by the m th drone is greater than or equal to W m , the drone successfully accesses and then returns a success feedback to the drone; When the drone fails to access due to sensing errors, or there is no applicable aggregated spectrum segment in the entire spectrum, resulting in the drone being unable to access, the drone receives negative feedback; therefore, the goal is to enable as many drones as possible to access the spectrum to optimize the system performance.
[0018] Preferably according to the present invention, each drone is equipped with an actor-critic network, and the maximum entropy method is introduced in the actor-critic network for optimization to form a MAEAC reinforcement learning model; including: To solve the dynamic spectrum sensing and aggregation problems, a distributed execution and training framework is designed, and a MAEAC algorithm is proposed; among them, the actor-critic network includes an actor network and a critic network; the input of each drone actor-critic network is the spectrum environment of the drone, that is, the channel state; The actor network selects the action of the m th drone according to the o m observed state of the m th drone , and executes it in the environment , and then obtains a reward r m feedback and a new (the observed state of the m th drone in the next time slot); The critic network evaluates the value of the action and outputs an evaluation of the current drone action to optimize the strategy of the actor network; Both the channel state and the observation state of the UAV include: the idle (1) state, i.e., the channel is not used by any user; the occupied (0) state, i.e., the channel is already occupied; the channel state is the actual state of the channel, and the observation state of the UAV is the channel state observed by the UAV; Let represent the channel state at each time slot t , where ; The UAV can only rely on its spectrum sensing ability to observe part of the channel state. Denote the observation states of M UAVs as ; Let be the observation state of the t -th UAV at time slot m , where represents the observation state of the n -th UAV on the L m -th channel (n ∈ L m ), and the observation length of the UAV is equal to the aggregated spectrum segment length Aggregated spectrum segment length L m < N , the spectrum is divided into N - L m +1 segments; Denote the actions of M UAVs as , let be the action of the m -th UAV at time slot t, where means that the m -th UAV selects the t -th spectrum segment of the i -th spectrum at time slot ; When the UAV executes an action, the spectrum will return a reward including the UAV access information. Denote the rewards of M UAVs as , and set the reward of the t -th UAV at each time slot m as: : ; (1) The ultimate goal of each UAV is to find an optimal policy to maximize the long-term discounted reward, expressed as: ; (2) Among them, is the discount factor (discount coefficient), Represents a long-term discount reward (cumulative discount reward), D Represents the spectrum environment; represents the channel state, E [ ] is the expected function, Represents performing deep reinforcement learning in the spectrum environment.
[0019] According to a preferred embodiment of the present invention, the maximum entropy method is introduced into the actor-critic network for optimization to form the MAEAC reinforcement learning model; including: The parameters of the MAEAC reinforcement learning model include the main network parameters and the target network parameters. The main network is the actor network and the critic network, and the target network is the target actor network and the target critic network; the main network and the target network have the same structure. The update process of the main network is to update all parameters each time, and the update process of the target network is to update only part of the parameters each time, and part of the parameters are copied from the main network; The update of the actor network is calculated through policy gradients. In order to add some random exploration during the reinforcement learning process and introduce the maximum entropy method into the policy gradients, that is, by introducing an entropy term to modify the policy, α represents the balance factor, which affects the balance between maximizing entropy and rewards. The policy gradient of the m-th drone is as follows: ; (3) Wherein, Represents the policy of the m-th drone π m Regarding the parameter Gradient of, E [ ] is the expected function, B represents the spectrum environment, Is the m Value obtained by the Q -th drone, Represents the policy network of the m-th drone, Represents the parameters of the policy network of the m-th drone, Represents the action of the m-th drone, o m Represents the observed state of the m-th drone, Represents the introduced entropy term; The critic network is updated by minimizing the regression loss function (minimizing the regression loss function to update , that is, updating the critic network). The m -th drone samples the experience unit and obtains the Q value through the critic network and the target network respectively, and then combines the target Q value and the introduced entropy term To calculate the output value y of the target network m, the critic network is updated by minimizing the regression loss function as follows: ; (4) ; (5) where denotes the loss function, is the expectation function, represents the m -th value of the Q -th drone, y m represents the predicted Q value (representing the output value of the target network, or the target Q value output by the target network), represents the set of experience units sampled by the r m -th drone, m represents the reward of the -th drone for the next time slot, represents the observed state of the -th drone for the next time slot, m represents the Q value of the target network of the B -th drone for the next time slot, and γ is the discount factor (discount coefficient); each drone stores the set of experience units in the memory pool ; in the MAEAC reinforcement learning model, each drone, under the channel state s , perceives the observed state according to the detection probability o m (the observed state of the m -th drone at all time slots), selects a spectrum segment that meets the bandwidth W m and correctly senses it for access; the drone marks the selected spectrum segment to avoid repeated selection by other drones and prevent conflicts; then the drone receives channel feedback and updates the action selection policy π m (i.e., the policy in the policy update and optimization process).
[0020] According to the preferred embodiment of the present invention, the MAEAC reinforcement learning model is trained to obtain a trained MAEAC reinforcement learning model, and the optimal policy for spectrum allocation and access is obtained; including: After the actor network of the UAV executes an action based on the observation state of the UAV, the environment updates the state and returns a reward. Then the UAV obtains a new observation state from the updated environment, and stores the observation state, action, reward, and new observation state as an experience unit group in the memory pool; Sample a mini-batch of experience unit groups from the memory pool, update the critic network by minimizing the regression loss function, and update the actor network through policy gradient; use soft update to update the target network; Loop the above steps. Finally, the MAEAC reinforcement learning model gradually converges to obtain a trained MAEAC reinforcement learning model; Among them, the target network is used to copy the main network parameters in each iteration, as follows: ; (6) Among them, 、 are the main network parameters, 、 are the target network parameters, τ is the soft update factor; represents the parameterized policy function, that is, the probability distribution of generating an action according to the current channel state; the goal of the actor network is to learn the optimal policy to maximize the long-term cumulative reward; represents the parameterized value function, such as the state value function or action value function used to evaluate the expected return of the current state or action; Through continuous update and training of the main network and the target network, a trained MAEAC reinforcement learning model is obtained, and the trained MAEAC reinforcement learning model is used for dynamic spectrum allocation and access of the low-altitude intelligent Internet of Things to obtain the optimal policy for spectrum allocation and access.
[0021] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for dynamic spectrum allocation and access of the low-altitude intelligent Internet of Things based on deep reinforcement learning.
[0022] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for dynamic spectrum allocation and access of the low-altitude intelligent Internet of Things based on deep reinforcement learning.
[0023] The second aspect of the present invention provides a dynamic spectrum allocation and access system for a low-altitude intelligent Internet of Things based on deep reinforcement learning, including: A UAV network construction module configured to: construct a UAV network for the low-altitude intelligent Internet of Things; The reinforcement learning algorithm construction module is configured to: each drone is equipped with an actor-critic network, and the maximum entropy method is introduced into the actor-critic network for optimization to form a MAEAC reinforcement learning model; The optimal policy acquisition module is configured to: train the MAEAC reinforcement learning model to obtain a trained MAEAC reinforcement learning model, and obtain the optimal policy for spectrum allocation and access.
[0024] The beneficial effects of the present invention are as follows: 1. The present invention provides a method for dynamic spectrum allocation and access based on deep reinforcement learning with excellent performance and strong adaptability to face the situation where the drone network in the low-altitude intelligent Internet needs to achieve stable communication under limited spectrum resources.
[0025] 2. The present invention proposes a multi-agent actor-critic algorithm based on maximum entropy to solve the problem of dynamic spectrum allocation and access. This algorithm optimizes the updates of the actor network and the critic network by adding an entropy term, increasing the random exploration of reinforcement learning and avoiding falling into local optimal solutions.
[0026] 3. In the present invention, the cooperation between drones effectively reduces the negative impact when drones mis-sense the channel state. When drones have different sensing capabilities, the proposed algorithm has better performance and faster convergence speed than DQN.
[0027] 4. The method for dynamic spectrum allocation and access of drones in the low-altitude intelligent Internet based on multi-agent actor-critic reinforcement learning proposed by the present invention solves the problem of how to achieve stable communication in the drone network in the low-altitude intelligent Internet under limited spectrum resources and achieves good results. Description of the Drawings
[0028] Figure 1 It is a model diagram of the low-altitude intelligent Internet communication system of the present invention; Figure 2 It is a structural diagram of the MAEAC reinforcement learning model of the present invention; Figure 3 For four drones with the same detection probability =0.9 and different aggregation lengths L m and bandwidth W m It is a schematic diagram of the cumulative discounted reward of the system; Figure 4 For four drones with the same detection probability =0.9 and different aggregation lengths L m and bandwidthW m Schematic diagram of the cumulative discount rewards of four UAVs Figure 5 For two UAVs having the same aggregation length L m = 8 and bandwidth W m = 4, different detection probabilities Schematic diagram of the cumulative discount rewards of two UAVs and the system Figure 6 Comparison chart of the cumulative discount rewards under different detection probabilities Figure 7 Schematic diagram of the successful access rate of UAVs under different bandwidths Specific implementation manners
[0029] The present invention will be further described below by way of examples in conjunction with the accompanying drawings, but is not limited thereto
[0030] Example 1 A method for dynamic spectrum allocation and access in a low-altitude intelligent network based on deep reinforcement learning, comprising: S1. Construct a UAV network of the low-altitude intelligent network S2. Each UAV is equipped with an actor-critic network, and the maximum entropy method is introduced in the actor-critic network for optimization to form a MAEAC reinforcement learning model S3. Train the MAEAC reinforcement learning model to obtain a trained MAEAC reinforcement learning model, and obtain an optimal policy for spectrum allocation and access
[0031] Example 2 A method for dynamic spectrum allocation and access in a low-altitude intelligent network based on deep reinforcement learning according to Example 1, wherein the difference lies in: Construct a UAV network of the low-altitude intelligent network; as Figure 1 shown, comprising: The UAV network includes M UAVs, wherein M UAVs dynamically access a communication satellite through N channels Each UAV has a bandwidth W m , spectrum sensing ability and spectrum aggregation ability, wherein the spectrum sensing ability determines the detection probability m of the th UAV , and the detection probability judges the channel state; the spectrum aggregation ability represents the length m of the spectrum segment aggregated by the th UAV L m ; The UAV randomly accesses multiple channels in each time slot, and in each time slot, the UAVs are sorted from largest to smallest according to the bandwidth to avoid conflicts; Each UAV successively relies on spectrum sensing to sense the channels, and aggregates the discrete idle channels into a spectrum segment through spectrum aggregation. Finally, according to the bandwidth W m selects a suitable aggregated spectrum segment for access; when the m th UAV correctly senses the channel state and the number of idle channels in the aggregated spectrum segment selected by the m th UAV is greater than or equal to W m , the UAV successfully accesses and then returns a success feedback to the UAV; When the UAV fails to access due to sensing errors, or there is no applicable aggregated spectrum segment in the entire spectrum, resulting in the UAV being unable to access, the UAV receives negative feedback; Therefore, the goal is to enable as many UAVs as possible to access the spectrum to optimize the system performance.
[0032] Each UAV is equipped with an actor-critic network, and the maximum entropy method is introduced for optimization in the actor-critic network to form a MAEAC reinforcement learning model; as Figure 2 shown, it includes: To solve the dynamic spectrum sensing and aggregation problems, a distributed execution and training framework is designed, and a MAEAC algorithm is proposed; among them, the actor-critic network includes an actor network and a critic network; the input of each UAV actor-critic network is the spectrum environment of the UAV, that is, the channel state; The actor network selects the action of the m th UAV according to the observation state of the o m th UAV m and executes it in the environment , and then obtains the reward feedback and the new r m (the observation state of the th UAV in the next time slot); m The critic network evaluates the value of the action and outputs the evaluation of the current UAV action to optimize the strategy of the actor network; The channel state and the observation state of the UAV both include: the idle (1) state, i.e., the channel is not used by any user; the occupied (0) state, i.e., the channel is already occupied; the channel state is the actual state of the channel, and the observation state of the UAV is the channel state observed by the UAV; Let denote the channel state at each time slot t where ; The UAV can only rely on its spectrum sensing ability to observe part of the channel state. Denote the observation states of M UAVs as ; Let be the observation state of the t -th UAV at time slot m where represents the observation state of the m-th UAV on the n -th channel (n ∈ L m ). The observation length of the UAV is equal to the aggregated spectrum segment length L m , which also represents the number of channels observed by the UAV; The aggregated spectrum segment length L m < N , and the spectrum is divided into N - L m +1 segments; Denote the actions of M UAVs as , and let be the action of the m -th UAV at time slot t, where means that the m -th UAV selects the t -th spectrum segment at time slot i ; When the UAV executes an action, the spectrum will return a reward including the UAV access information. Denote the rewards of M UAVs as , and set the reward of the t -th UAV at each time slot m as: ; (1) The ultimate goal of each UAV is to find an optimal policy to maximize the long-term discounted reward, which is expressed as: ; (2) where is the discount factor (discount coefficient), Represents a long-term discount reward (cumulative discount reward). D Represents the spectrum environment; s represents the channel state. E [ ] is the expectation function. Represents performing deep reinforcement learning in the spectrum environment. Introduce the maximum entropy method in the actor-critic network for optimization to form the MAEAC reinforcement learning model, including: The parameters of the MAEAC reinforcement learning model include the main network parameters and the target network parameters. The main network is the actor network and the critic network, and the target network is the target actor network and the target critic network. The main network and the target network have the same structure. The update process of the main network is to update all parameters each time, and the update process of the target network is to update only some parameters each time, and some parameters are copied from the main network. The update of the actor network is calculated through policy gradients. In order to add some random exploration during the reinforcement learning process, the maximum entropy method is introduced in the policy gradients, that is, by introducing an entropy term to modify the policy. α represents the balance factor, which affects the balance between maximizing entropy and rewards. The policy gradient of the m-th drone is as follows:[[]] ; (3) Among them,[[]] Represents the policy of the m-th drone π m Regarding the parameter Gradient of,[[]] E [ ] is the expectation function. B Represents the spectrum environment. Is the m Value obtained by the Q -th drone,[[]] Represents the policy network of the m-th drone. Represents the parameters of the policy network of the m-th drone. Represents the action of the m-th drone. o m Represents the observation state of the m-th drone. Represents the introduced entropy term. The critic network is updated by minimizing the regression loss function (minimize the regression loss function to update , that is, update the critic network). The m -th drone samples the experience unit and obtains the Q value through the critic network and the target network respectively, and then combines the target Q value and the introduced entropy term To calculate the output value y of the target network m, the critic network is updated by minimizing the regression loss function as follows: ; (4) ; (5) where represents the loss function, is the expectation function, represents the m -th value of the Q -th drone, y m represents the predicted Q value (which represents the output value of the target network, or the target Q-value output by the target network), represents the empirical unit group sampled by the r m -th drone, m represents the reward of the -th drone, represents the observed state of the -th drone in the next time slot, m represents the Q value of the next time slot of the target network of the -th drone, and γ is the discount factor (discount coefficient); each drone stores the empirical unit group in the memory pool B and randomly selects a part of the empirical unit group in each iteration; (the code setting is to sample 128 empirical unit groups each time), which helps to break the correlation between data and improve the convergence of training; In the MAEAC reinforcement learning model, each drone, under the channel state s , senses the observed state according to the detection probability o m (the observed state of the m -th drone in all time slots), selects a spectrum segment that meets the bandwidth W m and correctly senses it for access; the drone marks the selected spectrum segment to avoid repeated selection by other drones and prevent conflicts; then the drone receives channel feedback and updates the action selection policy π m (i.e., the policy in the policy update and optimization process).
[0033] Train the MAEAC reinforcement learning model to obtain a trained MAEAC reinforcement learning model and obtain the optimal policy for spectrum allocation and access; including: After the actor network of the UAV executes an action based on the observed state of the UAV, the environment updates the state and returns a reward. Then the UAV obtains a new observed state from the updated environment, and stores the observed state, action, reward, and new observed state as an experience unit group in the memory pool; Sample a mini-batch of experience unit groups from the memory pool, update the critic network by minimizing the regression loss function, and update the actor network by policy gradient; update the target network using soft update; Loop the above steps. Finally, the MAEAC reinforcement learning model gradually converges to obtain a trained MAEAC reinforcement learning model; Among them, the target network is used to copy the main network parameters in each iteration, as follows: ; (6) Among them, 、 are the main network parameters, 、 are the target network parameters, τ is the soft update factor; represents the parameterized policy function, that is, the probability distribution of generating actions according to the current channel state; the goal of the actor network is to learn the optimal policy to maximize the long-term cumulative reward; represents the parameterized value function, such as the state value function or the action value function used to evaluate the expected return of the current state or action; Through continuous update training of the main network and the target network, a trained MAEAC reinforcement learning model is obtained. The trained MAEAC reinforcement learning model is used for dynamic spectrum allocation and access in the low-altitude intelligent Internet of Things to obtain the optimal policy for spectrum allocation and access.
[0034] The performance of the MAEAC algorithm is analyzed in the form of simulation experiments. The performance of the MAEAC method is evaluated by comparing the simulation results of the MAEAC algorithm and the DQN algorithm. In the simulation, the maximum number of training rounds of the two algorithms is set to 200000, and the batch size is set to 128. The learning rate β of the Adam optimizer is 0.001, the discount factor γ is 0.9, and the exploration rate of DQN is 0.9→0.
[0035] First, when the number of channels is 32, the cumulative discounted rewards of two UAVs and the entire system in the two algorithms are compared. The reward of the system is the average of the rewards of all UAVs. From Figure 3 and Figure 4 it can be seen that the rewards of the two algorithms increase with the increase of the number of iterations. The rewards of the four UAVs and the system in MAEAC are higher than those in DQN, and the convergence speed of MAEAC is faster. It can also be concluded that in MAEAC, asL m As the [elevation] increases, the reward for the UAV is higher because when the [elevation] of the UAV is L m higher, more idle channels can be aggregated to meet W m .
[0036] For example, Figure 5 as shown, at different ( = 0.1 or = 0.9), the reward of MAEAC is higher than that of DQN.
[0037] Next, the performance comparison between MAEAC and DQN under different detection probabilities was studied. In Figure 6 , the number of channels is different, that is, N = {16, 32, 64}. As increases, that is, the correct rate of the UAV sensing the channel state increases, the system rewards of both algorithms show an upward trend. When the [elevation] and [azimuth] of the UAV in each system are L m and W m are the same and the detection probability = {0.3, 0.5, 0.7, 0.9} of each system, it can be clearly observed that as the probability decreases, the performance of DQN becomes increasingly inferior to that of MAEAC because the strategy of MAEAC effectively reduces the negative impact when the UAV mis-senses the channel state. When <0.5, the system with 64 channels obtains the highest reward because a large number of available channels may have more correct segments, which can make the system reward higher.
[0038] Set L m = {8, 9, 10}, W m = {3, 4, 5, 6} and is the same. In Figure 7 , as L m increases, the successful access rate increases, but as W m increases, the successful access rate decreases. This is because it is more difficult for the limited L m to meet the larger W m , while the longer L m may have more idle channels. The access success rate of MAEAC is greater than that of DQN, especially when L m = 8 and It should be noted that some of the specific terms like "[elevation]" and "[azimuth]" are placeholders in the original text and might need to be further defined according to the actual context.W m When it is {3, 6}, the advantages of MAEAC are more obvious. This once again proves that the performance of MAEAC is better than that of DQN.
[0039] Example 3 A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of the method for dynamic spectrum allocation and access of a low-altitude intelligent network based on deep reinforcement learning described in Example 1 or 2.
[0040] Example 4 A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the steps of the method for dynamic spectrum allocation and access of a low-altitude intelligent network based on deep reinforcement learning described in Example 1 or 2.
[0041] Example 5 A system for dynamic spectrum allocation and access of a low-altitude intelligent network based on deep reinforcement learning includes: A UAV network construction module configured to construct a UAV network of a low-altitude intelligent network; A reinforcement learning algorithm construction module configured to: each UAV is equipped with an actor-critic network, and the maximum entropy method is introduced for optimization in the actor-critic network to form a MAEAC reinforcement learning model; An optimal policy acquisition module configured to: train the MAEAC reinforcement learning model to obtain a trained MAEAC reinforcement learning model and obtain an optimal policy for spectrum allocation and access.
Claims
1. A method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning, characterized in that: include: S1. Build a drone network for low-altitude intelligent networking; S2, each drone is equipped with an actor-critic network, and the maximum entropy method is introduced into the actor-critic network for optimization to form a MAEAC reinforcement learning model; S3. Train the MAEAC reinforcement learning model to obtain the trained MAEAC reinforcement learning model and obtain the optimal strategy for spectrum allocation and access.
2. According to claim 1, a method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning is characterized in that: Construct a drone network for low-altitude intelligent networking; including: The drone network includes M drones, including M Drones pass through N Channels dynamically access the communication satellite; Each drone has bandwidth W m , spectrum sensing capability and spectrum aggregation capability, among which spectrum sensing capability determines the m The detection probability of a drone , the detection probability determines the channel state; the spectrum aggregation capability indicates the m The length of the spectrum segment aggregated by drones L m ; The drones randomly access multiple channels in each time slot, and in each time slot, the drones are sorted from large to small according to bandwidth; Each drone relies on spectrum sensing to sense the channel in turn, and aggregates discrete idle channels into a spectrum segment through spectrum aggregation, and finally allocates them according to bandwidth. W m Select the appropriate aggregate spectrum segment for access; m The first UAV correctly senses the channel status and m The number of idle channels in the aggregated spectrum band selected by the drones is greater than or equal to W m When the UAV is successfully connected, a success feedback is returned to the UAV; When a drone perception error causes access failure, or there is no applicable aggregate spectrum segment in the entire spectrum, resulting in the drone being unable to access, the drone receives negative feedback.
3. According to claim 2, a method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning is characterized in that: Each drone is equipped with an actor-critic network, and the maximum entropy method is introduced into the actor-critic network for optimization to form a MAEAC reinforcement learning model; including: The actor-critic network consists of an actor network and a critic network. The input of each drone actor-critic network is the drone’s spectrum environment, i.e., the channel state. The actor network is based on m The observation status of the drone o m Select m Drone action , and execute in the environment , then get rewarded r m Feedback and new ; The critic network evaluates the value of the action and outputs the evaluation of the current drone action. To optimize the strategy of actor network; Both the channel state and the drone’s observed state include: idle state, i.e., the channel is not used by any user; occupied state, i.e., the channel is occupied; the channel state is the actual state of the channel, and the drone’s observed state is the channel state observed by the drone; set up Indicates each time slot t The channel state at time ;Will M The observation state of a UAV is expressed as ;set up It is a time slot t Time m The observation status of the UAV, Indicates that the mth drone is in n The observation state on the channel, the observation length of the drone is equal to the length of the aggregated spectrum segment L m , which also represents the number of channels observed by the UAV; Aggregate spectrum segment length L m < N The spectrum is divided into N - L m +1 paragraph; M The action of a drone is represented as ,set up It is m The action of a UAV at time slot t, where It refers to m UAVs in time slot t Selected i a spectrum segment of a spectrum; When the drone performs an action, the spectrum will return a reward including the drone's access information. M The reward of a drone is expressed as , and each time slot t The m Drone Rewards Set to: ;(1) The ultimate goal of each drone is to find an optimal strategy , which maximizes the long-term discounted reward, expressed as: ;(2) in, is the discount factor, represents the long-term discount reward, D Indicates the spectrum environment; indicates the channel status, E [ ] is the expected function, It means deep reinforcement learning in a spectrum environment.
4. According to claim 3, a method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning is characterized in that: The maximum entropy method is introduced into the actor-critic network for optimization to form the MAEAC reinforcement learning model; including: The parameters of the MAEAC reinforcement learning model include the main network parameters and the target network parameters. The main network is the actor network and the critic network, and the target network is the target actor network and the target critic network. The main network has the same structure as the target network. The update process of the main network is to update all parameters each time, while the update process of the target network is to update only some parameters each time, and some parameters are copied from the main network. The update of the actor network is calculated through policy gradient, and the maximum entropy method is introduced in the policy gradient, that is, by introducing the entropy term To modify the strategy, α represents the balancing factor, and the policy gradient of the mth drone is as follows: ;(3) in, represents the strategy of the mth drone π m About parameters The gradient of E [] is the expected function, B Indicates the spectrum environment, It is m The drone got Q value, represents the policy network of the mth drone, represents the parameters of the policy network of the mth UAV, represents the action of the mth drone, o m represents the observation state of the mth UAV, represents the introduced entropy term; The critic network is updated by minimizing the regression loss function , No. m Each drone samples the experience unit and obtains the Q value through the critic network and the target network respectively, and then adds the target Q value and the introduced entropy term Combined to calculate the output value y of the target network m , update the critic network by minimizing the regression loss function as follows: ;(4) ;(5) in, represents the loss function, is the expectation function, Indicates m drone Q Value, y m Indicates the predicted Q value, represents the experience unit group sampled by the mth drone, r m Indicates m A drone reward. represents the observation state of the mth UAV in the next time slot, represents the action of the mth drone in the next time slot, Indicates m The next time slot of the target network of the drone Q value, γ is the discount factor; each drone stores the experience unit group in the memory pool B In each iteration, some experience unit groups are randomly selected; In the MAEAC reinforcement learning model, each drone is in the channel state s Next, according to the detection probability To perceive the observation state o m , select the bandwidth that satisfies W m The drone marks the selected spectrum segment and then receives channel feedback and updates the action selection strategy. π m .
5. According to claim 4, a method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning is characterized in that: Train the MAEAC reinforcement learning model to obtain the trained MAEAC reinforcement learning model and obtain the optimal strategy for spectrum allocation and access; including: After the drone's actor network performs an action based on the drone's observation state, the environment updates its state and returns a reward. The drone then obtains a new observation state from the updated environment and stores the observation state, action, reward, and new observation state as an experience unit group in the memory pool. Sample a small batch of experience units from the memory pool, update the critic network by minimizing the regression loss function, update the actor network by policy gradient; update the target network using soft update; The above steps are repeated, and finally the MAEAC reinforcement learning model gradually converges to obtain a trained MAEAC reinforcement learning model; Here, the target network is used to replicate the parameters of the main network in each iteration as follows: ;(6) in, , are the main network parameters, , are the target network parameters, τ is the soft update factor; Through continuous updating and training of the main network and the target network, a trained MAEAC reinforcement learning model is obtained. The trained MAEAC reinforcement learning model is used to perform dynamic spectrum allocation and access of the low-altitude intelligent network to obtain the optimal strategy for spectrum allocation and access.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning as described in any one of claims 1-5 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for dynamic spectrum allocation and access of low-altitude intelligent network based on deep reinforcement learning as described in any one of claims 1-5 are implemented.
8. A low-altitude intelligent network dynamic spectrum allocation and access system based on deep reinforcement learning, characterized in that: include: The reinforcement learning algorithm building module is configured as follows: each drone carries an actor-critic network, and the maximum entropy method is introduced into the actor-critic network for optimization to form a MAEAC reinforcement learning model; The optimal strategy acquisition module is configured to: train the MAEAC reinforcement learning model, obtain the trained MAEAC reinforcement learning model, and obtain the optimal strategy for spectrum allocation and access.
Citation Information
Cited By
Spectrum decision-making method, device and equipment based on double-graph structure and reconstruction cost
CN122269290A
Spectrum decision method and device based on double graph structure and reconstruction cost, and equipment
CN122269290B