Heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing
By introducing HASAC algorithm and entropy regularization in heterogeneous multiagent system, the suboptimal problem of spectrum allocation in heterogeneous scenarios is solved, efficient spectrum sharing and network coordination are achieved, and the performance and robustness of spectrum access are improved.
Patent Information
- Application Number
- CN202510781710.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing DRL-based spectrum access research has failed to effectively solve the problem of heterogeneous multiagent collaborative optimization, resulting in serious degradation of spectrum allocation strategies in complex heterogeneous scenarios and the inability to achieve efficient spectrum utilization.
The heterogeneous multiagent entropy regularization resource allocation method for dynamic spectrum sharing is adopted. Through the integration of the HASAC algorithm with the DSA scenario, combining entropy regularization and MaxEnt targets, the differentiation capabilities and transmission requirements of heterogeneous terminals are optimized, and network robustness and spectrum sharing efficiency are improved.
The spectrum allocation strategy optimization in complex heterogeneous scenarios is realized, which improves the long-term efficiency of dynamic spectrum sharing and network robustness, reduces inter-user interference, and improves the throughput and accuracy of spectrum access.
Smart Images

Figure CN120343728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless network communication, and in particular to a heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing. Background Art
[0002] To address the problem of scarce spectrum resources in communication development and achieve a higher spectrum utilization rate, Dynamic Spectrum Access (DSA) technology has become a research focus in recent years due to its environmental adaptability. Compared with traditional machine learning methods, deep reinforcement learning has shown significant advantages in dealing with channel state uncertainty due to its dynamic environment modeling and real-time autonomous decision-making capabilities. However, existing DRL-based spectrum access research is mostly limited to the assumption of homogeneous networks and fails to effectively solve the core challenge of heterogeneous multi-agent collaborative optimization: on the one hand, existing methods assume that agents have homogeneous sensing and computing capabilities, ignoring the heterogeneity of network layer protocols and terminal layer capabilities in actual systems; on the other hand, traditional DRL frameworks are prone to trap agents in suboptimal Nash equilibria, leading to a serious degradation of spectrum allocation strategies in complex heterogeneous scenarios. Summary of the Invention
[0003] The object of the present invention is to provide a heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing. By innovatively integrating the HASAC algorithm with the DSA scenario, a collaborative decision-making engine for heterogeneous DRL nodes is developed to support the dynamic spectrum on-demand allocation of differentiated terminals in multi-protocol coexistence scenarios; by enhancing the exploration ability of secondary users in complex environments through entropy regularization and using the joint strategy optimization driven by the MaxEnt objective, the differentiated capabilities and transmission requirements of heterogeneous terminals are effectively coordinated, thereby improving the long-term efficiency and network robustness of dynamic spectrum sharing.
[0004] To achieve the above object, the present invention provides a heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing, including the following steps: S1. Load a basic trainer to manage the subsequent reinforcement learning process; including environment interaction, data storage, training loop, evaluation, and logging processes; S2. Construct a simulation environment, establish a mathematical model according to the dynamic spectrum access environment of cognitive radio networks, initialize parameter parsing (including command-line parameters, algorithm parameters, and environment parameters), construct a training environment and an evaluation environment for supporting heterogeneous dynamic spectrum access (DSA), and create agents (policy network, evaluation network, and experience replay buffer, etc.), and perform entropy regularization configuration. Prepare for the subsequent reinforcement learning training; The parameter parsing uniformly manages all configurable items. It constructs a multi-agent reinforcement learning environment that supports heterogeneous dynamic spectrum access (DSA). Its core role is to provide a simulation platform for agents to interact with the wireless communication environment, and it is responsible for simulating underlying physical rules such as the wireless environment, changes in wireless channel states, interference calculations, protocol conflicts, and the agent behavior logic framework.
[0005] S3. Generate initial experience data with a random policy and fill the experience replay buffer to provide basic data for subsequent formal training of the agent to avoid conflicts, and return the final state after reaching the preset warm-up steps. After the warm-up ends, the formal training will sample data from the experience replay buffer to update the network; S4. Start the formal training process. The agent interacts with the environment and stores experience data in the buffer. Periodically sample experience data from the buffer, optimize the policy network through policy gradients, optimize the evaluation network with TD errors, and dynamically adjust the entropy coefficient , thereby enhancing the decision-making ability of the agent. At the same time, periodically run the test environment, evaluate the performance of the current policy, and record data such as spectral efficiency and conflict statistics.
[0006] Preferably, the basic trainer in S1 is the scheduling center for multi-agent off-policy learning, which is used to integrate the three major components of the environment, agents, and buffer, standardize the training process (warm-up → collection → update → evaluation), and support flexible algorithm expansion through parameterized configuration and registration mechanisms. Its design goal is to provide a reusable, high-performance, and easily expandable training framework for dynamic spectrum access.
[0007] Preferably, the training environment in S2 is in an urban environment. Consider a multi-user cognitive wireless communication system, which includes M primary user (PU) devices with spectrum priority usage rights; P secondary user (SU) devices configured with traditional spectrum sensing modules; N cognitive SU devices configured with deep reinforcement learning decision-making modules; L orthogonally divided wireless parallel channel resources; the primary users (PUs) and the secondary users (SUs) using traditional MAC protocols (such as TDMA / ALOHA) form an independent communication layer. The PU devices have transmission priority within the authorized frequency band, and their communication behaviors are not affected by the activity states of the cognitive SUs, and they only perform data transmission based on their own service requirements. The non-cognitive SU devices communicate based on preset static access rules (such as fixed time slots or probabilistic contention), and their MAC layer protocols do not have environmental perception adaptability.
[0008] In the urban dynamic spectrum access scenario, the wireless communication network consists of M primary users (PUs) and P heterogeneous secondary users (SUs). The secondary users are further divided into three types of nodes: ALOHA protocol node: Access through a contention mechanism, perform spectrum sensing in a preset dedicated channel group, and randomly select a detected idle channel for transmission; TDMA protocol node: Adopt the time-division multiple access mechanism and strictly follow the pre-allocated time slot-channel combination for conflict-free periodic transmission; DRL-based intelligent node: Continuously maintain the state of waiting for transmission, and dynamically select the access channel through the deep reinforcement learning strategy in each time slot; Interference in a wireless communication system comes from two situations: Co-channel interference between primary and secondary users: Occurs when the PU and SU use the same channel simultaneously; Co-channel conflict between secondary users: Generated when multiple SUs compete for the same channel resource.
[0009] Preferably, construct a general path loss model, channel gain model, and signal-to-noise ratio model for the desired signal and interference signal in a wireless communication system in an urban environment, as follows: Assume Respectively represent The transmitter, The receiver, The position coordinates of the transmitter and The receiver, And Respectively represent the transmitter and the receiver, Represents the th cognitive user, Represents the th primary user; where The link distance of the desired signal is calculated by , and at the same time, the propagation distance of the interference signal is defined by And , where ; Considering the channel characteristics in an urban environment, the path loss is generated based on the widely used WINNER II model in the B1-urban micro-cell scenario. The general path loss models for the desired signal and interference signal are as follows: ; Among them, Represents the basic path loss in a specific environment; Is the path loss coefficient related to distance; Is the path loss coefficient related to frequency; Is the carrier frequency; Generally, there is a strong line-of-sight LoS path between the transmitter and the receiver. Therefore, the Rician channel model is adopted to calculate the channel gain: ; Among them, is the ratio of the received signal power of the line-of-sight path to that of the scattered path, determined by the path loss, indicating the phase of the received signal on the line-of-sight path, taking values from a uniform distribution between 0 and 1, representing a circularly symmetric complex Gaussian random variable; due to limited spectrum resources in the transmission environment, it is assumed that all channel bandwidths are the same, and the entire spectrum is evenly divided into channels; at the same time, the transmission power of each channel is the same, and the carrier frequency is the only fixed value; according to the above settings, the channel gain of the th secondary user on the th channel is defined as: ; ; To quantify the channel quality, the signal-to-noise ratio and the transmission rate are set as the evaluation criteria for the channel quality. For the frequency band selected by the th secondary user, this secondary user selects channels to meet its bandwidth requirements , and at the same time, there are primary users respectively occupying channels, and there may also be channels in conflict due to the common selection of several secondary users. The signal-to-noise ratio obtained by the th secondary user is defined as: ; Among them, is the gain of the sub-channel in the frequency band selected by the th secondary user, is the gain of the th primary user on the sub-channel , represents the gains generated by the secondary user and the remaining secondary users on each conflict channel, represents the noise spectral density of the cognitive user , is the transmission power of the cognitive user , is the transmission power of other cognitive users interfering with the current cognitive user.
[0010] Preferably, the composition and status of the channel resources in the wireless communication system are as follows: The wireless communication system includes parallel channel resources. The sub-channels include two states: occupied state (1) or idle state (0); and are divided into the following three non-overlapping frequency band intervals according to the access protocol type: Random access frequency band: Consisting of channels, and the ALOHA protocol is used to achieve asynchronous access. Each secondary user equipment (SU) in this frequency band competes for channel access at any time with a probability of ; Time division multiplexing frequency band: Contains channels, and the time division multiple access (TDMA) protocol is used to achieve synchronous access. Each TDMA device performs periodic data transmission in the first time slots of a period of ; Primary user dedicated frequency band: Configured with channels, providing dedicated communication services for the primary user PU. The channel states of each PU follow a two-state Markov chain and evolve dynamically. Its state space is defined as: State 1: Channel idle (can be opportunistically accessed by the SU), State 0: Channel occupied (SU access prohibited); The Markov chain state transition probability matrix can be parameterized as: ; Construct heterogeneous DRL nodes based on the partially observable Markov decision process, and use the entropy-regularized heterogeneous multi-agent actor-critic algorithm to autonomously optimize the access strategy in the above three spectrum environments, The heterogeneous DRL nodes deployed in this system are modeled based on the partially observable Markov decision process, and autonomously optimize the access strategy in the above three spectrum environments through the entropy-regularized heterogeneous multi-agent actor-critic algorithm to achieve primary user interference avoidance, multi-protocol access coordination, and dynamic load balancing.
[0011] Preferably, based on the DRL nodes, the access channel is dynamically selected through a deep reinforcement learning strategy in each time slot, as follows: In each time slot t, the cognitive user (SU) performs spectrum sensing on all channels to detect the channel state. At the same time, the agent obtains the optimal access strategy by dynamically updating the policy network. The state space of the agent is: ; Considering the inherent defects in actual spectrum sensing, the channel state observation values obtained by the secondary user (SU) may have errors. Let the sensing result of the nth cognitive user on the th channel be , and the sensing error probability of the th SU on the th channel be , the channel state transition probability is as follows: ; Due to the hardware limitations of the cognitive user's sensing device, it is unable to sense all the channels in the environment; assuming that the sensing ability of the th SU is , the minimum required bandwidth is , , then each SU can sense channel blocks, and the observed channel blocks also have two cases according to whether the number of idle sub-channels meets the minimum required bandwidth : meeting the transmission requirement (1) and not meeting the transmission requirement (0). At time slot for sub-channel observation results are represented as follows: ; Each cognitive user will select consecutive sub-channels from channels for spectrum sensing and determine whether the aggregated frequency band meets the minimum bandwidth requirement , indicating a total of sensing access actions. Therefore, the action space of the agent can be represented as: The intelligent spectrum access method implemented in this system: Each secondary user (SU) generates a spectrum access decision based on a deep reinforcement learning policy network, selects a target aggregated channel block according to the real-time environmental state, performs spectrum conflict detection and minimum bandwidth guarantee verification on the selected aggregated channel block, so as to decide to access or idle. If the physical location non-overlapping area of the channel block selected by the current SU and the channel blocks of the already accessed SUs meets , the access is successful. Otherwise, a conflict warning is triggered. After taking an action, the system obtains an immediate reward.
[0012] Preferably, in an urban scenario, for the mixed service requirements of ordinary, voice, and image, the wireless communication system provides dynamic channel aggregation capabilities; users in the wireless communication system can select idle channels for aggregation and access to meet the transmission requirements, and the system transmission rate is specifically represented as follows: ; Among them, represents the cognitive user bandwidth, represents the signal-to-noise ratio loss coefficient, is the signal-to-noise ratio, and the user's reward is set as follows: a. The When a cognitive user chooses not to access the channel, the reward ; b. For the th cognitive user, if it selects the corresponding channel block, the basic reward is . If the channel block selected by this cognitive user meets the usage requirements and there is no conflict, , and respectively represent the reward and punishment mechanisms when the channel selection involves the selected protocol frequency band range; if there is a conflict between the channel blocks selected by this cognitive user and other cognitive users, and there are C sub-channels overlapping with the channel blocks accessed by other cognitive users, the overall reward is ; c. When the channel block selected by the th cognitive user has been occupied by the primary user or secondary users of other protocols, resulting in inability to transmit, the reward is set to .
[0013] Preferably, in S2, the agent adopts a modular design, including a policy network, an evaluation network, and an experience replay buffer; The policy network is implemented based on the HASAC algorithm and has the capabilities of action sampling, log probability calculation, and probability distribution output, and can be used in different stages such as policy exploration, policy optimization, and policy evaluation.
[0014] The evaluation network is based on the double Q-network structure and predicts the value for the input of shared observation information and joint actions. The module introduces an automatic temperature parameter adjustment mechanism to adaptively regulate the policy entropy by minimizing the temperature loss function, thereby dynamically balancing the exploration and exploitation of the policy. In addition, the module supports value normalization processing and Huber loss calculation, enhancing the robustness and adaptability in dealing with partially observable and cooperative control tasks in multi-agent reinforcement learning.
[0015] The experience replay buffer constructs an off-policy experience replay structure based on the global state information provided by the environment and is used to store key experience data in the process of interaction between multi-agents and the environment, including . By introducing the n-step temporal difference return calculation mechanism, the module can more effectively capture the long-term return signal and enhance the stability of policy training and the sample utilization efficiency.
[0016] Preferably, the role of the warm-up stage in S3 is that in the initial stage of training, to avoid unstable training caused by insufficient experience of the policy network, the system fills the experience replay buffer by executing random actions (instead of actions generated by the policy). Specifically, before each round of interaction, the agent obtains the local observation , the global shared observation of the entire system , and the set of currently available actions In each warm-up step, the agent samples actions according to the random policy and interacts with the environment to obtain the local observation at the next moment. Shared observation Immediate reward Termination flag And auxiliary statistical information, which constitute an experience data. The experience data obtained in the previous interaction process is: , where while . The obtained experience data is sequentially stored in the experience replay buffer to establish an initial experience library for subsequent policy training. The global shared observation is used as the input of the evaluation network in the centralized training stage to improve the estimation accuracy of the value function and the global collaborative modeling ability, while the local observation is used for each agent's policy network to make independent decisions.
[0017] Preferably, the formal training stage described in S4 adopts a three-stage loop architecture of "interaction-training-evaluation", which specifically includes the following steps: SA1. Environment interaction and sample collection. Each agent is based on the local observation space , and generates actions through the policy network , and interacts with the environment model constructed in S2.
[0018] SA2. Experience storage. The return value is combined with the action , state , local observation , available actions to form a complete experience data and stored in the experience buffer to support the training process of subsequent reinforcement learning.
[0019] SA3. Policy optimization. After accumulating K pieces of experience data, the evaluation network and the policy network are updated. After randomly sampling a batch of data from the buffer, each agent calculates the action at the next state and the corresponding logarithmic probability through the policy network. These values are used for subsequent calculation of the target value. The calculation method of the target value is as follows: ; is the reward at the current time step, is the discount factor, is a hyperparameter that controls the weight of the policy entropy term.
[0020] When starting to train the evaluation network, minimize the gap between the predicted value and the target value through the loss function (mean squared error or Huber loss): ; ; is the target Q - value, B is the batch size, is the Huber loss function, which can balance the stability of the MSE loss and the robustness to outliers.
[0021] Finally, calculate the gradient of the loss and update the parameters of the evaluation network using an optimizer.
[0022] ; are the parameters of the evaluation network, is the learning rate, is the gradient of the loss function with respect to the network parameters.
[0023] The update of the policy network is based on the maximum - entropy objective, aiming to optimize the policy while maintaining sufficient exploration: ; Each agent samples an action according to the current observation , and calculates the log - probability of this action. Subsequently, each agent is trained in a random order. During the training process, the current evaluation network is used to estimate the Q - value of the state - action - action pair.
[0024] The policy loss function consists of two parts: one is the expected Q - value under the policy, and the other is the entropy regularization term. The entropy coefficient controls the balance between exploration and exploitation and can be automatically adjusted according to the target entropy to maintain sufficient randomness of the policy and prevent the policy from falling into a sub - optimal solution prematurely.
[0025] ; After calculating the loss, update the policy network parameters through backpropagation to optimize the policy. Finally, update the target network through a soft - update mechanism (i.e., set the target evaluation network parameters as the exponential moving average of the current network parameters), so that its parameters slowly approach the parameters of the target network, thereby ensuring the stability and convergence of the training process.
[0026] SA4. Periodic evaluation. After every M training epochs are completed, the system performs the following standardized evaluation operations: Pause the parameter updates of the policy network and the evaluation network. In a test environment isolated from the training environment (to avoid data contamination), execute T complete episodes. Record the corresponding metrics and calculate the statistics to present the current training effect.
[0027] Therefore, the present invention adopts the above heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing, and has the following beneficial effects: (1) Aiming at the heterogeneity of each user in the dynamic spectrum access environment, heterogeneous agent deep reinforcement learning is introduced on the basis of extending the traditional SAC algorithm to a multi-agent deep reinforcement learning algorithm. During the training phase, the evaluation network can estimate the policy value according to the global state, guide the update of the policy network, so as to maximize the reward while maximizing the entropy of the policy, and achieve avoiding convergence to sub-optimal, maximizing the optimal access performance while minimizing the interference between users; (2) The HASAC used has the advantages of fast convergence speed, less collision interference, and good cooperative access effect; it enables cognitive users to improve the access performance as much as possible under the conditions of heterogeneous state space, action space, and time slots, solves the conflict problem of spectrum access when cognitive users are heterogeneous, reduces the interference collision caused by heterogeneity, and thus improves the throughput and accuracy of dynamic spectrum access.
[0028] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0029] Figure 1 is the method framework diagram of the embodiment of the present invention; Figure 2 is the schematic diagram of the dynamic spectrum access environment scenario of the embodiment of the present invention. Detailed Embodiments
[0030] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0031] The entropy-regularized heterogeneous multi-agent actor-critic algorithm (Heterogeneous-Agent Soft Actor-Critic, HASAC) introduces the maximum entropy reinforcement learning idea of the Soft Actor-Critic (SAC) algorithm. By adding an entropy regularization term to the policy update, it encourages the policy to maintain sufficient randomness, thereby improving the exploration ability and sample efficiency. This method is based on a heterogeneous agent architecture, and designs personalized policies for agents with different observation perspectives, channel states, and aggregation capabilities, enabling them to perform efficient learning and collaborative decision-making in a complex dynamic spectrum access environment.
[0032] In a spectrum environment with heterogeneous cognitive users, due to differences in physical location, sensing ability, etc. among users, they have different information inputs and policy structures. HASAC enables each agent to learn a policy adapted to its own conditions through an independent policy network and a shared evaluation mechanism, while maintaining the ability to cooperate and optimize with the overall system, thereby improving the convergence and performance of the multi-agent system in a large-scale state-action space.
[0033] Please refer to Figure 1 , a heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing, comprising the following steps: S1. Integrate three major components: the environment (urban dynamic spectrum access scenario), agents (generated based on the HASAC algorithm), and buffer (off-policy buffer), and standardize the training process (warm-up → collection → update → evaluation). In this embodiment, 20 parallel environments are run simultaneously to collect data.
[0034] S2. Build a simulation environment, establish a mathematical model according to the dynamic spectrum access environment of cognitive radio networks, and simulate underlying physical rules such as wireless environment, wireless channel state changes, interference calculation, protocol conflicts, and agent behavior logic frameworks.
[0035] Considering the urban dynamic spectrum access scenario, randomly distribute M = 16 primary users and P = 10 heterogeneous secondary users. The secondary users are further divided into three types of nodes: ALOHA nodes, TDMA nodes, and cognitive nodes of DRL. Establish a multi-user multi-channel wireless network. As Figure 2 shown, assume that in a typical urban environment, multiple primary users (PUs) and secondary users (SUs) are randomly distributed in a multi-user multi-channel wireless communication system. There are a total of L = 32 orthogonal licensed channels in the system for primary users to perform data transmission on demand, and primary users do not need to consider the existence of secondary users during communication.
[0036] In the urban environment, secondary users communicate in a Rician channel and are interfered by 802.11 series devices (such as 802.11b). Compared with the indoor scenario, the propagation path in the urban environment is more complex, with a large number of obstructions, multipath propagation, and reflection paths, resulting in a more dynamic and unpredictable channel condition. In order to accommodate the transmission requirements of bandwidth applications such as voice, image, and information transmission, secondary users need to sense the spectrum and aggregate multiple idle and better-quality channels for access, so as to obtain the required bandwidth. The system involves the coexistence relationship between the expected signal link and the interference link during data transmission. When wireless signals propagate in the urban scenario, obvious path loss will occur, and at the same time, the channel gain will also be significantly affected by factors such as multipath fading and shadowing effects.
[0037] Considering the channel complexity in the urban environment, the path loss model we adopted is the WINNER II model: .
[0038] Meanwhile, there is a strong line-of-sight (LoS) path between the transmitter and the receiver, so the Rician channel model is used to calculate the channel gain: ; Among them, is The factor represents The ratio of the received signal power of the LoS path to that of the scattered path, Determined by the path loss, Represents The phase of the received signal on the LoS path, Represents a circularly symmetric complex Gaussian random variable. Figure 2 In the wireless communication system, multiple devices (such as and ) share limited spectrum resources. To improve the spectrum utilization efficiency, these devices can adopt spectrum aggregation technology, that is, by coordinating their respective transmitters and receivers, so that they can transmit data simultaneously at different times or frequencies. The figure shows the spectrum occupancy of each device in different time periods. Due to the limited spectrum resources in the transmission environment, it is assumed that all channel bandwidths are the same, and the entire spectrum is evenly divided into = 32 channels; the transmission power of each channel is the same, and the carrier frequency is the only fixed value; according to the above settings, the channel gain of the th secondary user on the th channel is defined as , and the signal-to-interference-plus-noise ratio and transmission rate are set as the evaluation criteria for channel quality.
[0039] Given that there are multiple users in the system, there may be channels occupied by primary users and channels accessed by other SUs in the frequency band selected by a certain SU. For the frequency band selected by the th secondary user, this cognitive user senses that channels meet its minimum bandwidth requirement , and at the same time, there are primary users respectively occupying channels, and there may also be channels in conflict due to the common selection of several secondary users. The observation state and action space of cognitive users are heterogeneous due to their respective requirements and the complexity of the actual wireless communication environment, that is, they cannot be exactly the same.
[0040] Generate a dynamic spectrum access optimization strategy according to the above calculation model, that is, use the introduced heterogeneous agent deep reinforcement learning algorithm to perform task allocation and access planning for cognitive users in the basic framework scenario, and maximize the system throughput and minimize the collision between users while ensuring that the PUs in the network are not interfered, as Figure 1 shown, the goal of the entropy-regularized heterogeneous multi-agent actor-critic algorithm is to maximize the entropy of the policy while maximizing the cumulative reward in a given environment.
[0041] Then create agents, including a policy network (Actor network), a value network (Critic network), target networks corresponding to the Actor network and the Critic network respectively, and an experience replay buffer. The policy network is implemented based on the HASAC algorithm, and the value network is based on the double Q-network structure to predict the value for the shared observation information and the joint action input. The experience replay buffer is used to store key experience data during the interaction between multi-agents and the environment, including , in this embodiment, the size is set to 100,000. Then perform entropy regularization configuration. In this implementation, the temperature parameter is 0.1, and the learning rate of the temperature parameter is set to 0.0003.
[0042] S3. Warm-up stage, the system fills the experience replay buffer by executing random actions (instead of policy-generated actions). In this embodiment, the warm-up step is set to 1,000. Specifically as follows: Step 1: Input prior knowledge into the simulation environment. Therefore, each agent can obtain the environmental state information through perception. Then, the agent randomly selects the aggregated frequency band planned to be used for sending data in the next step. Considering the perception difference (heterogeneity) with other users, the environmental state perceived by user n at time t can be defined as , , where is the current frequency band state, = 32 is the number of channels observable by the user. Considering the user's bandwidth requirements, the observed channels are aggregated, and the aggregation length is , so the length of the system action space is , , , . is the selected access frequency band. In this embodiment, the number of aggregated channels , the number of cognitive users .
[0043] Step 2: Each cognitive user in the selected aggregated frequency band Access and send data. At the same time, the system evaluates the interference situation between the PU and the SU, observes whether the current data transmission is successful, and whether there are frequency band selection conflicts among users. Based on these observation results, the transmission situation of the system is statistically analyzed: success, collision with cognitive users, collision with primary users, collision with other protocol nodes, etc. Thus, the reward that the current action should obtain is calculated.
[0044] Step 3. Store the key experience data obtained in the previous interaction process , . Store it in the experience replay buffer.
[0045] S4. The formal training stage adopts a three-stage cyclic architecture of "interaction - training - evaluation". During the entire training process, the total number of steps is 500,000 steps. It specifically includes the following ordered steps: Step 1. Each agent generates an action based on the local observation space through the policy network and interacts with the environment model constructed in S2.
[0046] Step 2. Store the complete experience data in the experience buffer to support the training process of subsequent reinforcement learning.
[0047] Step 3. Execute a training process every 100 steps of interaction with the environment. The training process mainly includes the parameter update of the policy network and the value network. After randomly sampling a batch of data from the buffer, each agent calculates the action and the corresponding log probability at the next state through the policy network. These values are used to calculate the target value . Among them, is the reward at the current time step, is the discount factor, is a hyperparameter that controls the weight of the policy entropy term. When starting to train the evaluation network (in this embodiment, the learning rate of the evaluation network is 0.0003), the evaluation network is trained by minimizing the gap between the predicted value and the target value through the loss function: ; Among them, is the target Q value, and B = 200 is the batch size. Finally, calculate the gradient of the loss and update the parameters of the evaluation network using the optimizer.
[0048] Step 4. Each agent updates its policy network in a random order according to the maximum entropy objective (in this embodiment, the learning rate of the policy network is 0.0003): .
[0049] During the training process, the current evaluation network is used to estimate the Q-value of the state -action pair. When calculating the policy loss function , the policy network parameters are updated through backpropagation to optimize the policy. Finally, the target network is updated through a soft update mechanism (in this embodiment, the target network soft update coefficient parameter is 0.03), so that its parameters slowly approach the parameters of the target network.
[0050] Step Five: Periodic evaluation. After every M = 200 training cycles are completed, the system performs the following standardized evaluation operations: Pause the parameter updates of the policy network and the evaluation network. In a test environment isolated from the training environment (to avoid data contamination), perform T = 20 complete rounds. Record the corresponding metrics and calculate the statistics to present the current training effect.
[0051] Therefore, the present invention adopts the above heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing. Aiming at the heterogeneity of each user in the dynamic spectrum access environment, heterogeneous agent deep reinforcement learning is introduced on the basis of extending the traditional SAC algorithm to a multi-agent deep reinforcement learning algorithm. In the training stage, the evaluation network can estimate the policy value according to the global state and guide the update of the policy network, so as to maximize the reward while maximizing the entropy of the policy, achieving avoiding convergence to sub-optimal, optimal access performance while minimizing the interference between users. The HASAC used has the advantages of fast convergence speed, less collision interference, and good cooperative access effect; it enables cognitive users to improve the access performance as much as possible under the conditions of heterogeneous state space, action space, and time slots, solves the conflict problem that occurs in spectrum access when cognitive users are heterogeneous, reduces the interference collision caused by heterogeneity, and thus improves the throughput and accuracy of dynamic spectrum access.
[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing, characterized in that It includes the following steps: S1. Load a basic trainer to manage the subsequent reinforcement learning process; It includes processes of environment interaction, data storage, training loop, evaluation, and logging; S2. Build a simulation environment, establish a mathematical model according to the cognitive radio network dynamic spectrum access environment, initialize parameter parsing, build a training environment and an evaluation environment that support heterogeneous dynamic spectrum access (DSA), create an agent, and perform entropy regularization configuration; The agent includes a policy network, a value network, and an experience replay buffer; S3. Generate initial experience data with a random policy and fill the experience replay buffer to provide basic data for the formal training of the agent to avoid collisions. After reaching the preset warm-up steps, return the final state. After the warm-up ends, the formal training will sample data from the experience replay buffer to update the network; S4. Start the formal training process. The agent interacts with the environment and stores the experience data in the buffer. Periodically sample the experience data from the experience replay buffer, optimize the policy network through policy gradients, optimize the evaluation network with the TD error, and dynamically adjust the entropy coefficient , and at the same time, periodically run the test environment, evaluate the performance of the current policy, and record the spectral efficiency and conflict statistics 2. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 1, characterized in that: The basic trainer in S1 is the scheduling center for multi-agent off-policy learning, which is used to integrate the three major components of the environment, the agent, and the buffer, standardize the training process, and support the flexible extension of the algorithm through parameterized configuration and registration mechanisms.
3. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 2, characterized in that: The simulation environment in S2 is in an urban environment, considering a multi-user cognitive radio communication system, including M primary user (PU) devices with spectrum priority usage rights; P secondary user (SU) devices configured with traditional spectrum sensing modules; N cognitive SU devices configured with deep reinforcement learning decision-making modules; L orthogonally divided wireless parallel channel resources; In the urban dynamic spectrum access scenario, the wireless communication network consists of M primary users and P heterogeneous secondary users. The secondary users are divided into three types of nodes: ALOHA protocol nodes: Access through a contention mechanism, perform spectrum sensing in a preset dedicated channel group, and randomly select a detected idle channel for transmission; TDMA protocol nodes: Adopt a time-division multiple access mechanism and follow a pre-allocated time slot-channel combination for conflict-free periodic transmission; DRL-based intelligent nodes: Continuously maintain a pending transmission state, and dynamically select an access channel through a deep reinforcement learning strategy in each time slot; The interference in the wireless communication system comes from two situations: Primary-secondary user co-channel interference: Occurs when a PU and an SU use the same channel simultaneously; Secondary user co-channel conflict: Generated when multiple SUs compete for the same channel resource; 4. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 3, characterized in that, Build a general path loss model, channel gain model, and signal-to-noise ratio model for the desired signal and interference signal of the wireless communication system in the urban environment, as follows: Hypothesis respectively represent the transmitter,[[]] the receiver,[[]] the transmitter and the position coordinates of the receiver,[[]] and respectively represent the transmitter and the receiver,[[]] represents the th cognitive user. The cognitive user specifically refers to the secondary user with dynamic spectrum sensing ability. Among them, the DRL node optimizes the access strategy through intelligent algorithms, and the remaining SUs adopt fixed protocols.[[]] represents the th primary user; among them the link distance of the desired signal is calculated by while the propagation distance of the interference signal is defined by and where ; Considering the channel characteristics in the urban environment, the general path loss model for the desired signal and interference signal is as follows: ; Among them, represents the basic path loss in a specific environment; is the distance-related path loss coefficient; is the frequency-related path loss coefficient; is the carrier frequency; There is a line-of-sight (LoS) path between the transmitter and the receiver, and the Rician channel model is used to calculate the channel gain: ; Among them, is a factor, representing the ratio of the received signal power of the path to that of the scattering path; Determined by path loss; represents the phase of the received signal on the path, taking values from a uniform distribution between 0 and 1; represents a circularly symmetric complex Gaussian random variable; due to limited spectrum resources in the transmission environment, it is assumed that all channel bandwidths are the same, and the entire spectrum is evenly divided into channels; at the same time, the transmission power of each channel is the same, and the carrier frequency is the only fixed value; according to the above settings, the th secondary user's channel gain on the th channel is defined as: ; To quantify the channel quality, the signal-to-noise ratio and transmission rate are set as the evaluation criteria for channel quality. For the frequency band selected by the th secondary user, this secondary user selects channels that meet its bandwidth requirements . At the same time, there are primary users each occupying channels, and there are also channels that conflict due to being jointly selected by several secondary users. The signal-to-noise ratio obtained by the th secondary user is defined as: ; Among them, is the gain of the th sub-channel in the frequency band selected by the secondary user . is the gain of the th primary user in the sub-channel . represents the gain generated by the secondary user and each conflict channel of the remaining secondary users. represents the noise spectral density of the cognitive user . is the transmit power of the cognitive user . is the transmit power of other cognitive users that interfere with the current cognitive user.
5. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 4, characterized in that The composition and state of the channel resources in the wireless communication system are as follows: A wireless communication system includes parallel channel resources. The sub-channels include two states: occupied state 1 or idle state 0; and are divided into the following three non-overlapping frequency band intervals according to the access protocol type: Random access band: Consisted of channels, and uses the ALOHA protocol to achieve asynchronous access; each secondary user device within this band competes for channel access at any time with a probability of ; Time-division multiplexing frequency band: includes channels, and a time-division multiple access protocol is used to achieve synchronous access; each TDMA device performs periodic data transmission within the first cycles with a period of time slots; Dedicated frequency band for primary users: Configuration channels are configured to provide dedicated communication services for primary users; the states of each PU channel evolve dynamically according to a two-state Markov chain, and its state space is defined as: State 1: The channel is idle; State 0: The channel is occupied, and SU access is prohibited; The Markov chain state transition probability matrix is parameterized as: ; Build a heterogeneous DRL node based on the partially observable Markov decision process, and use the entropy-regularized heterogeneous multi-agent actor-critic algorithm to autonomously optimize the access strategy in the above three spectrum environments.
6. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 5, wherein Based on the DRL node, each time slot dynamically selects an access channel through a deep reinforcement learning strategy, as follows: At each time slot t, the cognitive primary user SU performs spectrum sensing on all channels to detect the channel state. Meanwhile, the agent obtains the optimal access strategy by dynamically updating the policy network. The state space of the agent is as follows: ; Considering the inherent defects in actual spectrum sensing, the channel state observations obtained by secondary users may have errors. Let the sensing result of the th cognitive user on the n th channel be , and the sensing error probability of the th SU on the th channel be . The channel state transition probability is as follows: ; Due to the hardware limitations of the cognitive user perception device, it is unable to perceive all channels in the environment; assume that the th SU has a sensing ability of , a minimum required bandwidth of , , each SU senses channel blocks, and the observed channel blocks have two cases according to whether the number of idle sub-channels meets the minimum required bandwidth : meeting the transmission requirement 1 and not meeting the transmission requirement 0; In time slot for the observation results of sub-channels are expressed as follows: ; Each cognitive user selects consecutive sub-channels from channels for spectrum sensing and determines whether the aggregated frequency band meets the minimum bandwidth requirement , indicating a total of sensing and access actions; the action space of the agent is represented as: ; Each secondary user generates a spectrum access decision based on the deep reinforcement learning policy network, selects a target aggregated channel block according to the real-time environmental state, performs spectrum conflict detection and minimum bandwidth guarantee verification on the selected aggregated channel block, and thus decides to access or idle; If the physical location of the channel block selected by the current SU and the channel blocks of the connected SUs are in non-overlapping regions and satisfy , then the access is successful; otherwise, a collision warning is triggered. After taking an action, the wireless communication system obtains a reward.
7. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that Suppose in an urban scenario, for the mixed service requirements of ordinary, voice, and image, the wireless communication system provides dynamic channel aggregation capabilities; users in the wireless communication system select idle channels for aggregation and access to meet the transmission requirements. The transmission rate of the wireless communication system is specifically expressed as follows: ; Among them, represents the cognitive user bandwidth, represents the signal-to-noise ratio loss coefficient, is the signal-to-noise ratio, and the user's reward is set as follows: a. The th cognitive user chooses not to access the channel, and the reward is ; b. The th cognitive user, if it selects the corresponding channel block, the basic reward is ; if the channel block selected by this cognitive user meets the usage requirements and there is no conflict, , and respectively represent the reward and punishment mechanisms when channel selection involves the selected protocol frequency band range; if there is a conflict between the channel blocks selected by this cognitive user and other cognitive users, there are C sub-channels overlapping with the channel blocks accessed by other cognitive users, and the overall reward is ; c. When the channel block selected by the th cognitive user has been occupied by the primary user or secondary users of other protocols, resulting in the inability to perform transmission, the reward is set to .
8. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that The agent in S2 adopts a modular design, including a policy network, a value network, and an experience replay buffer; The policy network is implemented based on the HASAC algorithm and has the capabilities of action sampling, log probability calculation, and probability distribution output, and is used for policy exploration, policy optimization, and policy evaluation; The value network is based on the double Q network structure and predicts the value for the shared observation information and joint action input; The experience replay buffer constructs an off-policy experience replay structure based on the global state information provided by the environment, and is used to store key experience data during the interaction between multiple agents and the environment, including .
9. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that The content of the warm-up stage in S3 is as follows: The prior knowledge is input into the simulation environment. Each agent obtains the environmental state information through perception. Before the start of each round of interaction, the agent obtains the local observation , the global shared observation of the entire system , and the set of currently available actions ; In each warm-up step, the agent samples an action according to a random policy and interacts with the environment to obtain the local observation at the next moment , the shared observation , the immediate reward , the termination flag and auxiliary statistical information to form the experience data; Obtain experience data in the previous round of interaction: ; Among them, while ; the obtained empirical data is sequentially stored in the experience replay buffer to establish an initial experience library.
10. The heterogeneous multi-agent entropy regularization resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that, The formal training stage in S4 adopts an interaction-training-evaluation three-stage loop architecture, specifically including the following steps: SA1, Environment Interaction and Sample Collection, where each agent is based on the local observation space , generates actions through the policy network , and interacts with the environment model constructed in S2; SA2, Experience storage, the return value and the action at the current moment , status , local observation , available actions , to form complete experience data and store it in the experience buffer; SA3. Policy optimization: After accumulating K pieces of experience data, update the value network and the policy network; After randomly sampling a batch of data from the buffer, each agent calculates the action at the next state through the policy network and the corresponding log probability for calculating the target value; the calculation method of the target value is as follows: ; is the reward at the current time step, is the discount factor, is a hyperparameter that controls the weight of the policy entropy term; When starting to train the value network, minimize the gap between the predicted value and the target value through the loss function: ; ; is the target Q value, and B is the batch size, is the Huber loss function; Finally, calculate the gradient of the loss and use the optimizer to update the parameters of the value network; ; is a parameter of the evaluation network, is the learning rate, is the gradient of the loss function with respect to the network parameters; The update of the policy network is based on the maximum entropy objective: ; Each agent samples an action according to the current observation and calculates the log probability of this action. Subsequently, each agent is trained in a random order, and during the training process, the current evaluation network is used to estimate the Q-value of the state - action pair; The policy loss function consists of two parts: the expected Q value under the policy; the entropy regularization term; ; After calculating the loss, update the policy network parameters through backpropagation to optimize the policy; finally, update the target network through the soft update mechanism to make its parameters slowly approach the parameters of the target network; SA4. Periodic evaluation: After each M training cycles are completed, the system performs the following standardized evaluation operations: Suspend the parameter update of the policy network and the value network, and execute T complete rounds in a test environment isolated from the training environment; record the corresponding metrics and calculate the statistics to present the current training effect.
Citation Information
Patent Citations
Distributed energy system cluster collaborative optimization method based on multi-agent reinforcement learning
CN117350423A
Heterogeneous unmanned aerial vehicle task unloading and resource optimization method
CN118612855A
Unmanned aerial vehicle auxiliary communication anti-interference method based on multi-agent reinforcement learning
CN118921099A
Dynamic spectrum access method based on heterogeneous agent near-end strategy optimization
CN119071791A
Cited By
Underwater network structure self-evolution method based on perception data conflict driving
CN120711365A
Radioactive source searching method based on reinforcement learning
CN121028795A
Reinforced learning training method and device based on ray framework and medium
CN121031709A
Entropy driving step length self-adaption-based diffusion reinforcement learning channel access method
CN121463143A
Low-altitude traffic infrastructure network layout optimization method based on multi-agent simulation
CN121562427A