Heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing

By introducing the integration of HASAC algorithm and DSA scenarios in the heterogeneous multiagent environment, a collaborative decision-making engine for heterogeneous DRL nodes is developed, and the suboptimal Nash equalization problem in heterogeneous multiagent collaborative optimization is solved, the long-term efficiency of spectrum sharing and network robustness is improved, the interference between users is reduced, and the throughput and accuracy of spectrum access is improved.

CN120343728BActive Publication Date: 2025-08-12ZHEJIANG SCI-TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510781710.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-12
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing DRL-based spectrum allocation method has suboptimal Nash equilibrium problem in heterogeneous multiagent collaborative optimization, which leads to severe degradation of spectrum allocation strategies in complex heterogeneous scenarios, and cannot effectively solve the core challenge of heterogeneous multiagent collaborative optimization.

Method used

The heterogeneous multiagent entropy regularization resource allocation method for dynamic spectrum sharing is adopted. By fusing the HASAC algorithm with the DSA scenario, a collaborative decision-making engine for heterogeneous DRL nodes is developed, and the joint strategy optimization driven by MaxEnt target is used to coordinate the differentiated capabilities and transmission needs of heterogeneous terminals, and the long-term efficiency of spectrum sharing and network robustness are improved.

Benefits of technology

It has achieved enhanced exploration capabilities of secondary users in complex environments, reduced interference between users, improved throughput and accuracy of dynamic spectrum sharing, solved the spectrum access conflict caused by heterogeneity, and improved spectrum utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343728B_ABST
    Figure CN120343728B_ABST
Patent Text Reader

Abstract

The present invention discloses a heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing, which belongs to the field of wireless network communication technology. The method includes loading a basic trainer and managing the subsequent reinforcement learning process; constructing a simulation environment, establishing a mathematical model, initializing parameter parsing, constructing a training environment and an evaluation environment, and creating an agent; using a random strategy to generate initial experience data and fill the experience replay buffer, returning to the final state after reaching a preset number of warm-up steps, and formal training will update the network by sampling data from the experience replay buffer; starting the formal training process, the agent interacts with the environment, and stores the experience data in the buffer; optimizing the policy network through policy gradient, and optimizing the evaluation network with TD error. The present invention adopts the above method to solve the performance bottleneck problem caused by the traditional DRL method in spectrum allocation due to the policy convergence to suboptimal, and achieves resource allocation efficiency close to the global optimal through random policy optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless network communications, and in particular to a heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing. Background Art

[0002] To address the scarcity of spectrum resources in the development of communications and achieve more efficient spectrum utilization, Dynamic Spectrum Access (DSA) technology has become a research focus in recent years due to its environmentally adaptive nature. Compared to traditional machine learning methods, deep reinforcement learning, with its dynamic environment modeling and real-time autonomous decision-making capabilities, has demonstrated significant advantages in coping with channel state uncertainty. However, existing DRL-based spectrum access research is mostly limited to the assumption of homogeneous networks and fails to effectively address the core challenge of heterogeneous multi-agent collaborative optimization: on the one hand, existing methods assume that agents have homogeneous perception and computing capabilities, ignoring the heterogeneity of network-layer protocols and terminal-layer capabilities in actual systems; on the other hand, traditional DRL frameworks easily cause agents to fall into suboptimal Nash equilibrium, resulting in severe degradation of spectrum allocation strategies in complex heterogeneous scenarios. Summary of the Invention

[0003] The purpose of the present invention is to provide a heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing. By innovatively integrating the HASAC algorithm with the DSA scenario, a collaborative decision engine for heterogeneous DRL nodes is developed to support on-demand dynamic spectrum allocation for differentiated terminals in multi-protocol coexistence scenarios. Entropy regularization is used to enhance the exploration capability of secondary users in complex environments. At the same time, the MaxEnt target-driven joint strategy optimization is used to effectively coordinate the differentiated capabilities and transmission requirements of heterogeneous terminals, thereby improving the long-term efficiency and network robustness of dynamic spectrum sharing.

[0004] To achieve the above objectives, the present invention provides a heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing, comprising the following steps:

[0005] S1, loads the basic trainer and manages the subsequent reinforcement learning process; including environment interaction, data storage, training loop, evaluation and logging processes;

[0006] S2. Build a simulation environment, establish a mathematical model based on the dynamic spectrum access environment of cognitive wireless networks, initialize parameter parsing (including command line parameters, algorithm parameters, and environment parameters), build a training and evaluation environment that supports heterogeneous dynamic spectrum access (DSA), and create an intelligent agent (such as a policy network, evaluation network, and experience replay buffer) and perform entropy regularization configuration. This prepares for subsequent reinforcement learning training.

[0007] Parameter parsing unifies all configurable items. A multi-agent reinforcement learning environment supporting heterogeneous dynamic spectrum access (DSA) is constructed. Its core function is to provide a simulation platform for agents to interact with the wireless communication environment. This platform is responsible for simulating the underlying physical rules of the wireless environment, wireless channel state changes, interference calculations, protocol conflicts, and the agent behavior logic framework.

[0008] S3. Generate initial experience data using a random strategy and fill the experience replay buffer, providing basic data for subsequent formal training to avoid conflicts. After reaching the preset number of warm-up steps, the final state is returned. After the warm-up is complete, formal training samples data from the experience replay buffer to update the network.

[0009] S4. Start the formal training process. The agent interacts with the environment and stores the experience data in the buffer. Regularly sample the experience data from the buffer, optimize the policy network through the policy gradient, optimize the evaluation network with the TD error, and dynamically adjust the entropy coefficient. , thereby improving the decision-making ability of the intelligent agent. At the same time, the test environment is run periodically to evaluate the performance of the current strategy and record data such as spectrum efficiency and conflict statistics.

[0010] The preferred base trainer in S1 serves as the scheduling hub for multi-agent offline policy learning. It integrates the three components of the environment, agents, and buffer, standardizes the training process (pre-training → collection → update → evaluation), and supports flexible algorithm expansion through parameterized configuration and registration mechanisms. Its design goal is to provide a reusable, high-performance, and easily scalable training framework for dynamic spectrum access.

[0011] Preferably, the training environment in S2 is an urban environment. Consider a multi-user cognitive wireless communication system consisting of M primary user (PU) devices with priority spectrum use rights; P secondary user (SU) devices equipped with traditional spectrum sensing modules; N cognitive SU devices equipped with deep reinforcement learning decision modules; L orthogonally divided wireless parallel channel resources; and independent communication layers consisting of primary users (PU) and secondary users (SU) using traditional MAC protocols (such as TDMA / ALOHA). PU devices enjoy transmission priority within the licensed frequency band, and their communication behavior is unaffected by the activity state of the cognitive SUs; data transmission is based solely on their own service needs. Non-cognitive SU devices communicate based on preset static access rules (such as fixed time slots or probabilistic contention), and their MAC layer protocols lack environmental awareness and adaptability.

[0012] In an urban dynamic spectrum access scenario, the wireless communication network consists of M primary users (PUs) and P heterogeneous secondary users (SUs). The secondary users are further divided into three types of nodes:

[0013] ALOHA protocol nodes: access through a contention mechanism, perform spectrum sensing in a preset dedicated channel group, and randomly select idle channels for transmission;

[0014] TDMA protocol node: uses the time division multiple access mechanism and strictly follows the pre-assigned time slot-channel combination for conflict-free periodic transmission;

[0015] DRL-based intelligent nodes: They continuously maintain a ready-to-transmit state and dynamically select access channels in each time slot using deep reinforcement learning strategies.

[0016] Interference in wireless communication systems comes from two situations:

[0017] Primary and secondary user co-channel interference: occurs when the PU and SU use the same channel at the same time;

[0018] Co-frequency conflict between secondary users: occurs when multiple SUs compete for the same channel resources.

[0019] Preferably, a universal path loss model, channel gain model, and signal-to-noise ratio model for the desired signal and interference signal of a wireless communication system in an urban environment are constructed, and the content is as follows:

[0020] Assumptions Respectively transmitter, Receiver, Transmitter and The receiver's location coordinates, and Represent the transmitter and receiver respectively, Indicates the Cognitive users, Indicates the primary users; The link distance of the desired signal is Calculate, and the interference signal propagation distance is calculated by and defined, where ;

[0021] Taking into account the channel characteristics in urban environments, the path loss is generated based on the widely used WINNER II model in the B1-Urban micro-cell scenario. The general path loss model for the desired signal and the interference signal is as follows:

[0022] ;

[0023] in, Indicates the basic path loss under specific conditions; is the distance-dependent path loss coefficient; is the frequency-dependent path loss coefficient; is the carrier frequency;

[0024] Typically, there is a strong line-of-sight LoS path between the transmitter and receiver, so the Rician channel model is used to calculate the channel gain:

[0025] ;

[0026] Among them, for Factor representation The ratio of the receiver signal power of the path to the scattering path, Determined by the path loss, express The phase of the received signal on the path takes values uniformly distributed between 0 and 1. represents a circularly symmetric complex Gaussian random variable; due to the limited spectrum resources in the transmission environment, it is assumed that all channels have the same bandwidth and the entire spectrum is evenly divided into channels; at the same time, the transmission power of each channel Same, the carrier frequency is a unique fixed value; according to the above settings, Secondary users in The channel gain on a channel is defined as: ;

[0027] In order to quantify the channel quality, the signal-to-noise ratio and transmission rate are set as the evaluation criteria of the channel quality. In the frequency band selected by the secondary user, the secondary user selected channels to meet its bandwidth requirements , there is The main users each occupied channels, there may also be The channel conflicts due to the joint selection of several secondary users. The signal-to-noise ratio obtained by a secondary user is defined as:

[0028] ;

[0029] in, For the A secondary user selects a sub-channel in the frequency band The gain, For the Primary users in subchannels The gain, Indicates secondary users With the remaining The gain generated by each conflicting channel of each secondary user, Represents cognitive users The noise spectral density, For cognitive users The transmission power, Other cognitive users who interfere with the current cognitive user The transmission power.

[0030] Preferably, the composition and status of channel resources in the wireless communication system are as follows:

[0031] Wireless communication systems include Parallel channel resources, sub-channels have two states: occupied state (1) or idle state (0); divided into the following three non-overlapping frequency band intervals according to the access protocol type:

[0032] Random access frequency band: Channels are composed and ALOHA protocol is used to achieve asynchronous access. Each secondary user equipment (SU) in the frequency band has a probability of Perform channel contention access;

[0033] Time Division Multiplexing Band: Included channels, using the Time Division Multiple Access (TDMA) protocol to achieve synchronous access. Each TDMA device has a period of Before Periodic data transmission is performed within a time slot;

[0034] Primary user dedicated frequency band: configuration The channel states of each PU follow a two-state Markov chain dynamic evolution, and its state space is defined as: State 1: channel idle (can be accessed opportunistically by SU), State 0: channel occupied (no SU access);

[0035] The Markov chain state transition probability matrix can be parameterized as:

[0036] ;

[0037] We construct heterogeneous DRL nodes based on partially observable Markov decision processes and use entropy regularized heterogeneous multi-agent actor-critic algorithm to achieve autonomous optimization of access strategies in the above three spectrum environments.

[0038] The heterogeneous DRL nodes deployed in this system are modeled based on partially observable Markov decision processes. Through the entropy regularized heterogeneous multi-agent actor-critic algorithm, they autonomously optimize access strategies in the above three spectrum environments to achieve primary user interference avoidance, multi-protocol access coordination, and dynamic load balancing.

[0039] Preferably, based on the DRL node, each time slot dynamically selects the access channel through a deep reinforcement learning strategy, as follows:

[0040] At each time slot t, the cognitive user (SU) The channel performs spectrum sensing to detect the channel status, and the agent obtains the optimal access strategy by dynamically updating the strategy network. The state space of the agent is:

[0041] ;

[0042] Considering the inherent defects of actual spectrum sensing, the channel state observations obtained by the secondary user (SU) may have errors. The perception result of the nth cognitive user on the channel is: , No. On the channel The probability of perception error of a SU is , the channel state transition probability is:

[0043] ;

[0044] Due to the hardware limitations of cognitive user perception devices, they cannot perceive all channels; assuming the The perception capability of a SU is , the minimum required bandwidth is , , then each SU can perceive channel blocks, and The observed channel block is based on whether the number of idle sub-channels meets the minimum required bandwidth There are also two cases: transmission requirements are met (1) and transmission requirements are not met (0). In the time slot right Observation results of sub-channels It is expressed as follows:

[0045] ;

[0046] Each cognitive user will start from Select consecutive Spectrum sensing is performed on each sub-channel, and it is determined whether the aggregated frequency band meets the minimum bandwidth requirement. , indicating a total Therefore, the action space of the agent can be expressed as:

[0047] ;

[0048] The intelligent spectrum access method implemented by this system is as follows: each secondary user (SU) generates spectrum access decisions based on a deep reinforcement learning strategy network, selects a target aggregate channel block according to the real-time environment status, performs spectrum conflict detection and minimum bandwidth guarantee verification on the selected aggregate channel block, and then decides whether to access or idle. If the physical location non-overlapping area of the channel block selected by the current SU and the channel block of the connected SU meets , the access is successful. Otherwise, a conflict warning is triggered. After taking an action, the system receives an immediate reward.

[0049] Preferably, in an urban scenario, the wireless communication system provides dynamic channel aggregation capabilities for general, voice, and image mixed service needs. In the wireless communication system, users can select idle channels for aggregation and access to meet transmission needs. The system transmission rate is specifically expressed as follows:

[0050] ;

[0051] in, represents the cognitive user bandwidth, represents the signal-to-noise ratio loss coefficient, is the signal-to-noise ratio, and the user’s reward is set as follows:

[0052] a. No. A cognitive user chooses not to access the channel, and the reward ;

[0053] b. If a cognitive user chooses the corresponding channel block, the basic reward is If the channel block selected by the cognitive user meets the usage requirements and there is no conflict, , and They represent the reward and punishment mechanism when channel selection involves the selected protocol frequency band interval; if the channel block selected by the cognitive user and other cognitive users conflicts, there are C sub-channels with other The cognitive users access channel blocks overlap, and the total reward is ;

[0054] c. When the channel block selected by a cognitive user has been occupied by the primary user or a secondary user of another protocol, resulting in the inability to transmit, the reward is set to .

[0055] Preferably, the agent in S2 adopts a modular design, including a policy network, an evaluation network, and an experience replay buffer;

[0056] The policy network is implemented based on the HASAC algorithm and has the capabilities of action sampling, logarithmic probability calculation, and probability distribution output. It can be used in different stages such as policy exploration, policy optimization, and policy evaluation.

[0057] The evaluation network, based on a dual-Q network structure, predicts the value of shared observation information and joint action inputs. The module incorporates an automatic temperature parameter adjustment mechanism to adaptively control policy entropy by minimizing a temperature loss function, thereby dynamically balancing the policy's exploratory and exploitative nature. Furthermore, the module supports value normalization and Huber loss calculation, enhancing robustness and adaptability in multi-agent reinforcement learning with partial observability and collaborative control tasks.

[0058] The experience replay buffer builds a different strategy experience replay structure based on the global state information provided by the environment, which is used to store key experience data in the process of interaction between multiple agents and the environment, including By introducing an n-step temporal difference return calculation mechanism, the module can more effectively capture long-term return signals, enhancing the stability of strategy training and sample utilization efficiency.

[0059] Preferably, the warm-up phase in S3 is to fill the experience replay buffer by executing random actions (rather than policy-generated actions) at the beginning of training to avoid instability of the policy network due to insufficient experience. Specifically, before each round of interaction begins, the agent obtains local observations. , global shared observation of the entire system , and the set of currently available actions In each warm-up step, the agent samples actions according to the random strategy and interacts with the environment to obtain the local observations of the next moment. , shared observation , instant rewards , termination sign And auxiliary statistical information, constitute an experience data. Experience data obtained during the previous round of interaction: ,in, and The obtained experience data is stored in the experience replay buffer in sequence to establish the initial experience library for subsequent strategy training. Global shared observation In the centralized training phase, it is used as the input of the evaluation network to improve the estimation accuracy of the value function and the global collaborative modeling ability, while the local observation The policy network of each agent makes independent decisions.

[0060] Preferably, the formal training phase described in S4 adopts a three-stage loop architecture of "interaction-training-evaluation", specifically comprising the following steps:

[0061] SA1, Environmental Interaction and Sample Collection. Each agent is based on the local observation space , through the policy network Generate Action , interacting with the environment model constructed in S2.

[0062] SA2, experience storage. Return value Action at the current moment ,state , local observation , Available Actions , forming a complete empirical data Store in the experience buffer to support the subsequent reinforcement learning training process.

[0063] SA3, Strategy Optimization. After accumulating K pieces of experience data, the evaluation network and the strategy network are updated. After randomly sampling a batch of data from the cache, each agent calculates the action in the next state through the strategy network. and the corresponding log-probability. These values are used to calculate the target value later. The target value is calculated as follows:

[0064] ;

[0065] is the reward at the current time step, is the discount factor, is a hyperparameter that controls the weight of the entropy term of the policy.

[0066] When starting to train the evaluation network, the difference between the predicted value and the target value is minimized through the loss function (mean square error or Huber loss):

[0067] ;

[0068] ;

[0069] is the target Q value, B is the batch size, is the Huber loss function, which balances the stability of MSE loss and the robustness to outliers.

[0070] Finally, the gradient of the loss is calculated and the optimizer is used to update the parameters of the evaluation network.

[0071] ;

[0072] are the parameters of the evaluation network, is the learning rate, is the gradient of the loss function with respect to the network parameters.

[0073] The policy network is updated based on the maximum entropy objective, aiming to optimize the policy while maintaining sufficient exploration:

[0074] ;

[0075] Each agent observes Sampling Action , and calculate the logarithmic probability of the action. Then, each agent is trained in a random order. During the training process, the current evaluation network is used to evaluate the state - action Estimate the Q value of .

[0076] The policy loss function consists of two parts: one is the expected Q value under the policy, and the other is the entropy regularization term. It controls the balance between exploration and exploitation and can automatically adjust according to the target entropy to maintain sufficient randomness of the strategy and prevent the strategy from falling into a suboptimal solution too early.

[0077] ;

[0078] After calculating the loss, backpropagation is used to update the policy network parameters, thereby optimizing the policy. Finally, a soft update mechanism is used (setting the target evaluation network parameters to an exponential sliding average of the current network parameters) to update the target network parameters, slowly approaching those of the target network, thus ensuring the stability and convergence of the training process.

[0079] SA4. Periodic Evaluation. After completing M training cycles, the system performs the following standardized evaluation operations: pausing parameter updates for the policy and evaluation networks and executing T full rounds in a test environment isolated from the training environment (to avoid data contamination). The corresponding metrics are recorded and statistics are calculated to present the current training results.

[0080] Therefore, the present invention adopts the above-mentioned heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing, which has the following beneficial effects:

[0081] (1) Aiming at the heterogeneity of users in a dynamic spectrum access environment, heterogeneous agent deep reinforcement learning is introduced based on the extension of the traditional SAC algorithm to the multi-agent deep reinforcement learning algorithm. During the training phase, the evaluation network can estimate the policy value according to the global state and guide the policy network to update, thereby maximizing the reward while maximizing the entropy of the policy, avoiding convergence to suboptimal and optimal access performance while minimizing interference between users;

[0082] (2) The HASAC used has the advantages of fast convergence speed, less collision interference, and good cooperative access effect; it enables cognitive users to improve access performance as much as possible when the state space, action space, and time slot are heterogeneous, solves the conflict problem of spectrum access when there is heterogeneity of cognitive users, reduces interference collisions caused by heterogeneity, and thus improves the throughput and accuracy of dynamic spectrum access.

[0083] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 A method framework diagram of an embodiment of the present invention;

[0085] Figure 2 A schematic diagram of a dynamic spectrum access environment scenario according to an embodiment of the present invention. DETAILED DESCRIPTION

[0086] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.

[0087] The entropy-regularized heterogeneous-agent soft actor-critic (HASAC) algorithm incorporates the maximum entropy reinforcement learning concept of the soft actor-critic (SAC) algorithm. By incorporating an entropy regularization term into policy updates, it encourages policies to maintain sufficient randomness, thereby improving exploration and sample efficiency. Based on a heterogeneous agent architecture, this approach designs personalized policies for agents with different observation perspectives, channel states, and aggregation capabilities, enabling efficient learning and collaborative decision-making in complex dynamic spectrum access environments.

[0088] In a spectrum environment with heterogeneous cognitive users, users possess different information inputs and policy structures due to differences in physical location, perception, and other factors. HASAC, through an independent policy network and a shared evaluation mechanism, enables each agent to learn a policy adapted to its own conditions while maintaining the ability to coordinate optimization with the overall system, thereby improving the convergence and performance of multi-agent systems in large-scale state-action spaces.

[0089] See also Figure 1 , a heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing, including the following steps:

[0090] S1 integrates the three components of the environment (urban dynamic spectrum access scenario), the agent (generated based on the HASAC algorithm), and the buffer (off-policy buffer), and standardizes the training process (preheat → collect → update → evaluate). In this example, 20 parallel environments are run simultaneously to collect data.

[0091] S2. Build a simulation environment and establish a mathematical model based on the dynamic spectrum access environment of cognitive wireless networks to simulate the underlying physical rules such as the wireless environment, wireless channel state changes, interference calculations, protocol conflicts, and the intelligent agent behavior logic framework.

[0092] Consider a scenario of dynamic spectrum access in an urban area, with M = 16 primary users and P = 10 heterogeneous secondary users randomly distributed. Secondary users are further divided into three types of nodes: ALOHA nodes, TDMA nodes, and DRL cognitive nodes. A multi-user multi-channel wireless network is established. Figure 2 As shown in the figure, in a typical urban environment, multiple primary users (PUs) and secondary users (SUs) are randomly distributed in a multi-user, multi-channel wireless communication system. The system has L = 32 orthogonal licensed channels for primary users to transmit data on demand, without having to consider the presence of secondary users during communication.

[0093] In urban environments, secondary users communicate in Rician channels and are subject to interference from 802.11 devices (such as 802.11b). Compared to indoor scenarios, propagation paths in urban environments are more complex, with numerous obstructions, multipath propagation, and reflection paths, resulting in more dynamic and unpredictable channel conditions. To meet the transmission requirements of bandwidth-demanding applications such as voice, image, and information transmission, secondary users need to sense the spectrum and aggregate multiple idle, high-quality channels for access to obtain the required bandwidth. This system involves the coexistence of the intended signal link and the interfering link during data transmission. Wireless signal propagation in urban scenarios experiences significant path loss, and channel gain is also significantly affected by factors such as multipath fading and shadowing.

[0094] Considering the channel complexity in urban environments, we use the WINNER II path loss model:

[0095] .

[0096] At the same time, there is a strong line-of-sight (LoS) path between the transmitter and the receiver, so the Rician channel model is used to calculate the channel gain:

[0097] ;

[0098] in, for Factor representation The ratio of the receiver signal power of the path to the scattering path, Determined by the path loss, express The phase of the received signal on the path, represents a circularly symmetric complex Gaussian random variable. Figure 2 In wireless communication systems, multiple devices (such as and ) share limited spectrum resources. To improve spectrum utilization efficiency, these devices can use spectrum aggregation technology, that is, by coordinating their respective transmitters and receivers, they can transmit data simultaneously at different times or frequencies. The figure shows the spectrum occupancy of each device in different time periods. Due to the limited spectrum resources in the transmission environment, it is assumed that all channels have the same bandwidth and the entire spectrum is evenly divided into =32 channels; transmission power per channel Same, carrier frequency is a unique fixed value; according to the above settings, Secondary users in The channel gain on a channel is defined as , and set the signal-to-interference-and-noise ratio and transmission rate as the evaluation criteria for channel quality.

[0099] Since there are multiple users in the system, the frequency band selected by a certain SU may contain channels occupied by the primary user and channels accessed by other SUs. In the frequency band selected by the secondary user, the cognitive user perceives channels meet their minimum bandwidth requirements , there is The main users each occupied channels, there may also be The channels conflict due to the joint selection of several secondary users. The observation states and action spaces of cognitive users are heterogeneous due to their respective needs and the complexity of the actual wireless communication environment, that is, they cannot be completely consistent.

[0100] Based on the above calculation model, a dynamic spectrum access optimization strategy is generated. That is, the task allocation and access planning of cognitive users in the basic framework scenario are performed by introducing a deep reinforcement learning algorithm of heterogeneous agents, so as to maximize the system throughput and minimize the collision between users while ensuring that the PUs in the network are not interfered with. Figure 1 The entropy-regularized heterogeneous multi-agent actor-critic algorithm shown aims to maximize the entropy of the policy while maximizing the cumulative reward in a given environment.

[0101] Next, we create an agent, which includes a policy network (Actor network), an evaluation network (Critic network), a target network corresponding to each of the Actor and Critic networks, and an experience replay buffer. The policy network is implemented based on the HASAC algorithm, and the evaluation network is based on a double-Q network structure to predict the value of shared observation information and joint action input. The experience replay buffer is used to store key experience data during the interaction between multiple agents and the environment, including , in this embodiment, the size is set to 100000. Then the entropy regularization configuration is performed. In this implementation, the temperature parameter is 0.1, and the temperature parameter learning rate is set to 0.0003.

[0102] S3, the warm-up phase, in which the system fills the experience replay buffer by executing random actions (rather than strategy-generated actions). In this example, the number of warm-up steps is set to 1000. The details are as follows:

[0103] Step 1: Input prior knowledge into the simulation environment so that each agent can obtain information about the environment state through perception. Then, the agent randomly selects the aggregate frequency band to be used for the next data transmission. Considering the perception difference (heterogeneity) with other users, the environment state perceived by user n at time t can be defined as , ,in is the current frequency band status, =32 is the number of channels that can be observed by the user. Considering the user bandwidth requirements, the observed channels are aggregated and the aggregation length is , therefore, the length of the system action space is , , , . is the selected access frequency band. In this embodiment, the number of aggregated channels , the number of cognitive users .

[0104] Step 2: Each cognitive user selects the aggregated frequency band Access is made and data is transmitted. Simultaneously, the system evaluates interference between the PU and SU, observes the success of the current data transmission, and observes whether there are frequency band selection conflicts between users. Based on these observations, the system calculates transmission statistics, including success, collisions with cognitive users, collisions with primary users, and collisions with other protocol nodes. This calculation calculates the reward for the current action.

[0105] Step 3: The key experience data obtained during the previous round of interaction , . Stored in the experience replay buffer.

[0106] S4, the formal training phase adopts a three-stage loop architecture of "interaction-training-evaluation". The total number of steps in the entire training process is 500,000. Specifically, it includes the following ordered steps:

[0107] Step 1: Each agent is based on the local observation space , through the policy network Generate Action , interacting with the environment model constructed in S2.

[0108] Step 2: Complete empirical data Store in the experience buffer to support the subsequent reinforcement learning training process.

[0109] Step 3: After every 100 steps of interaction with the environment, a training process is performed. The training process mainly includes updating the parameters of the policy network and the value network. After randomly sampling a batch of data from the buffer, each agent calculates the action in the next state through the policy network. and the corresponding logarithmic probability. These values are used to calculate the target value later .in, is the reward at the current time step, is the discount factor, is a hyperparameter that controls the weight of the entropy term of the policy. When training the evaluation network (in this example, the learning rate of the evaluation network is 0.0003), the evaluation network is trained by minimizing the difference between the predicted value and the target value through the loss function:

[0110] ;

[0111] in, is the target Q value, and B=200 is the batch size. Finally, the gradient of the loss is calculated and the optimizer is used to update the parameters of the evaluation network.

[0112] Step 4: Each agent updates its policy network in a random order according to the maximum entropy objective (in this example, the learning rate of the policy network is 0.0003):

[0113] .

[0114] During the training process, the current evaluation network is used to evaluate the state -action Estimate the Q value of . When calculating the strategy loss function , update the policy network parameters through back propagation to optimize the policy. Finally, update the target network through the soft update mechanism (in this embodiment, the target network soft update coefficient parameters is 0.03), so that its parameters slowly approach those of the target network.

[0115] Step 5: Periodic Evaluation. After completing M = 200 training cycles, the system performs the following standardized evaluation: Parameter updates for the policy and evaluation networks are suspended, and T = 20 full rounds are executed in a test environment isolated from the training environment (to avoid data contamination). The corresponding metrics are recorded and statistics are calculated to present the current training results.

[0116] Therefore, the present invention adopts the above-mentioned heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing. Aiming at the heterogeneity of each user in the dynamic spectrum access environment, the traditional SAC algorithm is extended to the multi-agent deep reinforcement learning algorithm and then introduced into the heterogeneous agent deep reinforcement learning. In the training phase, the evaluation network can estimate the policy value according to the global state and guide the policy network to update, thereby maximizing the reward while maximizing the entropy of the policy, avoiding convergence to suboptimal and optimal access performance while minimizing interference between users. The HASAC used has the advantages of fast convergence speed, less collision interference, and good collaborative access effect; it enables cognitive users to improve access performance as much as possible when the state space, action space, and time slot are heterogeneous, solves the conflict problem of spectrum access when there is heterogeneity among cognitive users, reduces interference collisions caused by heterogeneity, and thus improves the throughput and accuracy of dynamic spectrum access.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing, characterized by: The following steps are involved: S1, load the basic trainer and manage the subsequent reinforcement learning process; Including environment interaction, data storage, training loop, evaluation and logging process; S2. Build a simulation environment, establish a mathematical model based on the dynamic spectrum access environment of cognitive wireless networks, initialize parameter analysis, build a training environment and evaluation environment that supports heterogeneous dynamic spectrum access (DSA), and create an intelligent agent and perform entropy regularization configuration; the intelligent agent includes a policy network, an evaluation network, and an experience replay buffer; S3. Generate initial experience data using a random strategy and fill the experience replay buffer to provide basic data for formal training of the agent to avoid conflicts. After reaching the preset number of warm-up steps, the final state is returned. After the warm-up is complete, formal training will sample data from the experience replay buffer to update the network. S4. Start the formal training process. The agent interacts with the environment and stores the experience data in the buffer. Periodically sample the experience data from the experience replay buffer, optimize the policy network through policy gradient, optimize the evaluation network with TD error, and dynamically adjust the entropy coefficient. ,At the same time, the test environment is run periodically to evaluate the performance of the current strategy, and record spectrum efficiency and conflict statistics.

2. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 1, characterized in that: The basic trainer in S1 is the scheduling center for multi-agent offline policy learning. It is used to integrate the three major components of the environment, agents, and buffers, standardize the training process, and support flexible algorithm expansion through parameterized configuration and registration mechanisms.

3. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 2, characterized in that: The simulation environment in S2 is an urban environment. A multi-user cognitive wireless communication system is considered, including M primary user PU devices with spectrum priority; P secondary user SU devices equipped with traditional spectrum sensing modules; N cognitive SU devices equipped with deep reinforcement learning decision modules; and L orthogonally divided wireless parallel channel resources. In an urban dynamic spectrum access scenario, the wireless communication network consists of M primary users and P heterogeneous secondary users. The secondary users are divided into three types of nodes: ALOHA protocol node: Accesses through a contention mechanism, performs spectrum sensing in a preset dedicated channel group, and randomly selects idle channels for transmission; TDMA protocol node: uses the time division multiple access mechanism to perform conflict-free periodic transmission according to the pre-assigned time slot-channel combination; DRL-based intelligent nodes: They continuously maintain a ready-to-transmit state and dynamically select access channels in each time slot using deep reinforcement learning strategies. Interference in wireless communication systems comes from two situations: Primary and secondary user co-channel interference: occurs when the PU and SU use the same channel at the same time; Co-frequency conflict between secondary users: occurs when multiple SUs compete for the same channel resources.

4. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 3, characterized in that: Construct a universal path loss model, channel gain model, and signal-to-noise ratio model for the desired signal and interference signal of a wireless communication system in an urban environment. The content is as follows: Assumptions Respectively transmitter, Receiver, Transmitter and The receiver's location coordinates, and Represent the transmitter and receiver respectively, Indicates the Cognitive users are secondary users with dynamic spectrum sensing capabilities. DRL nodes optimize access strategies through intelligent algorithms, while other SUs use fixed protocols. Indicates the primary users; The link distance of the desired signal is Calculate, and the interference signal propagation distance is calculated by and defined, where ; Considering the channel characteristics in urban environments, the general path loss model for the desired signal and the interference signal is as follows: ; in, Indicates the basic path loss under specific conditions; is the distance-dependent path loss coefficient; is the frequency-dependent path loss coefficient; is the carrier frequency; There is a line-of-sight LoS path between the transmitter and the receiver. The Rician channel model is used to calculate the channel gain: ; in, for Factor, indicating The ratio of the receiver signal power of the path to the scattering path; Determined by path loss; express The phase of the received signal on the path takes values uniformly distributed between 0 and 1; represents a circularly symmetric complex Gaussian random variable; due to the limited spectrum resources in the transmission environment, it is assumed that all channels have the same bandwidth and the entire spectrum is evenly divided into channels; at the same time, the transmission power of each channel Same, the carrier frequency is a unique fixed value; according to the above settings, Secondary users in The channel gain on a channel is defined as: ; In order to quantify the channel quality, the signal-to-noise ratio and transmission rate are set as the evaluation criteria of the channel quality. In the frequency band selected by the secondary user, the secondary user selected channels to meet its bandwidth requirements , there is The main users each occupied channels, and also The channel conflicts due to the joint selection of several secondary users. The signal-to-noise ratio obtained by a secondary user is defined as: ; in, For the A secondary user selects a sub-channel in the frequency band The gain, For the Primary users in subchannels The gain, Indicates secondary users With the remaining The gain generated by each conflicting channel of each secondary user, Represents cognitive users The noise spectral density, For cognitive users The transmission power, Other cognitive users who interfere with the current cognitive user The transmission power.

5. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 4, characterized in that: The composition and status of channel resources in wireless communication systems are as follows: Wireless communication systems include Parallel channel resources, sub-channels have two states: occupied state 1 or idle state 0; divided into the following three non-overlapping frequency band intervals according to the access protocol type: Random access frequency band: channels, and uses ALOHA protocol to achieve asynchronous access; each secondary user device in the frequency band can access the network with probability at any time. Perform channel contention access; Time Division Multiplexing Band: Included channels, using time division multiple access protocol to achieve synchronous access; each TDMA device has a period of Before Periodic data transmission is performed within a time slot; Primary user dedicated frequency band: configuration The channel provides dedicated communication services for the primary user. The state of each PU channel follows the dynamic evolution of a two-state Markov chain, and its state space is defined as: State 1: Channel is idle; State 0: Channel is occupied, SU access is prohibited; The Markov chain state transition probability matrix is parameterized as: ; A heterogeneous DRL node based on a partially observable Markov decision process is constructed, and an entropy-regularized heterogeneous multi-agent actor-critic algorithm is used to achieve autonomous optimization of access strategies in the above three spectrum environments.

6. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that: Based on the DRL node, each time slot dynamically selects the access channel through a deep reinforcement learning strategy, as follows: In each time slot t, the cognitive primary user SU The channel performs spectrum sensing to detect the channel status, and the agent obtains the optimal access strategy by dynamically updating the strategy network. The state space of the agent is: ; Considering the inherent defects of actual spectrum sensing, the channel state observations obtained by secondary users may have errors; On the channel n The perception result of cognitive users is , No. On the channel The probability of perception error of a SU is , the channel state transition probability is: ; Due to the hardware limitations of cognitive user perception devices, they cannot perceive all channels; assuming the The perception capability of a SU is , the minimum required bandwidth is , , each SU perceives channel blocks, and The observed channel block is based on whether the number of idle sub-channels meets the minimum required bandwidth There are two cases: transmission requirement 1 is met and transmission requirement 0 is not met; In the time slot right Observation results of sub-channels It is expressed as follows: ; Each cognitive user Select consecutive Spectrum sensing is performed on each sub-channel, and it is determined whether the aggregated frequency band meets the minimum bandwidth requirement. , indicating a total Perception access action; the action space of the agent is expressed as: ; Each secondary user generates spectrum access decisions based on a deep reinforcement learning policy network, selects a target aggregate channel block based on real-time environmental conditions, performs spectrum conflict detection and minimum bandwidth guarantee verification on the selected aggregate channel block, and then decides whether to access or idle. If the physical location non-overlapping area of the channel block selected by the current SU and the channel block of the connected SU satisfies , the access is successful; otherwise, a conflict warning is triggered; After taking an action, the wireless communication system receives a reward.

7. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that: In an urban scenario, the wireless communication system provides dynamic channel aggregation for mixed general, voice, and image services. Users in the wireless communication system select idle channels for aggregation and access to meet transmission requirements. The transmission rate of the wireless communication system is specifically expressed as follows: ; in, represents the cognitive user bandwidth, represents the signal-to-noise ratio loss coefficient, is the signal-to-noise ratio, and the user’s reward is set as follows: a. No. A cognitive user chooses not to access the channel, and the reward ; b. If a cognitive user chooses the corresponding channel block, the basic reward is If the channel block selected by the cognitive user meets the usage requirements and there is no conflict, , and They represent the reward and punishment mechanism when channel selection involves the selected protocol frequency band interval; if the channel block selected by the cognitive user and other cognitive users conflicts, there are C sub-channels with other The cognitive users access channel blocks overlap, and the total reward is ; c. When the channel block selected by a cognitive user has been occupied by the primary user or a secondary user of another protocol, resulting in the inability to transmit, the reward is set to .

8. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that: The agent in S2 adopts a modular design, including a policy network, an evaluation network, and an experience replay buffer; The policy network is implemented based on the HASAC algorithm and has the capabilities of action sampling, logarithmic probability calculation, and probability distribution output, which is used for policy exploration, policy optimization, and policy evaluation. The evaluation network is based on a dual Q network structure and performs value prediction on shared observation information and joint action inputs. The experience replay buffer builds a different strategy experience replay structure based on the global state information provided by the environment, which is used to store key experience data in the process of interaction between multiple agents and the environment, including .

9. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that: The contents of the warm-up phase in S3 are as follows: Prior knowledge is input into the simulation environment, and each agent obtains the state information of the environment through perception. Before each round of interaction begins, the agent obtains local observations. , global shared observation of the entire system , and the set of currently available actions ; In each warm-up step, the agent samples actions according to the random strategy and interacts with the environment to obtain local observations at the next moment. , shared observation , instant rewards , termination sign and auxiliary statistical information to constitute empirical data; Experience data obtained during the last round of interaction: ; in, and ; The obtained experience data is stored in the experience playback buffer in sequence to establish the initial experience library.

10. The heterogeneous multi-agent entropy regularized resource allocation method for dynamic spectrum sharing according to claim 5, characterized in that: The formal training phase in S4 adopts a three-stage cycle architecture of interaction-training-evaluation, which specifically includes the following steps: SA1, environmental interaction and sample collection, each agent is based on the local observation space , through the policy network Generate Action , interact with the environment model constructed in S2; SA2, experience storage, will return value Action at the current moment ,state , local observation , Available Actions , forming a complete empirical data Deposit into experience buffer; SA3, strategy optimization, after accumulating K pieces of experience data, execute the update of the evaluation network and the strategy network; After randomly sampling a batch of data from the buffer, each agent calculates the action in the next state through the policy network and the corresponding log-probability are used to calculate the target value; the target value is calculated as follows: ; is the reward at the current time step, is the discount factor, is a hyperparameter that controls the weight of the entropy term of the policy; When starting to train the evaluation network, the difference between the predicted value and the target value is minimized by the loss function: ; ; is the target Q value, B is the batch size, is the Huber loss function; Finally, the gradient of the loss is calculated and the parameters of the evaluation network are updated using the optimizer; ; are the parameters of the evaluation network, is the learning rate, is the gradient of the loss function with respect to the network parameters; The update of the policy network is based on the maximum entropy objective: ; Each agent observes Sampling Action , and calculate the logarithmic probability of the action. Then, each agent is trained in a random order, using the current evaluation network to evaluate the state during the training process. - action Estimate the Q value of ; The policy loss function consists of two parts: the expected Q value under the policy; the entropy regularization term; ; After calculating the loss, the policy network parameters are updated through backpropagation to optimize the policy. Finally, the target network is updated through a soft update mechanism so that its parameters slowly approach those of the target network. SA4, Periodic Evaluation: After each M training cycles, the system performs the following standardized evaluation operations: Pause parameter updates for the policy network and evaluation network, and execute T complete rounds in a test environment isolated from the training environment. Record the corresponding indicators and calculate statistics to present the current training results.

Citation Information

Patent Citations

  • Distributed energy system cluster collaborative optimization method based on multi-agent reinforcement learning

    CN117350423A

  • Dynamic spectrum access method based on heterogeneous agent near-end strategy optimization

    CN119071791A