Multi-band spectrum resource allocation method based on A2-MADDPG
By adopting the A2-MADDPG method in multi-band resource allocation, combined with the multi-agent deep reinforcement learning and attention mechanism, the problems of inter-band interference and high computational complexity are solved, and efficient spectrum resource allocation and system performance improvement are achieved.
Patent Information
- Application Number
- CN202510342031.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-10
AI Technical Summary
When the existing multi-band resource allocation method faces the problems of inter-band interference and high computational complexity, it is difficult to effectively improve the spectrum efficiency and the real-time nature of the system.
A multi-band spectrum resource allocation method based on A2-MADDPG is adopted, and multi-band collaborative resource allocation is achieved by combining multi-agent deep reinforcement learning, adaptive multi-scale convolution, dynamic attention mechanism and evolutionary strategy.
It significantly improves spectrum efficiency, effectively reduces interference between frequency bands, enhances the algorithm's adaptability, reduces computing complexity and training data requirements, and improves the overall performance of the system.
Smart Images

Figure CN120129068A_ABST
Abstract
Description
Technical Field
[0001] The present invention designs a multi-band resource allocation method, belonging to the technical field of MAB attention mechanism, and is based on a multi-band spectrum resource allocation method adding an attention mechanism and an evolutionary strategy, and specifically relates to a multi-band spectrum resource allocation method based on A2-MADDPG. Background Art
[0002] With the rapid development of wireless communication technologies, especially the wide application of 5G and future 6G communication systems, the efficient utilization of spectrum resources has become a key challenge in network design. In a multi-band system, the effective allocation of spectrum resources plays a crucial role in the performance of the system. Existing multi-band resource allocation methods can be roughly divided into methods based on traditional optimization algorithms and methods based on machine learning. Traditional optimization algorithm methods, such as resource allocation strategies based on convex optimization, dynamic programming, or genetic algorithms, usually optimize using parameters such as channel conditions, interference constraints, and power allocation. As a technology to improve spectrum efficiency, NOMA has been widely studied and applied, especially in heterogeneous networks. However, this method still has significant interference problems, especially during the resource allocation process, and it is difficult to effectively manage the interference of multiple users. Even with the adoption of advanced algorithm optimization, such as the Successive Convex Approximation (SCA) algorithm, this interference problem still affects the overall efficiency and stability of the spectrum. In addition, traditional methods usually assume that the interference between frequency bands can be eliminated through an accurate interference management mechanism, but in practical applications, due to the uncertainty and variability of wireless channels, it is difficult to completely eliminate the interference between frequency bands, thus limiting the further improvement of spectrum efficiency.
[0003] In a Cooperative Non-Orthogonal Multiple Access (C-NOMA) system, different modulation modes are adopted for different users to reduce interference as much as possible and improve spectrum efficiency. However, this method relies on complex signal detection algorithms and relatively high computational overhead, and the improvement of spectrum efficiency is still limited, and there are still challenges that are difficult to completely solve in interference management. In addition, in a Cognitive Radio (CR) system, multi-band resource allocation has made certain progress in reducing spectrum utilization waste and improving efficiency. By jointly using spectrum sensing and power allocation strategies, the opportunity for data transmission can be significantly increased, but this method still faces challenges in interference management, especially under the conditions of limited spectrum and power budgets.
[0004] With the development of artificial intelligence, especially deep learning technology, the application of deep reinforcement learning (DRL) in the field of multi-band resource allocation has gradually become a research hotspot. Especially in multi-agent systems (MAS), deep reinforcement learning enables multiple agents to cooperate and achieve optimal allocation of spectrum resources in dynamic and complex environments. Multi-agent reinforcement learning (MARL) utilizes multiple agents to execute tasks in parallel and optimizes system performance through interaction, competition, and cooperation among them. In multi-band resource allocation, each agent usually represents different base stations, users, or frequency bands, and they jointly optimize the overall spectrum efficiency by sharing information or making collaborative decisions. Through MARL, the system can dynamically adapt to changes in the network environment and achieve real-time scheduling of spectrum resources and interference management. Although this method shows great potential, it still faces some challenges, such as how to achieve effective coordination under limited resources, how to handle conflicts and interference among agents, and how to balance exploration and exploitation to ensure convergence and stability. In particular, deep reinforcement learning methods often require a large amount of training data and computing resources, and it is difficult to effectively avoid interference in the face of severe interference between frequency bands in actual communication systems, thus affecting the improvement of spectrum efficiency.
[0005] However, although these methods overcome the deficiencies of traditional optimization algorithms to a certain extent, there are still some challenges and limitations in practical applications, which are specifically manifested in the following aspects:
[0006] (1) The computational complexity is still relatively high: Although resource allocation methods based on deep learning can reduce the computational overhead in some cases, due to the complexity of deep neural network models, the training process and online inference still require a large amount of computing resources. Especially in large-scale multi-band systems, the training of deep reinforcement learning algorithms requires a large number of samples and iterations, resulting in high time and computational costs for model training, which affects the real-time performance and deployment efficiency of the system. Even with the support of accelerated hardware, the increase in computational complexity remains a bottleneck restricting practical applications.
[0007] (2) It is difficult to effectively solve the problems of interference and loss between frequency bands: Although deep reinforcement learning methods can dynamically adjust spectrum resources through the learning and interaction of agents, due to the uncertainty of wireless channels and the complexity of interference between frequency bands, existing deep learning methods still perform inadequately in the face of frequency band overlap and interference management. Especially in the case of limited spectrum resources and severe interference, deep learning methods are difficult to completely eliminate interference or optimize spectrum utilization, thus restricting the further improvement of spectrum efficiency.
[0008] (3) High requirements for training data and computing resources: Deep learning methods rely on a large amount of training data to learn resource allocation strategies. Especially in dynamic and complex network environments, the collection and annotation of training data become an important challenge. In addition, the training process itself requires a large amount of computing resources, resulting in high costs for deploying and updating models. Although some studies have tried to solve this problem by simulating data, in practical applications, due to the randomness of wireless channels, there are still significant problems with the generalization ability of training data, affecting the performance and stability of the model.
[0009] Therefore, the multi-band resource allocation method of the present invention aims to further reduce losses and interference in the system, optimize the utilization efficiency of spectrum resources, and overcome the deficiencies of existing artificial intelligence technologies in interference management, computing efficiency, and training data requirements. Summary of the Invention
[0010] The main purpose of the present invention is to overcome the defects of the prior art and propose a multi-band resource allocation scheme based on multi-scale attention and evolutionary strategy. The system constructed by it can greatly reduce the losses caused by inter-band interference and finally improve the spectrum efficiency.
[0011] To achieve the above object, the present invention provides the following technical solutions:
[0012] A multi-band spectrum resource allocation method based on A2-MADDPG, the steps are as follows:
[0013] S1. Establish the basic framework of the communication system, define the environment and conditions for multi-band resource allocation; model according to the characteristics of each band to ensure that different bands can work together;
[0014] S2. Use adaptive multi-scale convolution and dynamic attention mechanism for feature extraction to extract multi-scale feature information from the system state;
[0015] S3. Use evolutionary strategy to optimize the hyperparameters in the model to ensure that the network can achieve optimal resource allocation in different network environments;
[0016] S4. Establish a dynamic reward mechanism to obtain spectrum efficiency rewards and interference suppression rewards based on the current action, so as to obtain the final reward, and adjust the reward weight in real time according to the current environment and the state of the agent, so as to balance short-term and long-term goals.
[0017] Further, in step S1, through multi-band collaborative modeling: perform targeted modeling according to the bandwidth, channel state information, and signal attenuation conditions of each band.
[0018] Further, in step S1, the state space includes the state of each access point, its own state, and the environmental state; APA1 The state S at time t t,A1 It is expressed as:
[0019]
[0020] where, represents APA 1 the channel gain of its associated user, is the large-scale fading coefficient, is APA 1 the current power allocation information for different users, is the in-band or inter-band interference power level from other access points, is the frequency band bandwidth information allocated to the user, is the association matrix describing the connection relationship between the base station and the access point.
[0021] Furthermore, in step S1, the action space of the agent is represented by a continuous multi-dimensional vector, where each dimension corresponds to a specific decision variable, including power allocation, frequency band selection, and user association; the Actor network takes the observation information of each agent about the environment and the implicit information of the strategies of other agents as inputs, and can further adjust its learning strategy according to the evaluation value feedback by the Critic network.
[0022] Furthermore, in step S2, the adaptive multi-scale convolution module is used to extract multi-level feature representations from channel data of different frequency bands; let the input channel state be s t , define the multi-scale convolution kernel size as k = {k 1 ,..., k m}; each convolution kernel extracts features from the input state s t at different scales to obtain multi-scale features F i , and then, these features are obtained by an adaptive weighting mechanism to get the final comprehensive feature F.
[0023]
[0024] where, α i is the weight coefficient, which is adjusted in real time through channel dynamic information, so that the algorithm can focus more on the characteristics of the current frequency band.
[0025] Furthermore, in step S2, after the feature extraction of each frequency band, the dynamic attention mechanism is used to further assign different weights to the features of different frequency bands; through the joint mechanism of channel attention and frequency band attention, this module dynamically and selectively focuses on important channel characteristics, while suppressing the channel information that causes significant interference to the current agent; the channel attention calculates the attention distribution W through the dot product of the input feature matrix F and the weight matrix W fc ofc ; Finally, the comprehensive attention feature Z is used as the input of the A network and the Critic network to generate an optimized resource allocation strategy and evaluate its performance.
[0026] Furthermore, in step S3, the agents form a population, and each agent has a unique neural network parameter configuration; subsequently, during the training process, the tournament selection mechanism is used to screen out the optimal individual in the population, and the adaptive mutation mechanism is used to optimize the key parameters, ensuring that the population can explore the optimal solution in each generation.
[0027] Furthermore, in step S4, the final reward function of the agent is shown in Equation (3).
[0028] R total = ω 1 R SE + ω 2 R total_inter (3)
[0029] where ω 1 represents the spectrum efficiency reward weight, R SE represents the spectrum efficiency reward, ω 2 represents the interference suppression reward weight, and R total_inter represents the interference suppression reward.
[0030] Advantages of the present invention:
[0031] Compared with the prior art, the present invention constructs a multi-band collaborative resource allocation scheme by combining multi-agent deep reinforcement learning, adaptive multi-scale convolution, attention mechanism, and evolutionary strategy; this scheme can reduce the loss between frequency bands to improve the spectrum utilization rate; in addition, the present invention integrates adaptive multi-scale convolution and dynamic attention mechanism, effectively improving the adaptive ability of the algorithm, enabling the system to maintain high performance under different network environments and demand changes, significantly reducing interference and improving the spectrum utilization rate; the multi-band spectrum resource allocation method based on A2-MADDPG of the present invention has the following technical characteristics and advantages:
[0032] (1) Significantly improve the spectrum efficiency: By constructing a multi-band spectrum resource allocation method based on A2-MADDPG, the present invention realizes the collaborative optimization among multi-agents. The agents can dynamically adjust decision variables such as power allocation, frequency band selection, and user association to adapt to the changing network environment, thereby significantly improving the utilization efficiency of spectrum resources.
[0033] (2) Effectively reduce interference between frequency bands: By introducing adaptive multi-scale convolution and dynamic attention mechanism, the present invention can extract multi-level and multi-scale feature representations from channel data of different frequency bands, dynamically and selectively focus on important channel characteristics, and at the same time suppress channel information that causes significant interference to the current agent. This refined feature extraction and interference management mechanism effectively reduces interference between frequency bands and reduces losses.
[0034] (3) Enhance the adaptive ability of the algorithm: The technical solution of the present invention integrates adaptive multi-scale convolution and dynamic attention mechanism, enabling the algorithm to adaptively adjust the strategies of feature extraction and decision-making according to changes in the network environment and requirements. This adaptive ability ensures that the system can maintain high performance in different scenarios, improving the robustness and stability of the system.
[0035] (4) Reduce computational complexity and training data requirements: Although deep reinforcement learning methods usually require a large amount of training data and computational resources, the present invention effectively reduces computational complexity and training data requirements by optimizing the neural network structure, optimizing hyperparameters using evolutionary strategies, and utilizing transfer learning and data augmentation techniques. This makes the solution of the present invention more feasible and efficient in practical applications.
[0036] (5) Improve the overall performance of the system: The multi-band resource allocation method of the present invention not only focuses on spectral efficiency and interference suppression, but also takes into account short-term and long-term goals by establishing a dynamic reward mechanism. This comprehensive optimization strategy enables the system to reserve space for future performance improvement while ensuring current performance, thus improving the overall performance of the system.
[0037] (6) Promote the development of wireless communication technology: The present invention provides a new solution to the resource allocation problem of wireless communication systems. By combining advanced technologies such as multi-agent deep reinforcement learning, adaptive multi-scale convolution, attention mechanism, and evolutionary strategy, the present invention not only solves the problems in the prior art, but also provides new ideas and methods for the design of future wireless communication systems.
[0038] The multi-band spectrum resource allocation method based on A2-MADDPG of the present invention shows significant beneficial effects in significantly improving spectral efficiency, effectively reducing interference between frequency bands, enhancing the adaptive ability of the algorithm, reducing computational complexity and training data requirements, and improving the overall performance of the system. These beneficial effects not only promote the development of wireless communication technology, but also provide strong support for the design of wireless communication systems in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0040] Figure 1 It is a schematic diagram of the MAB network structure of the present invention;
[0041] Figure 2 It is a flowchart of the present invention for optimizing hyperparameters using evolutionary strategies. Specific Embodiments
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The following combines the attached Figure 1-2 To further illustrate the multi-band spectrum resource allocation method based on A2-MADDPG.
[0043] Embodiment 1
[0044] This application proposes a multi-band resource allocation scheme based on multi-scale attention and evolutionary strategies. The specific technical solutions are as follows:
[0045] (1) Construct a multi-band collaborative system model in the downlink, perform targeted modeling according to the characteristics of different frequency bands, and obtain the final multi-band resource allocation strategy and optimization scheme. Effectively reduce inter-band interference and improve the spectral efficiency of the system. Through this modeling method, dynamic adjustment can be made according to the propagation characteristics, signal quality and interference conditions of each frequency band, enabling different frequency bands to work together and avoiding conflicts and interference between frequency bands.
[0046] (2) Construct a multi-agent deep reinforcement learning algorithm and introduce an attention mechanism into it to improve the learning efficiency of the model by precisely focusing on key features. Based on the constructed system model, experimental data is generated for training, and hyperparameters are optimized in combination with evolutionary strategies, thereby enhancing the adaptive ability and performance of the algorithm in complex environments. The effects that the present invention can achieve.
[0047] Based on the above-trained multi-agent deep reinforcement learning algorithm, the final optimized allocation scheme is obtained.
[0048] Embodiment 2
[0049] This solution provides a multi-band resource allocation scheme based on multi-scale attention and evolutionary strategies, and the specific implementation further elaborates on the present invention.
[0050] Step 1: Establish the basic framework of the communication system and define the environment and conditions for multi-band resource allocation. Model the characteristics of each frequency band to ensure that different frequency bands can work in coordination. This model provides a basis for subsequent agent design and resource allocation.
[0051] (1) Multi-band collaborative modeling: Conduct targeted modeling according to the characteristics of each frequency band, fully considering the bandwidth, channel state information (CSI), signal attenuation conditions, etc. of each frequency band. Since resource allocation is currently carried out in the context of multi-band collaboration, the present invention needs to fully consider the characteristics of each frequency band to conduct targeted modeling.
[0052] (2) State space: The state space includes the states of each access point, including its own state, the states of other access points, and the environmental state. APA 1 The state at time t can be expressed as:
[0053]
[0054] Among them, represents the channel gain between APA 1 and its associated user, is the large-scale fading coefficient, is APA 1 the current power allocation information for different users, is the in-band or inter-band interference power level from other access points, is the frequency band bandwidth information allocated to the user, is the association matrix describing the connection relationship between the base station and the access point.
[0055] (3) Action space: The action space of the agent is represented by a continuous multi-dimensional vector, where each dimension corresponds to a specific decision variable, including power allocation, frequency band selection, and user association. The Actor network takes the observation information of each agent about the environment and the implicit information of the strategies of other agents as inputs, and can further adjust its learning strategy according to the evaluation value feedback by the Critic network.
[0056] Step 2: This algorithm uses adaptive multi-scale convolution and dynamic attention mechanism for feature extraction. This module extracts multi-scale feature information from the system state. First, the adaptive multi-scale convolution module is used to extract multi-level feature representations from the channel data of different frequency bands. Let the input channel state be s t , and define the multi-scale convolution kernel size as k = {k 1 ,..., k m}. Each convolution kernel performs convolution on the input state s tFeature extraction is performed to obtain multi-scale features F i , and then these features are processed by an adaptive weighting mechanism to obtain the final comprehensive feature F.
[0057]
[0058] Among them, α i is the weight coefficient, which can be adjusted in real time through channel dynamic information, enabling the algorithm to focus more on the characteristics of the current frequency band.
[0059] After the feature extraction of each frequency band, the dynamic attention mechanism is used to further assign different weights to the features of different frequency bands.
[0060] Through the joint mechanism of channel attention and frequency band attention, this module dynamically and selectively focuses on important channel characteristics while suppressing channel information that causes significant interference to the current agent. Channel attention calculates the attention distribution W through the dot product of the input feature matrix F and the weight matrix c . Finally, the comprehensive attention feature Z is used as the input of the A network and the Critic network to generate an optimized resource allocation strategy and evaluate its performance.
[0061] The process implemented by this algorithm is shown in Algorithm 1.
[0062]
[0063]
[0064] Step 3: To further improve the algorithm performance, an evolutionary strategy is adopted to optimize the hyperparameters in the model, ensuring that the network can achieve optimal resource allocation in different network environments.
[0065] First, the agents form a population, and each agent has a unique neural network parameter configuration. Subsequently, during the training process, the best-performing individuals in the population are selected through the tournament selection mechanism, and an adaptive mutation mechanism is used to optimize the key parameters, ensuring that the population can explore the optimal solution in each generation.
[0066] Step 4: This algorithm designs a dynamic reward mechanism. The final reward is obtained by getting the spectrum efficiency reward and the interference suppression reward according to the current action, and the reward weight is adjusted in real time according to the current environment and the state of the agent, so as to balance short-term and long-term goals. The final reward function of the agent is shown in Equation (3).
[0067] R total = ω 1 R SE + ω 2 R total_inter (3)
[0068] where ω 1 represents the spectrum efficiency reward weight, and R SE represents the spectrum efficiency reward, and ω 2 represents the interference suppression reward weight, and R total_inter represents the interference suppression reward.
[0069] Through the above steps, the present invention constructs a multi-band resource allocation system with high fitness. This system can dynamically adjust the resource allocation strategy through a real-time feedback mechanism in a complex and changing wireless communication environment to reduce losses and interference and improve spectrum efficiency.
[0070] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A multi-band spectrum resource allocation method based on A2-MADDPG, characterized in that: Here are the steps: S1. Establish the basic framework of the communication system and define the environment and conditions for multi-band resource allocation; model the characteristics of each frequency band to ensure that different frequency bands can work together; S2, using adaptive multi-scale convolution and dynamic attention mechanism for feature extraction, extracting multi-scale feature information from the system state; S3, using evolutionary strategies to optimize the hyperparameters in the model to ensure that the network can achieve optimal resource allocation under different network environments; S4. Establish a dynamic reward mechanism to obtain spectrum efficiency rewards and interference suppression rewards based on the current action, thereby obtaining the final reward, and adjust the reward weight in real time according to the current environment and the state of the intelligent agent, so as to take into account both short-term and long-term goals.
2. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 1, characterized in that: In step S1, multi-band collaborative modeling is performed: targeted modeling is performed according to the bandwidth, channel state information and signal attenuation conditions of each frequency band.
3. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 1, characterized in that: In step S1, the state space includes the state of each access point, its own state, and the state of the environment; the state S of APA1 at time t t,A1 It is expressed as: in, represents the channel gain of APA1 and its associated users, is the large-scale fading coefficient, It is the power allocation information of APA1 to different users. is the intra-band or inter-band interference power level from other access points, Information on the frequency bandwidth allocated to users, The association matrix describes the connection relationship between the base station and the access point.
4. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 1, characterized in that: In step S1, the action space of the agent is represented by a continuous multi-dimensional vector, where each dimension corresponds to a specific decision variable, including power allocation, frequency band selection, and user association; the Actor network takes each agent's observation information of the environment and other agents' implicit strategy information as input, and can further adjust its learning strategy according to the evaluation value fed back by the Critic network.
5. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 1, characterized in that: In step S2, the adaptive multi-scale convolution module is used to extract multi-level feature representations from channel data of different frequency bands; assuming that the input channel state is s t , define the multi-scale convolution kernel size as k = {k1,…,k m }; Each convolution kernel processes the input state s at different scales t Perform feature extraction to obtain multi-scale features F i ,Then, these features are adaptively weighted to obtain the final comprehensive feature F. Among them, α i is the weight coefficient, which is adjusted in real time through channel dynamic information so that the algorithm can focus more on the characteristics of the current frequency band.
6. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 5, characterized in that: In step S2, after the features of each frequency band are extracted, the dynamic attention mechanism is used to further assign different weights to the features of different frequency bands; through the joint mechanism of channel attention and frequency band attention, the module dynamically and selectively focuses on important channel characteristics, while suppressing channel information that significantly interferes with the current agent; channel attention is achieved by combining the input feature matrix F with the weight matrix The dot product of the attention distribution W is calculated c ; Finally, the comprehensive attention feature Z is used as the input of the A network and the Critic network to generate an optimized resource allocation strategy and evaluate its performance.
7. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 1, characterized in that: In step S3, the agents form a population, each of which has a unique neural network parameter configuration; then during the training process, the best performing individuals in the population are selected through a bidding selection mechanism, and the key parameters are optimized using an adaptive mutation mechanism, ensuring that the population can explore the optimal solution in each generation.
8. The multi-band spectrum resource allocation method based on A2-MADDPG as claimed in claim 1, characterized in that: In step S4, the final reward function of the agent is shown in formula (3). R total =ω1R SE +ω2R total_inter (3) Where ω1 represents the spectrum efficiency reward weight, R SE represents the spectrum efficiency reward, ω2 represents the interference suppression reward weight, and R total_inter represents the interference rejection reward.
Citation Information
Cited By
Multi-agent-based spectrum sensing method and system, storage medium and terminal
CN122068987A
Multi-agent based spectrum sensing method, system, storage medium and terminal
CN122068987B