Multi-band network resource allocation method based on ML-MADDPG

By introducing a multi-agent deep reinforcement learning framework with meta-learning and layered reward mechanisms into multi-band networks, the problems of insufficient adaptability of dynamic environments and imbalance in multi-band networks are solved, and efficient spectrum resource utilization and throughput are achieved, suitable for large-scale and highly dynamic wireless communication scenarios.

CN120475414APending Publication Date: 2025-08-12DALIAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510665718.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has problems in the multi-band network with insufficient adaptability of dynamic environments, imbalance in the cooperation of multiple agents, weak training stability and generalization capabilities, resulting in low spectrum resource utilization and high packet loss rate, making it difficult to meet the network performance requirements in complex dynamic environments.

Method used

The meta-learning method is used to integrate the multi-agent deep reinforcement learning (MADDPG) framework, combined with the hierarchical reward mechanism, optimize multi-band resource allocation through meta-training and meta-adaptation stages, and build a hybrid integer nonlinear planning model to realize knowledge sharing and global coordination among multitasks, reduce packet loss rate and improve throughput.

Benefits of technology

Significantly reduce packet loss rate, improve network throughput, enhance dynamic environment adaptability and sample efficiency, optimize multi-agent collaboration and global coordination, improve training stability and generalization capabilities, support efficient resource scheduling under complex constraints, and reduce actual deployment costs and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475414A_ABST
    Figure CN120475414A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-band network resource allocation method based on ML-MADDPG, and belongs to the technical field of communication. According to the technical scheme, a basic framework of a communication system is established, and the environment and condition of multi-band resource allocation are defined; an integrated MBN network is adopted, and specific modeling is carried out by considering the characteristics of each frequency band; a meta learning method is fused into an MADDPG algorithm, and a meta learning MADDPG framework is constructed; according to task release of a multi-band network, each task corresponds to one network setting, specifically, the whole design is divided into a meta-training stage and a meta-adaptation stage; and target imbalance caused by direct weighted summation is avoided through a layered reward mechanism. The method has the beneficial effects that the packet loss rate is remarkably reduced, the network throughput is improved, the dynamic environment adaptability and sample efficiency are enhanced, the multi-agent cooperation and global coordination capability is optimized, the training stability and generalization capability are improved, efficient resource scheduling under complex constraints is supported, and the actual deployment cost and complexity are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technology, and relates to a multi-band resource allocation method, meta-learning, a multi-band network resource allocation method based on meta-learning, and specifically to a multi-band network resource allocation method based on ML-MADDPG. Background Art

[0002] In heterogeneous frequency band network architectures, efficient spectrum resource utilization faces multi-dimensional technical challenges. On the one hand, the sub-6 GHz, millimeter wave (mmWave), and terahertz (THz) frequency bands exhibit significant differences in propagation mechanisms, resulting in highly time-varying cross-band channel quality. On the other hand, the existing static allocation model based on Shannon capacity struggles to adapt to the multi-dimensional QoS requirements under dynamic service loads. Failure to properly address these challenges will not only lead to resource waste but may also hinder the promotion of ultra-high-speed services and new application scenarios. In this context, traditional single-band optimization strategies are no longer able to meet diverse service demands, and the contradiction between spectrum resource utilization and system performance is becoming increasingly prominent. To maximize downlink throughput and effectively mitigate high packet loss rates, a collaborative optimization solution is urgently needed that fully considers the characteristics of each frequency band and multiple constraints, thereby improving overall network performance in different link environments.

[0003] In multi-band networks, traditional optimization methods effectively improve spectrum utilization by rationally scheduling network resources. However, these methods typically rely on Shannon's capacity theory to construct static allocation models. Guided by optimization problems, they solve user association and power allocation problems through mathematical programming, convex optimization, game theory, or heuristic algorithms. However, these studies have encountered practical problems such as GPS interference, exposing the limitations of traditional static spectrum allocation methods for multi-band coordinated scheduling. The main drawback of these methods is their lack of adaptability to dynamic changes in the network environment, making it impossible to effectively balance the conflicts between network capacity, power budget, and channel conditions. Furthermore, these methods are often based on centralized optimization models, making it difficult to fully utilize the distributed information between base stations in the network. Consequently, in multi-user, multi-base station, and multi-connection scenarios, overall system performance is difficult to achieve.

[0004] With the advancement of artificial intelligence, particularly deep learning, a growing number of studies are exploring the application of these methods to multi-band resource allocation, aiming to achieve more efficient and adaptive scheduling strategies. In particular, deep reinforcement learning (DRL) leverages the collaborative efforts of multiple agents in multi-agent systems (MAS) to optimize spectrum resource allocation in dynamic and complex environments. In multi-band networks, a novel parameterized deep Q-network (P-DQN) was proposed to address the power allocation and user association issues in heterogeneous networks, significantly improving energy efficiency and dynamic scheduling performance. Furthermore, distributed deep reinforcement learning methods have been used to determine user association based on local channel information, optimizing user association and thereby improving overall throughput. However, these methods do not address global coordination or cross-network joint optimization. While these methods demonstrate significant potential, they still face challenges, including achieving effective coordination with limited resources, handling conflicts and interference between agents, and balancing exploration and exploitation to ensure convergence and stability. In particular, deep reinforcement learning methods often require a large amount of training data and computing resources, and still have shortcomings in the complex dynamic environment and multi-target coordination in actual communication systems. It is difficult to effectively reduce the packet loss rate, thus affecting the improvement of spectrum efficiency.

[0005] However, although these methods have overcome the shortcomings of traditional optimization algorithms to a certain extent, there are still some challenges and limitations in practical applications, which are specifically manifested in the following aspects:

[0006] (1) Low sample efficiency in dynamic environments: In multi-band networks, channel states change rapidly, user mobility is high, and the characteristics of each frequency band vary greatly. In particular, the signal attenuation in terahertz links is extremely significant due to the molecular absorption effect, resulting in a highly dynamic network environment. Traditional deep reinforcement learning algorithms require a large number of data samples for sufficient training, but in practical applications, the cost of obtaining high-quality and effective samples is very high. At the same time, the high attenuation characteristics of terahertz links require frequent adjustments to the power allocation strategy. Traditional methods are difficult to achieve rapid convergence under limited sample conditions, which in turn affects the overall performance of the system.

[0007] (2) Imbalance between multi-agent collaboration and competition: In a multi-base station scenario, each base station acts as an independent agent and makes distributed decisions, resulting in the widespread existence of local optimality. When competing for limited spectrum resources, adjacent base stations may interfere with each other's behavior, ultimately affecting the optimization of the overall network throughput. The contradiction between local and global interests cannot be effectively coordinated, causing the overall system operation to fall into a state similar to the prisoner's dilemma, limiting the optimization effect of network resource scheduling.

[0008] (3) Weak training stability and generalization ability: In multi-band networks, the frequent movement of users and switching between frequency bands make the network environment always in a non-stationary state, which has a significant impact on the training process of deep reinforcement learning algorithms. The training process is prone to oscillation or falling into a local optimal state, resulting in the model's adaptation speed and strategy stability in the new environment being unable to meet actual needs, thereby limiting the application effect of the algorithm in large-scale, dynamic networks.

[0009] Therefore, the multi-band network resource allocation method of the present invention aims to further reduce the packet loss rate in the system, improve the overall network throughput, optimize the utilization efficiency of spectrum resources, and overcome the shortcomings of existing artificial intelligence technology in sample efficiency, intelligent agent collaboration, and generalization capability requirements. Summary of the Invention

[0010] The main purpose of this invention is to overcome the shortcomings of the existing technology and propose a multi-band network resource allocation scheme based on meta-learning. The system constructed by the invention can greatly reduce the high packet loss rate caused by frequency band characteristics, and ultimately improve network throughput.

[0011] To achieve the above object, the present invention provides the following technical solutions:

[0012] A multi-band network resource allocation method based on ML-MADDPG, the steps are as follows:

[0013] S1. Establish the basic framework of the communication system and define the environment and conditions for multi-band resource allocation. Adopt an integrated downlink multi-band network (MBN) and conduct targeted modeling considering the characteristics of each frequency band.

[0014] S2. Integrate the meta-learning method into the MADDPG algorithm to construct the meta-learning MADDPG framework; according to the task release p(T) of the multi-band network, each task T k Specifically speaking, for a network setting, the entire design is divided into two stages: meta-training and meta-adaptation;

[0015] S3. Use a hierarchical reward mechanism to avoid imbalanced goals caused by direct weighted summation.

[0016] Furthermore, in step S1, in the multi-band collaborative modeling step, the unique characteristics of each frequency band include bandwidth, channel state information, and signal attenuation.

[0017] Furthermore, in step S1, the multi-band resource allocation problem is modeled as a mixed integer nonlinear programming by comprehensively considering power allocation, user association constraints, terahertz link molecular absorption constraints, and frequency band capacity constraints. The throughput is expressed as

[0018]

[0019] Among them, C1 is the upper limit constraint of the transmission power of each base station in different frequency bands, C2 is the user association constraint, UE uses multi-connection technology to simultaneously associate with multiple BSs, but not more than K; L in C3 th is the maximum acceptable absorption loss of the terahertz link, P represents the transmission power, X represents the relationship between the user and the base station, N represents the number of users, sub-6G represents the wireless communication frequency band below 6 GHz, n represents the nth user, mmWave represents millimeter wave, THz represents terahertz, i represents the i-th base station, j represents the j-th frequency band, and x i,n represents the connection status between base station i and user n, t represents the time index, K represents the maximum number of base stations that a user can associate with at the same time, d i,n,THz is the propagation distance from base station i to user n in frequency band j, and k(f) represents the molecular absorption coefficient.

[0020] Furthermore, in step S2, the meta-training phase is divided into two processes: an inner loop and an outer loop. First, in the inner loop task adaptation, the base parameter θ is used to generate the action a. i and interact with the environment to collect its trajectory data Performing a forward pass will then yield the task-specific loss Further adjust the parameters by gradient descent

[0021] Then, in the outer loop, we use the fine-tuned parameter θ′ k In the validation set Calculate the meta-loss The quadratic gradient update of the basis parameters θ is performed along the back propagation of the meta-loss.

[0022] Furthermore, in step S2,

[0023] In the meta-adaptation stage, fine-tuning is performed according to the new scene, using a small amount of samples D new Adjust the strategy parameters; initialize the adaptation parameters according to the base parameters θ obtained from meta-training use In D new Perform several gradient descents on

[0024]

[0025] Only the Actor network parameters θ are updated, while the Critic network parameters φ remain fixed to avoid overfitting; the intensity ò0 is high during the initial exploration, but decays exponentially with the number of adaptation steps t ε(t) = ε0·e -ηt ; In this stage, task-related noise is superimposed to generate a new action a new ; represents the new parameter θ at time t+1, represents the new parameter θ at time t, α represents the gradient update parameter, Represents task T new The policy loss, T new Represents a new task T.

[0026] Furthermore, in step S3, the total reward is composed of the basic reward layer, the penalty layer and the balance layer. The final reward function of the agent is shown in formula (3):

[0027]

[0028] in, represents the base reward layer, represents the penalty layer, Indicates the balancing layer.

[0029] Beneficial effects of the present invention:

[0030] Compared with the existing technology, the multi-band network resource allocation method based on ML-MADDPG in the present invention has the following technical features and beneficial effects:

[0031] (1) Significantly reduce packet loss and improve network throughput: By building a multi-band collaborative resource allocation model and combining the optimization of terahertz link characteristics with multi-connection constraints, the shortcomings of the traditional static allocation model under dynamic business loads are addressed. While maximizing downlink throughput, this solution effectively suppresses the high packet loss problem caused by frequency band characteristics differences (such as high attenuation of terahertz links).

[0032] (2) Enhanced adaptability to dynamic environments and sample efficiency: By integrating the meta-learning mechanism into the multi-agent deep reinforcement learning (MADDPG) framework, knowledge sharing and rapid migration between multiple tasks are achieved. The meta-training phase uses an internal and external loop gradient update mechanism to enable the model to quickly adapt to new environments based on a small number of samples, solving the problem of low sample efficiency of traditional deep reinforcement learning in dynamic multi-band networks.

[0033] (3) Optimizing multi-agent collaboration and global coordination capabilities: Distributed multi-agent collaborative decision-making, combined with meta-learning to optimize global strategies, effectively alleviates the local optimality problem caused by resource competition between base stations. A layered reward mechanism (basic reward, penalty layer, and balancing layer) coordinates local and global interests, avoids goal imbalance, and improves overall network throughput.

[0034] (4) Improving training stability and generalization: We introduce an evolutionary strategy to optimize hyperparameters and reduce the risk of overfitting during the meta-adaptation phase by fixing the critic network parameters and dynamically adjusting the actor network parameters. Combining the exploration strategy of task-related noise and exponential decay enhances the algorithm’s stability and generalization in non-stationary network environments.

[0035] (5) Supporting efficient resource scheduling under complex constraints: The multi-band resource allocation problem is modeled as a mixed integer nonlinear programming (MINLP), which comprehensively considers multi-dimensional constraints such as power allocation, user association, and terahertz link absorption loss, thereby achieving refined collaborative scheduling of heterogeneous frequency band network resources and significantly improving spectrum utilization and energy efficiency.

[0036] (6) Reduced deployment costs and complexity: The rapid fine-tuning capabilities of the meta-learning framework reduce reliance on large amounts of training data and computing resources. The system can dynamically adjust strategies based on real-time feedback, making it suitable for large-scale, highly dynamic wireless communication scenarios and reducing operational complexity and costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below in combination with the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] in:

[0039] Figure 1 This is the MBN scenario architecture diagram of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Figure 1 The multi-band network resource allocation method based on ML-MADDPG is further explained.

[0041] Example 1

[0042] This application proposes a multi-band network resource allocation solution based on meta-learning. The specific technical solution adopted is as follows:

[0043] (1) A multi-band network system model for the downlink was constructed. The characteristics of each band were specifically modeled, and a joint optimization model was constructed. This transformed the multi-band resource allocation problem into a collaborative optimization task that considers the characteristics of the terahertz link and multi-connection constraints. This effectively reduced the packet loss rate and improved the system throughput. This modeling approach maximized the downlink throughput while significantly reducing the packet loss rate.

[0044] (2) Construct a multi-agent deep reinforcement learning algorithm and integrate the meta-learning mechanism into the reinforcement learning framework. Share knowledge between multiple tasks, so that the model can quickly adapt to new environments. Improve the learning efficiency of the model by precisely focusing on key features. Based on the constructed system model, generate experimental data for training, and optimize hyperparameters in combination with evolutionary strategies to improve the algorithm's adaptability and performance in complex environments. The effects that can be achieved by the present invention.

[0045] The final optimized allocation plan is obtained based on the multi-agent deep reinforcement learning algorithm after the above training.

[0046] The above technical solutions adopted by the present invention have the following advantages compared with the prior art:

[0047] Multi-agent deep reinforcement learning and meta-learning construct a multi-band collaborative resource allocation scheme; this scheme reduces packet loss and improves system throughput. Furthermore, the invention incorporates meta-learning to effectively enhance the algorithm's transferability and generalization capabilities in dynamic environments, enabling the system to maintain high performance in dynamic environments, significantly reducing packet loss and improving system throughput.

[0048] Example 2

[0049] The present invention provides a multi-band network resource allocation solution based on meta-learning, and the specific implementation methods are further explained to further illustrate the present invention.

[0050] Step 1: Establish the basic framework of the communication system and define the environment and conditions for multi-band resource allocation. Use an integrated MBN network and model each frequency band based on its characteristics. This model provides the foundation for subsequent agent design and resource allocation.

[0051] (1) Multi-band collaborative modeling: Based on the unique characteristics of each frequency band, this invention requires the establishment of a targeted model that fully considers key factors such as bandwidth, channel state information, and signal attenuation. Since current research focuses on multi-band collaborative resource allocation, customized modeling based on the characteristics of different frequency bands is particularly important.

[0052] (2) Joint problem optimization modeling: In order to maximize the downlink throughput, it is necessary to comprehensively consider heterogeneous decision variables such as power allocation, user association constraints, terahertz link molecular absorption constraints, and frequency band capacity constraints. The multi-band resource allocation problem is modeled as a mixed integer nonlinear programming (MINLP). The throughput can be expressed as

[0053]

[0054] Among them, C1 is the upper limit constraint of the transmission power of each base station in different frequency bands, C2 is the user association constraint, and the UE can associate with multiple BSs at the same time through multi-connection technology, but the number cannot exceed K. th is the maximum acceptable absorption loss of the terahertz link.

[0055] Step 2: This algorithm integrates the meta-learning method into MADDPG and constructs a meta-learning MADDPG framework. This invention publishes p(T) based on the task of the multi-band network and converts each task T k Corresponding to a network setting. Specifically, the entire design is divided into two stages: meta-training and meta-adaptation.

[0056] In the meta-training stage, the present invention divides it into two processes: inner loop and outer loop. First, in the inner loop task adaptation, the basis parameters θ are used to generate action a i and interact with the environment to collect its trajectory data Performing a forward pass will then yield the task-specific loss Further adjust the parameters by gradient descent

[0057] Then, in the outer loop, we use the fine-tuned parameter θ′ k In the validation set Calculate the meta-loss The quadratic gradient update of the basis parameters θ is performed along the back propagation of the meta-loss.

[0058] In the meta-adaptation phase, the model will be fine-tuned according to the new scene, and a small number of samples D will be used. new Adjust the strategy parameters. Initialize the adaptation parameters according to the base parameters θ obtained from meta-training use In D new Perform several gradient descents on

[0059]

[0060] Only the Actor network parameters θ are updated, while the Critic network parameters φ remain fixed to avoid overfitting. The intensity of ò0 is high during the initial exploration, but it decays exponentially with the number of adaptation steps t ε(t) = ε0·e -ηtIn this stage, task-related noise is superimposed to generate a new action a new .

[0061] The implementation process of this algorithm is shown in Algorithm 1.

[0062]

[0063]

[0064]

[0065] Step 3: To further improve the performance of the algorithm, a hierarchical reward mechanism is designed to avoid the imbalance of the target caused by direct weighted summation. The total reward consists of a basic reward layer, a penalty layer, and a balance layer. The final reward function of the agent is shown in Equation 3

[0066]

[0067] Based on the above process, the present invention successfully designed a highly adaptable multi-band resource allocation system. This system can dynamically adjust resource scheduling strategies in a constantly changing wireless communication environment using a real-time feedback mechanism, effectively reducing packet loss and improving data throughput.

[0068] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A multi-band network resource allocation method based on ML-MADDPG, characterized in that: Here are the steps: S1. Establish the basic framework of the communication system and define the environment and conditions for multi-band resource allocation. Adopt an integrated downlink multi-band network and conduct targeted modeling considering the characteristics of each frequency band. S2. Integrate the meta-learning method into the MADDPG algorithm to construct the meta-learning MADDPG framework; according to the task release p(T) of the multi-band network, each task T k Specifically speaking, for a network setting, the entire design is divided into two stages: meta-training and meta-adaptation; S3. Use a hierarchical reward mechanism to avoid imbalanced goals caused by direct weighted summation.

2. The multi-band network resource allocation method based on ML-MADDPG according to claim 1, characterized in that: In step S1, in the multi-band collaborative modeling step, the unique characteristics of each frequency band include bandwidth, channel state information, and signal attenuation.

3. The multi-band network resource allocation method based on ML-MADDPG according to claim 1, characterized in that: In step S1, the multi-band resource allocation problem is modeled as a mixed integer nonlinear programming by comprehensively considering power allocation, user association constraints, terahertz link molecular absorption constraints, and frequency band capacity constraints. The throughput is expressed as Among them, C1 is the upper limit constraint of the transmission power of each base station in different frequency bands, C2 is the user association constraint, UE uses multi-connection technology to simultaneously associate with multiple BSs, but not more than K; L in C3 th is the maximum acceptable absorption loss of the terahertz link, P represents the transmission power, X represents the relationship between the user and the base station, N represents the number of users, sub-6G represents the wireless communication frequency band below 6 GHz, n represents the nth user, mmWave represents millimeter wave, THz represents terahertz, i represents the i-th base station, j represents the j-th frequency band, and x i,n represents the connection status between base station i and user n, t represents the time index, K represents the maximum number of base stations that a user can associate with at the same time, d i,n,THz is the propagation distance from base station i to user n in frequency band j, and k(f) represents the molecular absorption coefficient.

4. The multi-band network resource allocation method based on ML-MADDPG according to claim 1, wherein: In step S2, the meta-training phase is divided into two processes: inner loop and outer loop. First, in the inner loop task adaptation, the basis parameters θ are used to generate the action a. i and interact with the environment to collect its trajectory data Performing a forward pass will then yield the task-specific loss Further adjust the parameters by gradient descent Then, in the outer loop, we use the fine-tuned parameter θ k ′ in the validation set Calculate the meta-loss The quadratic gradient update of the basis parameters θ is performed along the back propagation of the meta-loss.

5. The multi-band network resource allocation method based on ML-MADDPG according to claim 4, characterized in that: In step S2, In the meta-adaptation stage, fine-tuning is performed according to the new scene, using a small amount of samples D new Adjust the strategy parameters; initialize the adaptation parameters according to the base parameters θ obtained from meta-training use In D new Perform several gradient descents on Only the Actor network parameters θ are updated, while the Critic network parameters φ remain fixed to avoid overfitting; the intensity ò0 is high during the initial exploration, but decays exponentially with the number of adaptation steps t ε(t) = ε0·e -ηt ; In this stage, task-related noise is superimposed to generate a new action a new ; represents the new parameter θ at time t+1, represents the new parameter θ at time t, α represents the gradient update parameter, and L Tnew Represents task T new The policy loss, T new Represents a new task T.

6. The multi-band network resource allocation method based on ML-MADDPG according to claim 4, characterized in that: In step S3, the total reward consists of the basic reward layer, the penalty layer and the balance layer. The final reward function of the agent is shown in formula (3): in, represents the base reward layer, represents the penalty layer, Indicates the balancing layer.

Citation Information

Cited By

  • Multi-service QoS cross-band bandwidth allocation method and system based on deep reinforcement learning

    CN119052861A

  • Multi-service qos cross-bandwidth allocation method and system based on deep reinforcement learning

    CN119052861B

  • Online multi-task optimization method for communication network base station resource scheduling

    CN120812758A