Multi-service qos cross-bandwidth allocation method and system based on deep reinforcement learning

By building an effective capacity model and a multi-task hierarchical architecture, combined with the MADDPG and DDPG algorithms, the problem of insufficient consideration of multi-service QoS differences and frequency band characteristics in wireless communication systems is solved, achieving more efficient bandwidth resource allocation, and improving user experience and system throughput.

CN119052861BActive Publication Date: 2025-10-17DALIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411263649.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-10-17
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

Wireless communication systems fail to fully consider the QoS differences of various services and the characteristics of different frequency bands, resulting in poor user experience and low effective capacity.

Method used

A multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning is adopted. By constructing an effective capacity model and a multi-task hierarchical architecture, combined with the MADDPG and DDPG algorithms, a shared representation layer, policy head and value head are introduced in the high-level and low-level networks respectively to achieve reasonable allocation of cross-band bandwidth resources.

Benefits of technology

It improves user experience and effective capacity, can maintain a high average effective capacity under different numbers of users, reduce latency, improve SINR, balance global and local optimization needs, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052861B_ABST
    Figure CN119052861B_ABST
Patent Text Reader

Abstract

The application discloses a multi-service QoS cross-frequency band bandwidth allocation method and system based on deep reinforcement learning, and relates to the technical field of wireless communication; comprising: constructing an effective capacity model to clarify the overall optimization target of bandwidth resources; taking the effective capacity model as the basis, a multi-task hierarchical architecture based on deep reinforcement learning is constructed, which preliminarily allocates the overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through a high-level network; then the specific bandwidth resources inside each frequency band are finely allocated through a low-level network; a shared representation layer, a policy head and a value head are introduced in the high-level network and the low-level network respectively, wherein the shared representation layer captures the common features of different frequency bands, the policy head generates a specific allocation strategy, and the value head evaluates the allocation effect. Through the combination of global and local strategies, the application successfully balances the global optimization and local optimization requirements in resource allocation, and performs outstandingly in reducing delay, improving SINR and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless communication, and particularly relates to a multi-service QoS cross-frequency band bandwidth allocation method and system based on deep reinforcement learning. BACKGROUND

[0002] With the rapid growth of global mobile data traffic, modern communication networks are facing unprecedented challenges. Users' demand for high-speed data transmission, ultra-low latency and high reliability continues to increase, while traditional spectrum resources have been difficult to meet this demand. In this context, 5G communication systems are constantly exploring new spectrum resource boundaries and seeking more efficient bandwidth utilization strategies [1] to meet this challenge. Especially in complex and variable multi-service environments, ensuring the quality of service (QoS) of various services has become a core task. Sub-6GHz frequency bands are widely used in the industry due to their excellent coverage capabilities, but their limited bandwidth is difficult to support future demand for ultra-high data rates. In order to break this bottleneck, millimeter wave (mmWave) communication technology has emerged, with its extremely high available bandwidth opening up a new era for large-scale data transmission. However, millimeter wave signals are easily affected by physical obstacles, and their coverage range is relatively limited. Therefore, the clever combination of Sub-6GHz and millimeter wave technology to build a cross-frequency band cooperative networking architecture [2] has become an effective strategy to solve the current dilemma. Under this architecture, how to intelligently allocate bandwidth resources between different frequency bands to accurately meet users' quality of service (QoS) requirements [3] has become a key problem that the industry needs to overcome. The significant differences between millimeter wave and Sub-6GHz frequency bands in terms of propagation characteristics, transmission distance, and interference sensitivity greatly increase the complexity and challenge of cross-frequency band resource allocation. Effective allocation strategies not only need to consider the physical characteristics of the frequency bands, but also need to address the differentiated QoS needs of different services to optimize user experience.

[0003] When implementing a QoS-based bandwidth allocation strategy in a cross-frequency band environment, the main challenge faced by current technology is that it fails to fully consider the complex QoS differences between multiple services, making it difficult to effectively guarantee the service quality requirements of certain services, especially the delay and reliability indicators of some key services. Literature [4] focuses on the transmission of specific services, improving the reliability and rate of information transmission through spectrum sharing, but does not fully consider important delay performance requirements. However, in actual application scenarios, network environments often need to meet the QoS requirements of multiple services simultaneously. This not only requires optimization of the performance of a single service, but also needs to consider the transmission rate, packet delay, and packet loss rate of multiple services. Literature [5-8]The weighted QoS bias utility function integrates multiple indicators to meet the diverse QoS needs of users, but the need to adjust multiple weight parameters increases algorithm complexity and makes it difficult to ensure the accuracy of the configuration. [9]

[10] A delay bias-based utility function is then designed, which performs well in handling delay-sensitive services but ignores other key QoS indicators. Overall, QoS utility function-based methods have made some progress in improving multi-service QoS guarantees, but they usually only focus on a single performance indicator and cannot fully describe user needs. In contrast, the effective capacity model exhibits unique advantages.

[11] Effective capacity is used to evaluate users' diverse QoS needs, and a spectrum allocation algorithm is designed to meet users' multi-faceted QoS requirements. The effective capacity model is characterized by its comprehensiveness, covering not only transmission rate but also key performance indicators such as latency and packet loss rate. This makes it more valuable in practical applications than utility functions that only focus on a single indicator, such as Shannon capacity or delay.

[0004] In addition to building efficient QoS models, researchers have done a lot of work on specific allocation schemes.

[12] An improved bald eagle search algorithm (IBES) is proposed, which comprehensively considers user service QoS, channel state information, and resource periodicity. However, the QoS analysis in this literature is based on fixed parameters set in advance, which lacks flexibility and is difficult to adapt to the changing conditions in practical applications.

[13] The ideal fractional programming method (Ideal FP) proposed assumes that global channel state information (CSI) can be obtained instantaneously, which is difficult and cannot adapt to the complex real environment.

[14] Effective capacity is used to optimize the access right allocation of relay nodes through a forward auction algorithm, although this research has not been thoroughly explored in the context of multiple concurrent auction requests in ultra-dense networks. With the development of artificial intelligence, machine learning techniques have been widely applied in bandwidth allocation in wireless communication systems. [15-17] Combining multi-agent collaboration and reinforcement learning to optimize bandwidth allocation, but not fully considering the characteristics of the frequency band, limiting the space for improving the allocation effect.

[18] An improved deep deterministic policy gradient (DDPG) algorithm is proposed, which takes into account resources such as computation, storage, and network bandwidth. This method effectively handles complex resource allocation problems, but due to its high complexity, it may face real-time and computational resource limitations in practical deployment.

[19] Multi-Agent Double Deep Q Network (MADDQN) and Prioritized Multi-agent Deep Deterministic Policy Gradient (P-MADDPG) are developed to solve the bandwidth allocation sub-problem for each user, but the research mainly focuses on specific network settings and conditions, and does not consider the scalability and adaptability of the algorithm in other environments or different scale networks. Literature

[20] A hybrid deep deterministic policy gradient and deep Q network (Deep Q Network, DQN) method is proposed, which can more comprehensively deal with resource allocation problems, but the generalization and stability in different network sizes or configurations are not fully verified.

[0005] In summary, although these methods perform well in single frequency bands or uniform spectrum environments, they often fail to fully consider the characteristics and needs of different frequency bands in complex multi-frequency band environments. These methods usually treat all frequency bands as homogeneous resources, ignoring the differences in coverage, propagation loss, and data rate between frequency bands. At the same time, the detailed modeling of frequency band characteristics is lacking in the optimization process, and personalized resource allocation strategies for each frequency band cannot be achieved. Therefore, although there is overall improvement, it is still difficult to fully tap the potential and advantages of resource allocation in a multi-frequency band environment.

[0006] [1] N. Prodromos, D. Diasakos, V. Kokkinos, A. Gkamas, C. Bouras and P. Pouyioutas, "Dynamic Bandwidth Allocation in MIMO 5G Networks," 2024 International Wireless Communications and Mobile Computing (IWCMC), Ayia Napa, Cyprus, 2024, pp. 97-102, doi: 10.1109 / IWCMC61514.2024.10592577.

[0007] [2] Lopatka J, Paso T, Massin R, et al. Multi band efficient networks for ad hoc communications [J]. Procedia Computer Science, 2022, 205: 88-96.

[0008] [3] T. Eckert and S. Bryant, “Quality of Service (QoS),” in Future Networks, Services and Management: Underlay and Overlay, Edge, Applications, Slicing, Cloud, Space, AI / ML, and Quantum Computing, M. Toy, Ed. Cham: Springer International Publishing, 2021, pp. 309-344. doi: 10.1007 / 978-3-030-81961-3_11.

[0009] [4] L. Liang, S. Xie, G. Y. Li, Z. Ding, and X. Yu, “Graph-Based Resource Sharing in Vehicular Communication,” in IEEE Transactions on Wireless Communications, vol. 17, no. 7, pp. 4579-4592, July 2018, doi: 10.1109 / TWC.2018.2827958.

[0010] [5] L. Li, N. Deng, W. Ren, B. Kou, W. Zhou, and S. Yu, “Multi-Service Resource Allocation in Future Network With Wireless Virtualization,” in IEEE Access, vol. 6, pp. 53854-53868, 2018, doi: 10.1109 / ACCESS.2018.2871506.

[0011] [6] Liu, Q.; Li, R.; Li, Y.; Wang, P.; Sun, J. Adaptive Bandwidth Allocation for Massive MIMO Systems Based on Multiple Services. Appl. Sci. 2023, 13, 9861. https: / / doi.org / 10.3390 / app13179861

[0012] [7] Bendaoud F. Network selection in a heterogeneous wireless environment based on path prediction and user mobility [M] / / Comprehensive Guide to Heterogeneous Networks. Academic Press, 2023: 59-86.

[0013] [8] Huang, J.; Yang, Z.; Xie, J.; Zhang, H.; Li, Z. Joint Power and Bandwidth Allocation in Collocated MIMO Radar Based on the Quality of Service Framework. Electronics 2023, 12, 2567.

[0014] https: / / doi.org / 10.3390 / electronics12122567

[0015] [9] W. Huang, L. Ding, D. Meng, J.-N. Hwang, Y. Xu and W. Zhang, "QoE-Based Resource Allocation for Heterogeneous Multi-Radio Communication in Software-Defined Vehicle Networks," in IEEE Access, vol. 6, pp. 3387-3399, 2018, doi: 10.1109 / ACCESS.2018.2800036.

[0016]

[10] P. V. P K, K. Jagannathan, D. B and K. Milleth, "Resource Allocation for QoS Enforcement in 5G: Trading off Fairness and Delay-Awareness," 2024 National Conference on Communications (NCC), Chennai, India, 2024, pp. 1-6, doi: 10.1109 / NCC60321.2024.10485929.

[0017]

[11] X. Mi, L. Xiao, M. Zhao, X. Xu and J. Wang, "Statistical QoS-Driven Resource Allocation and Source Adaptation for D2D Communications Underlaying OFDMA-Based Cellular Networks," in IEEE Access, vol. 5, pp. 3981-3999, 2017, doi: 10.1109 / ACCESS.2017.2679113.

[0018]

[12] Liu, Q.; Li, R.; Li, Y.; Wang, P.; Sun, J. Adaptive Bandwidth Allocation for Massive MIMO Systems Based on Multiple Services. Appl. Sci. 2023, 13, 9861. https: / / doi.org / 10.3390 / app13179861

[0019]

[13] K. Shen and W. Yu, "Fractional Programming for Communication Systems—Part I: Power Control and Beamforming," in IEEE Transactions on Signal Processing, vol. 66, no. 10, pp. 2616-2630, 15 May 15, 2018, doi: 10.1109 / TSP.2018.2812733.

[0020]

[14] X. Huang, F. She, K. Wu and M. Jiang, "Device-to-Device Content Placement and Delivery Exploiting Joint Matching and Auction Theories," 2020 International Conference on Wireless Communications and Signal Processing (WCSP), Nanjing, China, 2020, pp. 795-800, doi: 10.1109 / WCSP49889.2020.9299732.

[0021]

[15] J. Wang, X. Zhang, X. He and Y. Sun, "Bandwidth Allocation and Trajectory Control in UAV-Assisted IoV Edge Computing Using Multiagent Reinforcement Learning," in IEEE Transactions on Reliability, vol. 72, no. 2, pp. 599-608, June 2023, doi: 10.1109 / TR.2022.3192020.

[0022]

[16] H. Albinsaid, K. Singh, S. Biswas and C.-P. Li, "Multi-Agent Reinforcement Learning-Based Distributed Dynamic Spectrum Access," in IEEE Transactions on Cognitive Communications and Networking, vol. 8, no. 2, pp. 1174-1185, June 2022, doi: 10.1109 / TCCN.2021.3120996.

[0023]

[17] B. Lim and M. Vu, "Distributed Multi-Agent Deep Q-Learning for Load Balancing User Association in Dense Networks," in IEEE Wireless Communications Letters, vol. 12, no. 7, pp. 1120-1124, July 2023, doi: 10.1109 / LWC.2023.3250492.

[0024]

[18] Q. Liu, H. Zhang, X. Zhang and D. Yuan, "Improved DDPG Based Two-Timescale Multi-Dimensional Resource Allocation for Multi-Access Edge Computing Networks," in IEEE Transactions on Vehicular Technology, vol. 73, no. 6, pp. 9153-9158, June 2024, doi: 10.1109 / TVT.2024.3360943.

[0025]

[19] Zhang, C.; Lv, T.; Huang, P.; Lin, Z.; Zeng, J.; Ren, Y. Joint Optimization of Bandwidth and Power Allocation in Uplink Systems with Deep Reinforcement Learning. Sensors 2023, 23, 6822. https: / / doi.org / 10.3390 / s23156822.

[0026]

[20] H. Hu, D. Wu, F. Zhou, X. Zhu, R. Q. Hu and H. Zhu, "Intelligent Resource Allocation for Edge-Cloud Collaborative Networks: A Hybrid DDPG-D3QN Approach," in IEEE Transactions on Vehicular Technology, vol. 72, no. 8, pp. 10696-10709, Aug. 2023, doi: 10.1109 / TVT.2023.3253905. SUMMARY

[0027] In order to solve the problem that the bandwidth allocation in a wireless communication system does not fully consider the QoS differences of multiple services and the characteristics of different frequency bands (Sub-6GHz and millimeter wave), resulting in poor user experience and low effective capacity, the application provides a multi-service QoS cross-frequency band bandwidth allocation method and system based on deep reinforcement learning.

[0028] According to a first aspect of the embodiments of the present disclosure, a multi-service QoS cross-frequency band bandwidth allocation method based on deep reinforcement learning is provided, comprising the following steps:

[0029] An effective capacity model is constructed, and key performance indicators such as transmission rates, delays and packet loss rates of different services are comprehensively considered to determine the overall optimization target of bandwidth resources.

[0030] Based on the effective capacity model, a multi-task hierarchical architecture based on deep reinforcement learning is constructed. The architecture performs preliminary allocation of overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through a high-level network to ensure reasonable allocation of bandwidth resources across frequency bands. Then, the architecture performs fine allocation of specific bandwidth resources within each frequency band through a low-level network to meet the QoS requirements of different services.

[0031] In the high-level network and the low-level network, a shared representation layer, a policy head and a value head are introduced respectively. The shared representation layer captures common features of different frequency bands, the policy head generates a specific allocation strategy, and the value head evaluates the allocation effect to achieve more accurate resource allocation.

[0032] According to a second aspect of the embodiments of the present disclosure, a multi-service QoS cross-frequency bandwidth allocation system based on deep reinforcement learning is provided, comprising:

[0033] A model construction module is configured to construct an effective capacity model, and key performance indicators such as transmission rates, delays and packet loss rates of different services are comprehensively considered to determine the overall optimization target of bandwidth resources.

[0034] An architecture design module is configured to construct a multi-task hierarchical architecture based on deep reinforcement learning based on the effective capacity model. The architecture performs preliminary allocation of overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through a high-level network to ensure reasonable allocation of bandwidth resources across frequency bands. Then, the architecture performs fine allocation of specific bandwidth resources within each frequency band through a low-level network to meet the QoS requirements of different services.

[0035] An allocation module is configured to introduce a shared representation layer, a policy head and a value head in the high-level network and the low-level network respectively. The shared representation layer captures common features of different frequency bands, the policy head generates a specific allocation strategy, and the value head evaluates the allocation effect to achieve more accurate resource allocation.

[0036] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which includes a memory, a processor and a computer program stored on the memory and running on the memory. The processor implements the multi-service QoS cross-frequency bandwidth allocation method based on deep reinforcement learning when executing the program.

[0037] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program. The program is executed by a processor to implement the multi-service QoS cross-frequency bandwidth allocation method based on deep reinforcement learning.

[0038] Compared with the prior art, the above technical scheme adopted by the present application has the advantages that the present application proposes a multi-service QoS cross-frequency band bandwidth allocation method and system based on deep reinforcement learning, aiming to solve the problems of poor user experience and low effective capacity caused by insufficient consideration of multi-service QoS differences and different frequency characteristics in the current wireless communication system. By constructing an effective capacity model, a double-layer deep reinforcement learning architecture is designed, and a shared representation layer, a policy head and a value head are introduced in the high layer and the low layer respectively, realizing effective management of Sub-6GHz and millimeter wave frequency band resources. The high layer strategy performs cross-frequency band rough bandwidth allocation through the MADDPG algorithm, and the low layer strategy performs fine allocation of bandwidth in each frequency band through the DDPG algorithm, further meeting the QoS requirements of each service.

[0039] The present application not only can maintain a high average effective capacity under different user quantities, but also can better cope with the network pressure brought by the increase of user quantity, and exhibits more stable performance. In addition, by combining global and local strategies, the global optimization and local optimization requirements in resource allocation are successfully balanced, which performs outstandingly in reducing delay, improving SINR and improving user experience, and provides a solid foundation for further improving system throughput in the future. BRIEF DESCRIPTION OF DRAWINGS

[0040] The drawings accompanying the specification of this application form a part thereof, serve to provide further understanding of the present application, and together with the specification of the present application, serve to explain the present application, and do not constitute an improper limitation on the present application.

[0041] Figure 1 Sub-6GHz and millimeter wave cooperative networking scenario diagram;

[0042] Figure 2 Multi-task hierarchical architecture principle diagram based on deep reinforcement learning;

[0043] Figure 3 Actor network action decision flowchart;

[0044] Figure 4 Critic network value evaluation flowchart. DETAILED DESCRIPTION

[0045] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0046] It should be pointed out that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the present application have the same meaning as generally understood by those skilled in the art to which the present application belongs.

[0047] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0048] It should be noted that the flow diagrams and block diagrams in the drawings are representative of the architecture, functionality, and operation of possible implementations of methods and systems according to various embodiments of the present disclosure. It should also be noted that each block in the flow diagrams and block diagrams can represent a module, a segment, or a portion of code, which includes one or more executable instructions for implementing the specified logical functions ("instructions"). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each of the blocks of the flow diagrams and / or block diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and

[0049] Embodiment One:

[0050] The embodiment provides a multi-service QoS cross-frequency bandwidth allocation method based on deep reinforcement learning, including the following steps:

[0051] S1. Construct an effective capacity model, comprehensively consider the key performance indicators of different services, and clearly define the overall optimization target of bandwidth resources;

[0052] Specifically, in order to realize multi-service QoS guarantee, low-frequency base stations (Low-frequency Base Station, LF-BS) and millimeter wave high-frequency base stations (How-frequency Base Station, HF-BS) below 6GHz are fused, and high-low frequency collaborative networking is realized, as shown in Figure 1 The Sub-6GHZ frequency band of the LF-BS is equally divided into N s bandwidth resource blocks, and the millimeter wave frequency band of the HF-BS is equally divided into N mvEach bandwidth resource block hosts services with different QoS requirements. Users can generate communication links with either LF-BSs or HF-BSs, and link m consists of receiver m and its transmitter m. Transmission links between users and base stations use orthogonal channels, so there is no inter-cell downlink interference in each HF-BS microcell coverage range during downlink transmission.

[0053] Sub-6GHz and millimeter wave frequency bands have different requirements for antennas, which is reflected in the calculation of signal-to-noise ratio. Sub-6GHz frequency bands usually only consider antenna power gain because it is sufficient to describe relatively uniform signal radiation and reception characteristics. However, millimeter wave frequency bands need to consider the directional antenna gain of the transmitting and receiving ends to ensure effective signal transmission and reception in the case of short wavelengths and complex propagation environments, thereby improving system reliability and performance.

[0054] The signal-to-interference-plus-noise ratio (SINR) at the receiver m of the low-frequency base station LF-BS can be expressed as:

[0055]

[0056] wherein, is the low-frequency cellular antenna power gain of Sub-6GHz, N s is the additive white Gaussian noise (AWGN) power spectral density of the cellular spectrum resource, B s is the bandwidth of the cellular user spectrum resource; p m represents the transmission power on link m, h m is the channel gain from user m to the LF-BS, is the interference from other links to link m.

[0057] The signal-to-interference-plus-noise ratio at the receiver m of the high-frequency base station HF-BS is expressed as:

[0058]

[0059] wherein, and are the directional antenna gains of the HF-BS transmitter and receiver, respectively. N mv is the additive white Gaussian noise power spectral density of the millimeter wave spectrum resource, B mv is the bandwidth of the millimeter wave user spectrum resource;

[0060] The effective capacity model comprehensively considers the key indicators of transmission rate, delay and packet loss rate, quantifies the average data transmission rate of the link under different frequency bands through the effective capacity, and describes the average data transmission rate that can be achieved under the given wireless channel conditions to meet different QoS requirements of services:

[0061]

[0062] Where T represents the time slot length, T D represents the data packet delay, and t delay is set as the maximum tolerable delay of the data packet (when the data packet exceeds t delay , it is considered lost), R represents the link transmission rate, and ε represents the maximum allowed packet loss rate. θ is a QoS guarantee parameter, which covers key performance indicators such as transmission rate, delay and packet loss rate. For real-time data traffic transmission services, the delay needs to be strictly controlled, i.e. θ→0, and the effective capacity is expressed as the outage capacity; while for non-real-time data transmission services, the focus is on achieving high throughput while allowing loose delay constraints, i.e. θ→∞, and the effective capacity will become the throughout capacity. The relationship between the QoS guarantee parameter θ and the data packet arrival rate, delay and packet loss rate is expressed as follows:

[0063]

[0064] Therefore, based on the above definition, the effective capacity model of the LF-BS and the HF-BS is constructed:

[0065]

[0066]

[0067]

[0068] Wherein, is the delay QoS constraint parameter of the link. α m represents the communication mode selected by the link, α m =0 represents selecting LS-BS communication, α m =1 represents selecting HF-BS communication, and 0<α m <1 represents the proportion of the millimeter wave bandwidth allocated by the HF-BS on the link.

[0069] As a preferred embodiment, in order to further ensure the transmission reliability, a data overflow / loss constraint is introduced on the basis of the effective capacity model, which is an important limiting condition to ensure that the resource allocation scheme meets the reliability requirements of the actual system;

[0070] Assuming that the data packet is infinitely small, the source node of the link m transmits data packets at a constant rate λ mGeneration / arrival; these data packets will be temporarily stored in a First In First Out (FIFO) queue before transmission. The link m data packet loss probability can be expressed as:

[0071]

[0072] Where T m represents the data delay of link m, Q m represents the buffer queue length of link m, represents the maximum allowable data delay of service b, represents the maximum allowable buffer queue length of service b, data transmission exceeds the maximum delay threshold or the data in the node buffer exceeds the maximum queue length threshold , data loss. l(λ m ) represents the probability of non-empty in the link m buffer:

[0073]

[0074] Therefore, the constraint condition of the data overflow / loss probability of the link is expressed as:

[0075]

[0076] Where ε b represents the maximum data overflow / loss probability threshold allowed by service b.

[0077] It should be noted that the present application formulates the bandwidth resource allocation problem as an optimization problem, which is described as follows: in order to guarantee the QoS requirements of multiple services while improving the upper limit of system throughput, the present application takes the maximization of link effective capacity as the overall optimization target of bandwidth resource, and constructs a cross-frequency bandwidth allocation problem based on multi-service QoS, which takes the overflow / loss probability of transmission delay exceeding the service delay threshold, the link capacity higher than the data arrival rate, and the minimum communication quality guarantee as constraint conditions;

[0078]

[0079] Where, and respectively represent the allocation matrix of link m on the Sub6-GHz bandwidth resource block n s and the millimeter wave bandwidth resource block n mv ; θ = [θ m ] is the QoS guarantee parameter configuration matrix.

[0080] S2. Based on the effective capacity model, a multi-task hierarchical architecture based on deep reinforcement learning is constructed, which preliminarily allocates the overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through a high-level network, ensures the reasonable allocation of bandwidth resources across frequency bands, and then finely allocates the specific bandwidth resources in each frequency band through a low-level network to meet the QoS requirements of different services.

[0081] The optimization problem described in the application is a non-convex optimization problem, and the multi-service QoS cross-frequency bandwidth allocation method aims to maximize the total effective capacity of the network while meeting the QoS requirements of different service types. To achieve this goal, a multi-task hierarchical deep reinforcement learning architecture is designed, which combines MADDPG and DDPG algorithms to optimize high-level and low-level decision-making requirements respectively; each communication link is regarded as an independent agent.

[0082] In the high-level network decision-making stage, the main task is to preliminarily allocate the total bandwidth resources of the Sub-6GHz and millimeter wave frequency bands. This stage not only requires coordination of resource allocation among multiple communication links, but also fully considers the characteristics of the frequency bands and the dynamic needs of each link, and preferentially allocates resources to those links that most need bandwidth. Therefore, the MADDPG algorithm is selected. MADDPG is good at handling cooperation and competition in multi-agent environment, and can generate optimal allocation strategies across links based on global information through the strategy network and value network of each link agent, to preliminarily allocate the total bandwidth.

[0083] In the low-level network decision-making stage, the task changes to finely allocate specific bandwidth resources within each frequency band. The decision-making in this stage is more localized, mainly focusing on the optimization of resource allocation within each link's frequency band. Since each link's decision-making at this time focuses more on the bandwidth allocation within its own frequency band, the more simple and efficient DDPG algorithm is selected. DDPG is suitable for single-agent environment and can quickly generate action decisions to further optimize bandwidth allocation for each link's resource needs in a specific frequency band. Through independent strategy network and value network, DDPG refines the bandwidth resources allocated by the high-level and accurately meets the quality of service requirements.

[0084] To implement multi-task hierarchical bandwidth resource allocation, the state space, action space and reward function are first defined:

[0085] The state space S includes: the demand matrix θ of the receiver m , the instantaneous channel state information h m , the current service b tolerable delay threshold of the link transmitter and the maximum acceptable data overflow / loss probability The instantaneous traffic arrival rate λ of the linkm The state space of high-level and low-level decisions is the same, but the low-level is more detailed to the state of specific links and frequency bands.

[0086] The action space A of it includes: the action space of high-level network is defined as the total bandwidth resource allocation strategy of links in different frequency bands, ρ m is the allocation proportion of Sub6-GHz bandwidth resources of LF-BS, α m represents the allocation proportion of millimeter wave bandwidth resources of HF-BS. The action space of low-level network is represented as A = {ρ, α}, and the specific resource allocation of each frequency band is optimized independently to adapt to the characteristics and needs of different frequency bands.

[0087] The reward function R of it includes: the instantaneous reward of high-level network is designed as a global effective capacity indicator, aiming to provide as high a system transmission rate as possible under the premise of meeting the QoS requirements of different services. Therefore, the instantaneous reward of the tth time slot can be represented as:

[0088]

[0089] The instantaneous reward of low-level network is designed as the average effective capacity that the agent and its neighbor agents can achieve. Therefore, the instantaneous reward of the tth time slot can be represented as:

[0090]

[0091] wherein, is the sum of the current agent and the observed number of neighbor agents.

[0092] Multi-Task Learning (MTL) is a method that improves the generalization ability of a model by simultaneously learning multiple related tasks. In the bandwidth resource allocation problem, multi-task learning can effectively handle the QoS requirements of multiple services and improve resource utilization efficiency. Traditional single-task training methods often ignore valuable information in related tasks, while multi-task learning allows the model to learn multiple related tasks, extract common information between tasks, and share representations between related tasks, so that the model can better generalize to the original task and improve performance on specific indicators.

[0093] The core of MTL is to share the underlying representation learning and handle multiple tasks in parallel, thereby improving the performance of each task, capturing the commonality and relationship between tasks, and ultimately improving the overall performance of the model.

[0094] S3. A shared representation layer, a policy head and a value head are introduced in the high-level network and the low-level network respectively, wherein the shared representation layer captures the common features of different frequency bands, the policy head generates specific allocation strategies, and the value head evaluates the allocation effect to achieve more accurate resource allocation.

[0095] Specifically, to achieve the goal of multi-task learning, the present application designs an architecture with a shared representation layer, a policy head, and a value head, as shown in Figure 2 .

[0096] The shared representation layer allows multiple tasks or agents to use the same feature representation in the feature learning process by extracting and sharing basic features. This sharing mechanism improves the learning efficiency and generalization ability of the model, reduces the risk of overfitting, and makes the model not only perform well on training data but also better adapt to new data. In the bandwidth resource allocation task, the shared representation layer helps to improve task performance, capture the inherent rules of data, and improve the robustness to noise and abnormal data. In addition, it also promotes transfer learning, enabling general-purpose features to be more effectively applied to new tasks or data distributions, further enhancing the model's generalization ability. The shared representation layer extracts useful features from the input state through a neural network layer. Specifically, for the input state s, the shared representation layer extracts features through the parameterized function .

[0097]

[0098] Here, represents the features extracted from the input state s, represents the neural network parameters of the shared representation layer.

[0099] ① Shared representation layer of high-level network: In high-level decision-making, MADDPG needs to optimize the spectrum resource allocation strategy from a global perspective. The shared representation layer extracts common features from the input state of different links, which are used for high-level resource allocation decisions.

[0100]

[0101] ② Shared representation layer of low-level network: In low-level decision-making, DDPG is responsible for specific intra-band resource allocation optimization. Adding a shared representation layer can help the low-level model better understand the features and patterns within the frequency band and improve the optimization effect of the specific allocation strategy. The shared representation layer extracts features within the specific frequency band, which are used for intra-band resource allocation optimization.

[0102]

[0103] DDPG and MADDPG both use two main networks: Actor network and Critic network. By sharing the features extracted from the representation layer, the policy head and the value head are fed to generate specific policies and estimate the value of actions. The policy head and the value head are the specific implementation parts of the Actor and Critic networks. They are responsible for generating policies (decisions) and evaluating values (evaluations), respectively.

[0104] The role of the Actor network is to generate specific actions according to the current state. After introducing the shared representation layer, the Actor network can use the shared feature representation to generate more consistent and coordinated action decisions. The shared representation layer can be seen as a general feature extractor, and these features are fed to the policy head to generate actions, as shown in Figure 3 .

[0105] The role of the Critic network is to evaluate the value of state-action pairs. After introducing the shared representation layer, the Critic network can use the shared feature representation to make more accurate value evaluations. The shared representation layer is also used to extract state features, which are fed to the value head along with actions to evaluate Q values, so that the Critic network can better evaluate the value of the current action, as shown in Figure 4 .

[0106] The DDPG used by the low-level network decision combines the advantages of deep Q network and deterministic policy gradient (DPG), which directly outputs deterministic actions through the policy network and uses the Q value network to evaluate these actions. The policy network and the Q value network use neural networks to approximate, respectively, where the policy network generates the best action a according to the current state s (by generating a specific resource allocation policy through the policy head), and the Q value network evaluates the value of the state-action pair (s, a) (by evaluating the value of the current action through the value head). DDPG uses an experience replay mechanism to break the data correlation by randomly sampling from the experience (s, a, r, s') stored in the experience replay pool D, improving the training stability. At the same time, DDPG uses a soft update method to update the target network, further improving the stability and performance of the algorithm. The network parameters of the Actor network and the Critic network of DDPG are defined as θ μ and θ Q , and the target network parameters are θ μ' and θ Q' , respectively. In the Actor network, let μ(a|s) be the probability of taking action a under the current state s, and the Critic network gives the Q value according to s and a. The learning purpose is to maximize the expected sum:

[0107]

[0108] where γ ∈ (0, 1) is a discount factor that reflects how important future rewards are relative to the current policy.

[0109] Through the policy head Generate resource allocation strategies for specific frequency bands.

[0110] Through the value head Evaluate the expected total return of the current action.

[0111] Therefore, the total expected value under the policy μ can be written as:

[0112]

[0113] β is the behavior decision, which is stored in the experience replay D in the form of a quadruple (s, a, r, s') during the exploration process, s' is the next state that s is transferred to by action a. During the training process, a small batch of samples are extracted from D to update the parameters.

[0114] The loss function of the Critic network is updated as:

[0115]

[0116] where the target

[0117] In the MADDPG used by the high-level network decision, each agent has its own Actor and Critic network, and the Critic can access the information of all agents during training, which helps to make better decisions in the environment with other agents. First, initialize the four network parameters as DDPG, and the network parameters for agent m can be represented as

[0118] Through the policy head Generate resource allocation strategies for specific frequency bands.

[0119] Through the value head Evaluate the expected total return of the current action.

[0120] The loss function of the Critic network is updated as:

[0121]

[0122] where the target

[0123] Embodiment two:

[0124] The embodiment provides a multi-service QoS cross-frequency band bandwidth allocation system based on deep reinforcement learning, which comprises:

[0125] A model construction module constructs an effective capacity model, comprehensively considers transmission rates, delays, and packet loss rates of different services as key performance indicators, and explicitly defines an overall optimization target of bandwidth resources.

[0126] An architecture design module constructs a multi-task hierarchical architecture based on deep reinforcement learning, taking the effective capacity model as a basis. The architecture preliminarily allocates overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through a high-level network, ensures reasonable allocation of bandwidth resources across frequency bands, and then finely allocates specific bandwidth resources within each frequency band through a low-level network to meet QoS requirements of different services.

[0127] An allocation module introduces a shared representation layer, a policy head, and a value head in the high-level network and the low-level network, respectively. The shared representation layer captures common features of different frequency bands, the policy head generates a specific allocation strategy, and the value head evaluates the allocation effect to achieve more accurate resource allocation.

[0128] Embodiment three

[0129] An electronic device includes a memory, a processor, and a computer program stored on the memory and running on the memory. When the processor executes the program, the method for allocating bandwidth across frequency bands based on deep reinforcement learning for multiple services QoS is realized, including:

[0130] An effective capacity model is constructed, and key performance indicators of different services are comprehensively considered to explicitly define an overall optimization target of bandwidth resources.

[0131] A multi-task hierarchical architecture based on deep reinforcement learning is constructed, taking the effective capacity model as a basis. The architecture preliminarily allocates overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through a high-level network, ensures reasonable allocation of bandwidth resources across frequency bands, and then finely allocates specific bandwidth resources within each frequency band through a low-level network to meet QoS requirements of different services.

[0132] A shared representation layer, a policy head, and a value head are introduced in the high-level network and the low-level network, respectively. The shared representation layer captures common features of different frequency bands, the policy head generates a specific allocation strategy, and the value head evaluates the allocation effect to achieve more accurate resource allocation.

[0133] Embodiment four

[0134] A computer-readable storage medium stores a computer program, which is executed by a processor to realize the method for allocating bandwidth across frequency bands based on deep reinforcement learning for multiple services QoS, including:

[0135] An effective capacity model is constructed, and key performance indicators of different services are comprehensively considered to explicitly define an overall optimization target of bandwidth resources.

[0136] Based on the effective capacity model, a multi-task hierarchical architecture based on deep reinforcement learning is constructed, which preliminarily allocates the overall bandwidth resources of the Sub-6GHz and millimeter wave frequency bands through the high-level network, ensures the reasonable allocation of bandwidth resources across frequency bands, and then further allocates the specific bandwidth resources within each frequency band through the low-level network to meet the QoS requirements of different services.

[0137] The shared representation layer, policy head and value head are introduced in the high-level network and low-level network respectively, wherein the shared representation layer captures the common features of different frequency bands, the policy head generates specific allocation strategies, and the value head evaluates the allocation effect to realize more accurate resource allocation.

[0138] Before implementing the above-mentioned multi-service QoS cross-frequency band bandwidth allocation method based on deep reinforcement learning, the basic parameters of the wireless communication system are set, including frequency band division, user demand, channel state, etc. The parameters of the multi-task hierarchical architecture of deep reinforcement learning are initialized, including the weights of the high-level and low-level policy networks.

[0139] It should be noted that the high-level reward function is designed as a global effective capacity indicator, aiming to optimize the total transmission rate of the system. The low-level reward function is designed as the average effective capacity of the specific link and the neighbor link. The experience replay mechanism is used to randomly sample historical data for model training, and the policy network and value network are optimized. The soft update mechanism is used to update the target network to enhance the stability and performance of the algorithm. In each time slot, the bandwidth resources are preliminarily allocated according to the high-level policy, and then the specific frequency band is fine-tuned by the low-level policy.

[0140] Taking a specific scenario as an example, assuming that multiple users simultaneously request different types of services (such as video streaming, online gaming, etc.), the steps are as follows:

[0141] 1) User demand analysis: Collect the QoS requirements (such as bandwidth, delay requirements) of each user.

[0142] 2) Preliminary resource allocation: The high-level policy preliminarily allocates bandwidth resources according to user demand and channel state.

[0143] 3) Fine allocation: The low-level policy further adjusts the bandwidth allocation of each frequency band to meet the QoS requirements of specific users.

[0144] 4) Performance evaluation: Real-time monitoring of network performance, evaluation of effective capacity and delay, and optimization of strategy according to feedback.

[0145] In summary, the present application can efficiently realize intelligent bandwidth resource management across frequency bands, and is suitable for various application scenarios of modern wireless communication systems.

[0146] Those skilled in the art should understand that the modules or steps of the present disclosure described above can be realized by a general computer device, and alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be respectively manufactured into individual integrated circuit modules, or a plurality of modules or steps among them can be manufactured into a single integrated circuit module. The present disclosure is not limited to any specific combination of hardware and software.

[0147] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0148] The above describes the specific embodiments of the present disclosure in conjunction with the accompanying drawings, but is not intended to limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present disclosure without creative labor are still within the protection scope of the present disclosure.

Claims

1. A multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning, characterized by: The following steps are involved: Build an effective capacity model, comprehensively consider the key performance indicators of different services, and clarify the overall optimization goal of bandwidth resources; Based on the effective capacity model, a multi-task hierarchical architecture based on deep reinforcement learning was constructed. This architecture uses the high-level network to initially allocate overall bandwidth resources for the Sub-6 GHz and millimeter wave bands, ensuring a reasonable distribution of bandwidth resources across the bands. The low-level network then fine-tunes the allocation of specific bandwidth resources within each band to meet the QoS requirements of different services. A shared representation layer, a policy head, and a value head are introduced into the high-level network and the low-level network, respectively. The shared representation layer captures the common characteristics of different frequency bands, the policy head generates specific allocation strategies, and the value head evaluates the allocation effect, achieving more accurate resource allocation. For the input state s, the shared representation layer is parameterized by the function To extract features: in, represents the features extracted from the input state s, Represents the neural network parameters of the shared representation layer; The shared representation layer of the high-level network extracts common features from the input states of different links. These features are used for high-level resource allocation decisions: The shared representation layer of the lower-level network extracts features within a specific frequency band, which are used to optimize resource allocation within the specific frequency band. The features extracted by the shared representation layer are passed to the policy head and the value head. The policy head and the value head correspond to the actor network and the critic network in deep reinforcement learning, which are responsible for generating strategies and evaluating values ​​respectively. The DDPG algorithm is used in the low-level network, and the network parameters of the Actor network and the Critic network are defined as θ μ and θ Q , the target network parameters are θ μ' and θ Q' In the Actor network, let μ(a|s) be the probability of taking action a in the current state s, and the Critic network gives a Q value based on s and a. The learning goal is to maximize the expected sum: Where γ∈(0,1) is the discount factor of the relative importance of future rewards to the current strategy; Through the policy header Generate resource allocation strategies for specific frequency bands; By value header Evaluate the expected total reward of the current action; The total expected value under strategy μ can be written as: Among them, β is the behavioral decision, in the exploration process (s,a,r,s ′ ) is stored in the experience playback D in the form of a four-tuple, s ′ is the next state that s transitions to through action a; The loss function of updating the Critic network is expressed as: The target The MADDPG algorithm is used in the high-level network. Each agent has a corresponding Actor network and Critic network. The network parameters for agent m are expressed as Through the policy header Generate resource allocation strategies for specific frequency bands; By value header Evaluate the expected total reward of the current action; The loss function of updating the Critic network is expressed as: The target 2. The multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning according to claim 1 is characterized in that: The effective capacity model is constructed as follows: The low-frequency base station LF-BS below 6 GHz and the millimeter-wave high-frequency base station HF-BS are integrated to achieve high- and low-frequency collaborative networking. The signal-to-interference-and-noise ratio at the receiver m of the low-frequency base station LF-BS is expressed as: in, is the power gain of the Sub-6GHz low-frequency cellular antenna, N s is the power spectral density of additive white Gaussian noise of cellular spectrum resources, B s is the bandwidth of cellular user spectrum resources; p m represents the transmission power on link m, h m is the channel gain from user m to LF-BS, is the sum of interferences from other links on link m; The signal-to-interference-and-noise ratio at the receiver m of the high-frequency base station HF-BS is expressed as: in, and are the directional antenna gains of HF-BS transmitting and receiving users respectively; N mv is the power spectral density of additive white Gaussian noise of millimeter wave spectrum resources, B mv The bandwidth of the spectrum resources for millimeter wave users; Effective capacity describes the average data transmission rate achieved under given wireless channel conditions while meeting the QoS requirements of different services: Where T represents the time slot length, T D Indicates the packet delay, set t delay is the maximum tolerable delay of the data packet, R is the link transmission rate, and ε is the maximum allowable packet loss rate. θ is the QoS guarantee parameter, covering the key performance indicators of transmission rate, delay, and packet loss rate. The relationship between them is shown as follows: Therefore, the effective capacity model of LF-BS and HF-BS is: in, is the link delay QoS constraint parameter; α m represents the communication mode selected by the link, α m =0 means LS-BS communication is selected, α m =1 means HF-BS communication is selected, 0<α m <1 indicates the ratio of millimeter wave bandwidth allocated to HF-BS on the link.

3. The multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning according to claim 2 is characterized in that: The data overflow / loss constraint is introduced into the effective capacity model. Specifically, the probability of data packet loss on link m is expressed as: Among them, T m represents the data delay of link m, Q m represents the buffer queue length of link m, Indicates the maximum allowable data delay of service b, Indicates the maximum allowable cache queue length for service b, the data transmission exceeds the maximum delay threshold, or the node cached data exceeds the maximum queue length threshold. When , data is lost; l(λ m ) represents the probability that the cache of link m is not empty: Therefore, the constraint on the link data overflow / loss probability is expressed as: Among them, ε b Indicates the maximum data overflow / loss probability threshold allowed for service b.

4. The multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning according to claim 1 is characterized in that: Maximizing effective link capacity is the overall optimization goal for bandwidth resources. This solves the cross-band bandwidth allocation problem for multi-service QoS. This problem is subject to the following constraints: the overflow / loss probability of transmission delay exceeding the service delay threshold, the link capacity exceeding the data arrival rate, and the minimum communication quality guarantee: in, and They represent link m in Sub6-GHz bandwidth resource block n. s and millimeter wave bandwidth resource block n mv The distribution matrix on θ=[θ m ] is the QoS guarantee parameter configuration matrix.

5. The multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning according to claim 1 is characterized in that: The state space, action space, and reward function are defined separately in a multi-task hierarchical architecture based on deep reinforcement learning; Its state space S includes: the receiver's demand matrix θ m , instantaneous channel state information h m , the tolerable delay threshold of the current service b of the link transmitter and the maximum acceptable data overflow / loss probability The instantaneous traffic arrival rate of the link λ m The state space for high-level and low-level network decisions is the same, but the low-level network is refined to the state of specific links and frequency bands. Its action space A includes: The action space of the high-level network is defined as the total bandwidth resource allocation strategy of the link in different frequency bands, ρ m is the allocation ratio of Sub6-GHz bandwidth resources of LF-BS, α m represents the allocation ratio of millimeter wave bandwidth resources to the HF-BS; the action space of the low-layer network is expressed as A = {ρ, α}, which independently optimizes the specific resource allocation of each frequency band to adapt to the characteristics and requirements of different frequency bands; Its reward function R includes: The instantaneous reward of the t-th time slot of the high-level network is expressed as: The instantaneous reward of the tth time slot of the low-level network is expressed as: in, is the sum of the number of the current agent and the observed neighboring agents.

6. A multi-service QoS cross-band bandwidth allocation system based on deep reinforcement learning, characterized by: include: The model building module builds an effective capacity model, comprehensively considering key performance indicators such as transmission rate, latency, and packet loss rate of different services, and clearly defines the overall optimization goal of bandwidth resources; The architecture design module uses the effective capacity model as its foundation and constructs a multi-task hierarchical architecture based on deep reinforcement learning. This architecture uses the high-level network to initially allocate overall bandwidth resources for the Sub-6 GHz and millimeter wave bands, ensuring a reasonable distribution of bandwidth resources across the bands. The low-level network then fine-tunes the allocation of specific bandwidth resources within each band to meet the QoS requirements of different services. The allocation module introduces a shared representation layer, a policy head, and a value head in the high-level network and the low-level network, respectively. The shared representation layer captures the common characteristics of different frequency bands, the policy head generates specific allocation strategies, and the value head evaluates the allocation effect, achieving more accurate resource allocation. For the input state s, the shared representation layer is parameterized by the function To extract features: in, represents the features extracted from the input state s, Represents the neural network parameters of the shared representation layer; The shared representation layer of the high-level network extracts common features from the input states of different links. These features are used for high-level resource allocation decisions: The shared representation layer of the lower-level network extracts features within a specific frequency band, which are used to optimize resource allocation within the specific frequency band. The features extracted by the shared representation layer are passed to the policy head and the value head. The policy head and the value head correspond to the actor network and the critic network in deep reinforcement learning, which are responsible for generating strategies and evaluating values ​​respectively. The DDPG algorithm is used in the low-level network, and the network parameters of the Actor network and the Critic network are defined as θ μ and θ Q , the target network parameters are θ μ' and θ Q' In the Actor network, let μ(a|s) be the probability of taking action a in the current state s, and the Critic network gives a Q value based on s and a. The learning goal is to maximize the expected sum: Where γ∈(0,1) is the discount factor of the relative importance of future rewards to the current strategy; Through the policy header Generate resource allocation strategies for specific frequency bands; By value header Evaluate the expected total reward of the current action; The total expected value under strategy μ can be written as: Among them, β is the behavioral decision, in the exploration process (s,a,r,s ′ ) is stored in the experience playback D in the form of a four-tuple, s ′ is the next state that s transitions to through action a; The loss function of updating the Critic network is expressed as: The target The MADDPG algorithm is used in the high-level network. Each agent has a corresponding Actor network and Critic network. The network parameters for agent m are expressed as Through the policy header Generate resource allocation strategies for specific frequency bands; By value header Evaluate the expected total reward of the current action; The loss function of updating the Critic network is expressed as: The target 7. An electronic device comprising a memory, a processor, and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning is implemented as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-service QoS cross-band bandwidth allocation method based on deep reinforcement learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-band network resource allocation method based on ML-MADDPG

    CN120475414A