Resource allocation method of multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning
By adopting a resource allocation method of layered multi-agent reinforcement learning in the multi-hop access backhaul integrated network, dynamically allocating resource scheduling information and combining local information for spectrum and power allocation, the problem of complexity and low efficiency of resource allocation in the network is solved, and more efficient resource utilization and communication quality are achieved.
Patent Information
- Application Number
- CN202510287679.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-20
AI Technical Summary
In the integrated multi-jump access backhaul integrated network, the existing technology is difficult to effectively solve the complexity and efficiency of resource allocation, especially in large-scale networks. Traditional centralized and distributed solutions have the risk of large overhead, low efficiency and resource allocation conflicts.
The resource allocation method based on hierarchical multi-agent reinforcement learning is adopted, and the user satisfaction and resource utilization of each IAB-donor are obtained through IAB-donor, resource scheduling information is dynamically allocated, spectrum and power are allocated, and distributed distributed dynamic resource optimization is achieved.
Effectively respond to the ever-changing traffic demands and link states in the network, reduce interference, optimize power control, and improve the resource utilization and communication quality of the overall network.
Smart Images

Figure CN120186771A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of integrated access and backhaul, and particularly to a resource allocation method for a multi-hop integrated access and backhaul network based on hierarchical multi-agent reinforcement learning. Background Art
[0002] In the related art, the 3rd Generation Partnership Project has launched the Integrated Access and Backhaul (IAB) technology, which is a wireless backhaul solution that allows base stations to transmit data without traditional wired backhaul links, enabling rapid network expansion and dynamic adjustment according to user needs. The IAB technology has advantages in areas with high density and large coverage requirements, but also faces challenges in resource allocation. Especially in a multi-hop IAB network, the multi-hop transmission between IAB-nodes makes the network topology complex, increasing the difficulty of resource allocation. Existing centralized allocation schemes need to rely on global network information, resulting in large information transmission overhead in large-scale networks, low efficiency, and inability to effectively handle bursty traffic; while fully distributed schemes may lead to resource allocation conflicts or interference due to only using local information, reducing network performance.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The present invention provides a resource allocation method for a multi-hop integrated access and backhaul network based on hierarchical multi-agent reinforcement learning, and a computer program product, thereby overcoming the defects existing in the prior art to a certain extent.
[0005] Other features and advantages of the present invention will become apparent through the following detailed description, or be learned in part through the practice of the present invention.
[0006] According to a first aspect of the present invention, there is provided a resource allocation method for a multi-hop integrated access and backhaul network based on hierarchical multi-agent reinforcement learning, the method comprising:
[0007] The IAB-donor obtains the user satisfaction and resource utilization rate of each IAB-node during the current centralized scheduling period; wherein, the user satisfaction is used to describe the data transmission rate, and the resource utilization rate is used to describe the spectrum resource utilization efficiency;
[0008] The IAB-donor allocates resource scheduling information corresponding to the next centralized scheduling period for each IAB-node according to the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period; wherein, the resource scheduling information includes schedulable sub-channels and the maximum transmit power of the sub-channels.
[0009] In some exemplary embodiments, the centralized scheduling period includes N distributed scheduling periods; the method further includes:
[0010] In each distributed scheduling period, the IAB-node combines local information with the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor to allocate spectrum and power for the sub-nodes and UEs associated with the IAB-node, realizing distributed dynamic resource optimization of the IAB-node based on local information.
[0011] In some exemplary embodiments, the local information of the IAB-node includes at least one of the following:
[0012] The configuration results of the spectrum resources and power resources of the sub-nodes and UEs by the IAB-node in the previous distributed scheduling period;
[0013] The channel gains between the IAB-node and the sub-nodes and associated UEs in the current distributed scheduling period;
[0014] The channel interference between the IAB-node and the sub-nodes and associated UEs in the previous distributed scheduling period; wherein, the channel interference includes at least one of co-layer interference, cross-layer interference, and self-interference generated due to the use of the same sub-channel.
[0015] In some exemplary embodiments, the method further includes:
[0016] At the end of each centralized scheduling period, the IAB-node calculates the user satisfaction and resource utilization rate in the centralized scheduling period and uploads them to the IAB-donor;
[0017] The IAB-donor obtains the user satisfaction and resource utilization rate of the IAB-node in the current centralized scheduling period in the last frame of each centralized scheduling period.
[0018] In some exemplary embodiments, the IAB-node calculates the user satisfaction in the centralized scheduling period, including:
[0019] During a distributed scheduling period, determine the satisfaction degree of the UE group in the distributed scheduling period according to the ratio of the average data transmission rate of the UE group associated with the IAB-node to the minimum data transmission rate threshold set according to the basic service requirements of the users.
[0020] Determine the user satisfaction degree of the UE group associated with the IAB-node in a centralized scheduling period according to the satisfaction degrees of the UE groups in multiple distributed scheduling periods.
[0021] In some exemplary embodiments, the available spectrum resources of the IAB network include: M available orthogonal sub-channels; the orthogonal sub-channels include: sub-channels schedulable by the master relay node and the secondary relay node.
[0022] The method further includes:
[0023] Define the signal-to-noise ratio constraints of each type of node in the sub-channel of the IAB network, including: configure the signal-to-noise ratio of the master relay node on the sub-channel according to the co-layer interference, cross-layer interference, and residual self-interference; configure the signal-to-noise ratio of the UE group connected to the master relay node on the sub-channel according to the co-layer interference and cross-layer interference; configure the signal-to-noise ratio of the UE group connected to the secondary relay node on the sub-channel according to the co-layer interference and cross-layer interference.
[0024] Determine the link rates between various types of nodes, including: between the IAB-donor and the master relay node; the access link rate constraint from the master relay node to the UE group it connects; the backhaul link rate between the secondary relay nodes directly associated with the master node; the access link rate from the secondary relay node to the UE group it connects.
[0025] Optimize the total network rate in the centralized scheduling period based on the signal-to-noise ratio constraint, link rate constraint, and power allocation constraint.
[0026] In some exemplary embodiments, the IAB-node calculates the resource utilization rate in the centralized scheduling period, including:
[0027] The IAB-node determines the actual utilization rate of the spectrum resources of the IAB-node in a distributed scheduling period according to the ratio of the spectrum resources allocated to the IAB-node by the IAB-donor to the spectrum resources actually used by the IAB-node.
[0028] Determine the resource utilization rate of the IAB-node in the centralized scheduling period according to the actual utilization rates of the spectrum resources of each IAB-node in multiple distributed scheduling periods.
[0029] In some exemplary embodiments, the IAB-node combines local information with the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor, and allocates spectrum and power for the child nodes and UEs associated with the IAB-node, including:
[0030] The IAB-node obtains local information corresponding to the distributed scheduling period;
[0031] Input the local information and the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor into the trained second resource allocation model, and obtain the spectrum and power allocation data of each associated child node and UE output by the model; wherein, the second resource allocation model is constructed based on the deep deterministic policy gradient algorithm.
[0032] In some exemplary embodiments, the IAB-donor allocates the resource scheduling information corresponding to the next centralized scheduling period for each IAB-node according to the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period, including;
[0033] Input the user satisfaction and resource utilization rate of the IAB-node in the current centralized scheduling period into the trained first resource allocation model, and obtain the resource scheduling information corresponding to the next centralized scheduling period of each IAB-node output by the model, so as to send the resource scheduling information to the corresponding IAB-node; wherein, the first resource allocation model is constructed based on the twin-delayed deep deterministic policy gradient algorithm.
[0034] According to a second aspect of the present invention, there is provided a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method for resource allocation of a multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning as described above is implemented.
[0035] The method for resource allocation of a multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning provided by the embodiments of the present invention, in a multi-hop IAB network environment, the IAB-donor can dynamically allocate the resource scheduling information corresponding to the next centralized scheduling period for each IAB-node by obtaining the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period. Thus, it can effectively cope with the changing traffic demands and link states in the network, reduce interference, optimize power control, and improve the overall network resource utilization rate and communication quality. And it can quickly adjust the spectrum and power resource allocation.
[0036] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0038] Figure 1 Schematic diagram showing the system structure of a two-hop IAB network according to an exemplary embodiment of the present invention;
[0039] Figure 2 Schematic diagram showing a resource allocation method for a multi-hop access-backhaul integrated network based on hierarchical multi-agent reinforcement learning according to an exemplary embodiment of the present invention;
[0040] Figure 3 Schematic diagram showing another resource allocation method for a multi-hop access-backhaul integrated network based on hierarchical multi-agent reinforcement learning according to an exemplary embodiment of the present invention;
[0041] Figure 4 Schematic diagram showing a hierarchical resource allocation task framework according to an exemplary embodiment of the present invention;
[0042] Figure 5 Schematic diagram showing a hierarchical multi-time scale resource allocation method framework according to an exemplary embodiment of the present invention;
[0043] Figure 6 Schematic diagram showing a high-level TD3 algorithm framework according to an exemplary embodiment of the present invention;
[0044] Figure 7 Schematic diagram showing a DDPG algorithm framework according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0046] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0047] In the related art, in a multi-hop IAB network environment, the network structure presents unique complexity and dynamics. With the continuous expansion of the network scale and the increasing richness of service types, the resource allocation problem it faces becomes more and more prominent. From the perspective of network architecture characteristics, in a multi-hop IAB network, the multi-hop transmission mode between IAB-nodes forms a complex mesh topology. The link states and traffic demands of different nodes change at all times, and traditional simple resource allocation methods are difficult to adapt to such a complex and changeable network environment. In such a situation, reasonable spectrum and power resource allocation becomes the key to ensuring the efficient operation of the network. The allocation of spectrum resources needs to minimize the interference between multiple nodes and at the same time meet the ever-changing traffic demands in the network. Power control is a key factor in ensuring the quality of multi-hop transmission links. The power output of each node needs to be adjusted according to its distance from other nodes, link quality, and signal interference situation. Appropriate power allocation helps to improve the reliability of signals, reduce interference, and save energy.
[0048] Most of the existing studies on resource allocation problems in the IAB scenario are based on traditional optimization methods. By modeling the optimization problem, a mixed-integer non-linear programming problem is obtained, which is non-convex. Then, the traditional optimization theory is used to allocate and iteratively solve the problem. However, this method usually requires collecting complete information of the network, and it is difficult to collect all this information in practice. As the network scale expands, traditional optimization methods often face problems of too high computational complexity and poor real-time performance.
[0049] In view of the shortcomings and deficiencies of the prior art, in this exemplary embodiment, a resource allocation method for a multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning is provided. In a multi-hop IAB network environment, it can effectively cope with the ever-changing traffic demands and link states in the network, reduce interference, optimize power control, and improve the overall network resource utilization rate and communication quality.
[0050] Next, the resource allocation method for a multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0051] In this exemplary embodiment, the resource allocation method for a multi-hop integrated access and backhaul network based on hierarchical multi-agent reinforcement learning can be applied to a multi-hop IAB (Integrated Access and Backhaul) network. For example, referring to Figure 1 the two-hop IAB network shown in the figure, in the downlink communication scenario, the IAB-node needs to serve multiple user devices. The UEs connected to the same node are regarded as a user group, and there is no direct access link from the UE to the IAB-donor in the network. All nodes are uniformly deployed within the coverage area of the IAB-donor and operate in the in-band full-duplex mode. There is still some residual self-interference after self-interference cancellation. In this IAB network, the IAB-nodes are divided into two categories. One category is the primary relay nodes, such as IAB-node1 and IAB-node2; the other category is the secondary relay nodes, such as IAB-node3 and IAB-node4, which further expand the communication coverage area of the IAB network through multi-hop relay connections to serve remote UEs. Denote the set of primary relay nodes as FN and the set of secondary relay nodes as SN. Assume there are I nodes in FN, FN = {fn1, fn2,... fn I}, and there are J nodes in SN, SN = {sn1, sn2,... sn J}. For each primary relay node, several secondary relay nodes can be connected to it through potential backhaul links. Each fn i connects to i c i secondary relay nodes, and the subset of the children nodes of fn is denoted as 1,i In addition, use u i to represent the connection between fn 2,j and its UE group. Similarly, the connection between the j-th secondary relay node and its UE group is denoted as u d0 For the IAB-donor, denoted as d0, its connection set is denoted as C I = {fn1, fn2,..., fn 1,i}. The complete set of UE groups in the network is U = U1 ∪ U2, where U1 = {u 2,j | i = 1, 2,..., I}, U2 = {u
[0052] Considering the centralized coordination ability of the IAB-donor and the distributed processing ability of the IAB-node, a hierarchical resource allocation scheme (Hierarchical Resource Allocation, HRA) is proposed, and HRA complies with the 3GPP standard. Referring to Figure 2As shown in the figure, a hierarchical resource allocation task framework is provided, which models resource allocation from long and short time granularities as high-level long-term tasks and low-level short-term tasks, and is managed by the IAB-donor and all IAB-nodes in the network respectively.
[0053] Specifically, time is divided into periodic centralized scheduling periods T, and each centralized scheduling period is further composed of multiple fixed-length distributed scheduling periods t. Each centralized scheduling period is managed by the IAB-donor, and each distributed scheduling period is managed by each IAB-node. The high-level task is that the IAB-donor collects the average user satisfaction and resource utilization rate within the centralized scheduling period fed back by all IAB-nodes at the beginning of the last frame of each centralized scheduling period, and determines the resource pool that each IAB-node can schedule in the next scheduling period from a global perspective. Before the start of the next scheduling period, the IAB-donor has enough time to calculate and transmit this communication scheduling decision. This centralized coordination mechanism can clarify the schedulable resource range of each IAB-node, thus effectively reducing the probability of resource conflicts. The low-level task is that after the start of a centralized scheduling period, within each distributed scheduling period, the IAB-node will allocate radio resources to its child nodes and associated UEs within the schedulable resource range delimited by the IAB-donor according to the local information it collects, and calculate the average user satisfaction and resource utilization rate within this period at the end of the centralized scheduling period, and upload it to the IAB-donor, thus starting the next stage of hierarchical resource management.
[0054] Exemplarily, for the IAB-donor, referring to Figure 2 As shown in the figure, a resource allocation method for a multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning, the specific method may include:
[0055] Step S11, the IAB-donor obtains the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period; wherein, the user satisfaction is used to describe the data transmission rate, and the resource utilization rate is used to describe the spectrum resource utilization efficiency;
[0056] Step S12, the IAB-donor allocates resource scheduling information corresponding to the next centralized scheduling period to each IAB-node according to the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period; wherein, the resource scheduling information includes the schedulable subchannels and the maximum transmit power of the subchannels.
[0057] Exemplarily, the method further includes: at the end of each centralized scheduling period, the IAB-node calculates the user satisfaction and resource utilization rate within the current centralized scheduling period and uploads them to the IAB-donor;
[0058] At the last frame of each centralized scheduling period, the IAB-donor obtains the user satisfaction and resource utilization rate of the IAB-node within the current centralized scheduling period.
[0059] Specifically, time is divided into periodic centralized scheduling periods T, and each centralized scheduling period is further composed of multiple distributed scheduling periods t with a fixed length. Each IAB-node in the network can calculate the user satisfaction and resource utilization rate of its own IAB-node within the current centralized scheduling period at the end of the centralized scheduling period and upload the data result to the IAB-donor.
[0060] After receiving the user satisfaction and resource utilization rate uploaded by each IAB-node, the IAB-donor can estimate and allocate the resource scheduling information of each IAB-node in a centralized scheduling period based on this data.
[0061] Alternatively, in some embodiments, for the IAB-donor, it can start the second resource allocation model to actively collect the average user satisfaction and resource utilization rate of all IAB-nodes within the centralized scheduling period at the last frame of each centralized scheduling period.
[0062] Exemplarily, the hierarchical multi-time scale resource allocation problem can be modeled first. Among them, for high-level tasks, it is defined as the centralized resource pre-allocation of the IAB-donor.
[0063] Specifically, before each centralized scheduling period starts, the IAB-donor determines the resource scheduling range for all IAB-nodes, that is, the schedulable sub-channels and the maximum transmit power.
[0064] The IAB network operates in an in-band mode. The total available spectrum resources in the network are divided into M available orthogonal sub-channels, denoted as M = {1, 2,..., M}. The sub-channel scheduling matrix is defined as S ∈ {0, 1} (I+J)×M : The first I rows correspond to the sub-channel scheduling vectors of I master relay nodes , where the element indicates that the master relay node fn i can schedule sub-channel m, otherwise it is 0. The last J rows correspond to the sub-channel scheduling vectors of J secondary relay nodes , where the element indicates that the master relay node sn jThe sub-channel m can be scheduled, otherwise it is 0. Similarly, define the power scheduling matrix P: where the element represents the maximum transmit power of the primary relay node fn i , and the element represents the maximum transmit power of the secondary relay node sn j .
[0065] Specifically, for low-layer tasks, it can be defined as the distributed dynamic resource optimization of the IAB-node based on local information.
[0066] Specifically, assume that there are N distributed scheduling cycles in a centralized scheduling cycle. In each distributed scheduling cycle, the IAB-node only needs to collect local information, and within the schedulable resource pool divided by the IAB-donor, allocate spectrum and power for its associated child nodes and UEs, and upload the relevant information within this cycle after the centralized scheduling cycle ends.
[0067] Based on the centralized resource pre-allocation of the IAB-donor, the primary relay node and the secondary relay node will allocate spectrum and power resources for their downlink links in each distributed scheduling cycle t. Specifically, for the primary relay node fn i , the maximum transmit power of fn i is , and the schedulable sub-channel range is determined by the sub-channel scheduling vector . Assume that the power allocated to the child node sn i,k is , the power allocated to the associated UE group u 1,i is , and it needs to satisfy the maximum transmit power constraint Define the sub-channel allocation matrix i of the primary relay node fn , where represents the sub-channel vector allocated by the primary relay node fn i to its child node sn i,k , represents the sub-channel vector allocated to the associated UE group u 1,i . If the element represents the allocation of this sub-channel m, otherwise it is 0. The allocated sub-channel must be schedulable, that is, only when , the sub-channel m can be allocated. Similarly, for the secondary relay node sn j , the maximum transmit power of sn j is , and the schedulable sub-channel range is determined by the sub-channel scheduling vector . Assume that the power allocated to the UE group u 2,j is , and it needs to satisfy the maximum transmit power constraint Define the secondary relay node sn j 's sub-channel allocation matrix , where represents the sub-channel vector allocated to the associated UE group u 2,j . If the element represents the allocation of this sub-channel m, otherwise it is 0. The allocated sub-channel must be schedulable, that is, only when , sub-channel m can be allocated. In the distributed scheduling period, the IAB-donor needs to act as an ordinary IAB-node at the same time to allocate sub-channels and power to its associated child node fn i . The power allocated to the child node fn i is , and it needs to satisfy The sub-channel allocation matrix is defined as , where the element represents the allocation of sub-channel m to the child node fn i , otherwise it is 0.
[0068] Exemplarily, the available spectrum resources of the IAB network include: M available orthogonal sub-channels; the orthogonal sub-channels include: the sub-channels schedulable by the primary relay node and the secondary relay node.
[0069] The method further includes:
[0070] Define the signal-to-noise ratio constraints of each type of node in the IAB network on the sub-channels, including: configuring the signal-to-noise ratio of the primary relay node on the sub-channels according to the co-layer interference, cross-layer interference, and residual self-interference; configuring the signal-to-noise ratio of the UE group connected to the primary relay node on the sub-channels according to the co-layer interference and cross-layer interference; configuring the signal-to-noise ratio of the UE group connected to the secondary relay node on the sub-channels according to the co-layer interference and cross-layer interference;
[0071] Determine the link rates between various types of nodes, including: between the IAB-donor and the primary relay node; the access link rate constraint from the primary relay node to the UE group it connects; the backhaul link rate between the secondary relay nodes associated under the primary node; the access link rate from the secondary relay node to the UE group it connects;
[0072] Optimize the total network rate during the centralized scheduling period based on the signal-to-noise ratio constraint, link rate constraint, and power allocation constraint.
[0073] Specifically, an optimization problem model can be defined.
[0074] Specifically, considering a centralized scheduling period, by jointly optimizing network spectrum allocation and power control, an optimization problem aiming to maximize the total network capacity is established for the two-hop IAB scenario under the premise of meeting the user QoS requirements. The channel gain between transmitter i and receiver j is defined as g i,j 。
[0075] In the t-th (t = 1, 2,..., N) distributed scheduling period, the main relay node fn i The SINR received on sub-channel m is denoted as , considering the possible co-layer, cross-layer, and residual self-interference, can be expressed as:
[0076]
[0077] where I inter is the cross-layer interference. The first term is generated by other main relay nodes transmitting data to their child nodes or associated UE groups using the same sub-channel m; the second term is generated when the secondary relay node transmits to its associated UE group; is the residual self-interference, represents the noise power.
[0078] The main relay node fn i The UE group u connected to it 1,i The SINR on sub-channel m is denoted as can be expressed as:
[0079]
[0080] where I intra is the co-layer interference, generated by other main relay nodes transmitting data to their child nodes or associated UE groups using the same sub-channel m; I inter is the cross-layer interference. The first term is generated by the IAB-donor transmitting data to other main relay nodes using the same sub-channel m, and the second term is generated by the secondary relay nodes that are not the child nodes of fn i transmitting data to their associated UE groups.
[0081] The main node fn i The secondary relay node sn connected below i,k The SINR received on sub-channel m is denoted as , considering co-layer interference, cross-layer interference, and self-interference, can be expressed as:
[0082]
[0083] The secondary relay node sn j The UE group u connected to it2,j The SINR on sub-channel m is denoted as Considering co-tier and cross-tier interference, it can be expressed as:
[0084]
[0085] where I intra is co-tier interference, which is generated by other secondary relay nodes transmitting data to their associated UE groups using the same sub-channel m; I inter is cross-tier interference. The first term is generated by the IAB-donor transmitting data to other primary relay nodes using the same sub-channel m, and the second term is generated when the primary relay node transmits data to its child nodes or associated UE groups.
[0086] Exemplarily, in the t-th distributed scheduling period, the rate of the backhaul link between the IAB-donor and the primary relay node fn i is:
[0087]
[0088] The access link rate from the primary relay node fn i to the UE group u 1,i it is connected to is:
[0089]
[0090] The backhaul link rate between the secondary relay node sn i associated with the primary node fn i,k is:
[0091]
[0092] The access link rate from the secondary relay node sn j to the UE group u 2,j it is connected to is:
[0093]
[0094] where B is the bandwidth of a single sub-channel. Then the instantaneous rates of u 1,i , u 2,j can be respectively expressed as:
[0095]
[0096] Based on the above analysis, the optimization objective is to maximize the total network rate within a centralized scheduling period while meeting the user QoS requirements. Therefore, the optimization problem can be expressed as:
[0097]
[0098] Among them, C1 and C2 ensure that within each distributed scheduling period, the rate of the UE group u i connected to the master relay node fn 1,i and the UE group u j connected to the secondary relay node sn 2,j is not lower than their respective set minimum rate requirements. C2, C3, and C4 are power constraints to ensure that the total allocated power does not exceed their respective maximum power limits. C5 is the sub-channel allocation constraint for the IAB-donor. C6, C7, and C8 are the sub-channel allocation constraints for the master relay node fn i to ensure that at most one object is allocated to the same sub-channel, and the allocated sub-channels must be within the schedulable range defined by the IAB-donor during resource pre-allocation. Similarly, C9 is the sub-channel allocation constraint for the secondary relay node sn j .
[0099] For example, referring to the architecture shown in Figure 4 , in the scheduling period T3, the IAB-donor allocates the sub-channels in the range of 0-5 to the IAB-node1 and stipulates that the maximum transmit power of the IAB-node1 is P1. Therefore, within each distributed scheduling period in T3, taking the distributed scheduling period t1 as an example, when the IAB-node1 allocates sub-channels to the link between its child node IAB-node3 and UE1, it can only select from the sub-channel range defined by the IAB-donor for the IAB-node1 at T3, and the sum of the transmit powers allocated to all connected links cannot exceed the maximum transmit power P1.
[0100] Exemplarily, due to the computational complexity of the above optimization problem and the combinatorial nature of long and short time slots, it is computationally expensive and difficult to solve this problem using traditional optimization methods, and there are limitations in dealing with dynamic changes in the network environment. Therefore, a reinforcement learning method is adopted to solve this problem. However, traditional reinforcement learning methods often face many challenges when dealing with complex problems. Due to the large scale of the state space and action space, and the relatively scarce reward signals, the training process becomes complex and the convergence speed is slow. To solve these problems, Hierarchical Reinforcement Learning (HRL) is proposed as an effective solution. HRL decomposes a complex overall problem into several relatively simple sub-problems through the conceptualization of states or time. Each sub-problem can be independently modeled by MDP, and the complex Markov decision process is optimized by decomposing it into smaller and more manageable sub-problems. First, HRL decomposes the problem into multiple levels, each level focusing on tasks at different time scales or different levels of abstraction, reducing the complexity of learning. The flexibility and adaptability of HRL are also reflected in allowing different policies to be learned at different levels, and these policies can be quickly adjusted according to changes in the environment. The low-level policy can quickly respond to immediate changes in the environment, while the high-level policy can be slowly adjusted to adapt to long-term goals.
[0101] Therefore, for the proposed resource allocation optimization problem on different time scales, a hierarchical multi-agent reinforcement learning method is proposed to analyze and solve this problem. This method divides resource allocation into high-level and low-level problems on the time scale to meet the needs of different decision-making levels and time granularities.
[0102] Specifically, referring to Figure 5As shown in the figure, high-level problems focus on macro-decisions with a long time scale. The IAB-donor acts as a high-level agent, responsible for allocating schedulable spectrum and power resources to each IAB-node according to network status and policy selection. Low-level problems correspond to micro-operations with a short time scale. Each IAB-node acts as a low-level agent. Based on the resource division completed by the IAB-donor, it further allocates spectrum and power resources to its child nodes and associated UE groups, and generates rewards on a short time scale. After the low-level agent completes resource allocation, the IAB-node collects key performance indicators such as user satisfaction and resource utilization rate within a short-term scheduling cycle, integrates this information, and feeds it back to the high-level agent, the IAB-donor. The high-level agent uses this feedback information for learning and policy optimization, and then adjusts the resource pre-allocation decision for the next cycle. Through this closed-loop feedback mechanism, a close collaborative relationship is formed between the high-level and low-level agents. High-level decisions guide low-level operations, and the results of low-level operations in turn affect the policy optimization of the high-level, so as to achieve the collaborative optimization of resource allocation on long and short time scales, thereby improving the overall network performance.
[0103] Exemplarily, the IAB-node calculates the user satisfaction within a centralized scheduling cycle, including:
[0104] Step S21, within a distributed scheduling cycle, determine the satisfaction of the UE group in the distributed scheduling cycle according to the ratio of the average data transmission rate of the UE group associated with the IAB-node to the minimum data transmission rate threshold set by the basic service requirements of the user;
[0105] Step S22, determine the user satisfaction of the UE group associated with the IAB-node within a centralized scheduling cycle according to the satisfaction of the UE groups in multiple distributed scheduling cycles.
[0106] Exemplarily, the IAB-node calculates the resource utilization rate within a centralized scheduling cycle, including:
[0107] Step S31, the IAB-node determines the actual utilization rate of the spectrum resources of the IAB-node within a distributed scheduling cycle according to the ratio of the spectrum resources allocated to the IAB-node by the IAB-donor to the spectrum resources actually used by the IAB-node;
[0108] Step S32, determine the resource utilization rate of the IAB-node within a centralized scheduling cycle according to the actual utilization rate of the spectrum resources of each IAB-node in multiple distributed scheduling cycles.
[0109] Specifically, in order to quantify the network performance and clarify the information that the IAB-node needs to feedback to the IAB-donor, this section defines the evaluation metrics of user satisfaction and resource utilization. Let T be the time of a centralized scheduling period, t be the time of a distributed scheduling period, and it is set that there are N distributed scheduling periods in a centralized scheduling period.
[0110] Specifically, in a distributed scheduling period t, the satisfaction of the user group under the IAB-node is defined as:
[0111]
[0112] where represents the average data transmission rate obtained by the user group in this scheduling period, and R min is the minimum data transmission rate threshold set to ensure the basic service requirements of users. By the ratio of the actual rate to the minimum required rate, it reflects the satisfaction degree of the user group with the current network service. The closer the ratio is to or greater than 1, the higher the satisfaction degree.
[0113] Define the average satisfaction of the user group under the IAB-node in a centralized scheduling period T as:
[0114]
[0115] Comprehensively considering the change of user satisfaction in the entire centralized scheduling period, it reflects the average satisfaction degree of the user group with the network service provided by the IAB-node in a relatively long time period.
[0116] Define the network user satisfaction as:
[0117]
[0118] where I + J is the number of user groups in the network, and R min is the minimum data transmission rate threshold, and φ(T) measures the overall satisfaction of all users in the network in a centralized scheduling period.
[0119] Specifically, in a distributed scheduling period t, the resource utilization rate of the IAB-node is defined as:
[0120]
[0121] where B allocated represents the spectrum resources allocated by the IAB-donor to the IAB-node, It represents the spectrum resources actually used by the IAB-node, and reflects the actual utilization of spectrum resources by the IAB-node through the ratio of their quantities.
[0122] Define that in a centralized scheduling period T, the average resource utilization rate of the IAB-node is:
[0123]
[0124] It measures the average utilization efficiency of the IAB-node for spectrum resources in a centralized scheduling period as a whole.
[0125] Define the network resource utilization rate as:
[0126]
[0127] Among them, It is an index used to measure the average utilization efficiency of the entire network for spectrum resources in the centralized scheduling period T.
[0128] After each centralized scheduling period T ends, the IAB-node needs to feedback the following information to the IAB-donor: 1. The average satisfaction degree of the users associated with the IAB-node, so that the IAB-donor can understand the overall satisfaction of the users served by this IAB-node. 2. The average resource utilization rate of the IAB-node, which helps the IAB-donor master the usage efficiency of the allocated resources by the IAB-node.
[0129] In addition, in some exemplary embodiments, a service cycle can be defined. This service cycle can include a centralized scheduling period, or multiple consecutive centralized scheduling periods. It can be the IAB-donor that statistics the network user satisfaction degree and / or the network resource utilization rate within this service cycle. If the corresponding parameters meet the preset standard thresholds, the current state can be maintained. Or, if one or more parameters within the current service cycle are lower than the preset standard thresholds, it means that the current resource allocation method is unreasonable and cannot meet the actual requirements of the UE for network resource allocation; at this time, an optimization and update task for the first resource allocation model deployed on the IAB-donor side and / or the second resource allocation model deployed on the IAB-node side can be triggered to achieve the optimization of the resource allocation model.
[0130] Exemplarily, the reinforcement learning elements of high-level problems can be defined as <S h ,A h ,R h ,P h>, where the IAB-donor, as an agent in high-level problem reinforcement learning, has the ability to macroscopically control the entire multi-hop IAB network. It can collect network status information uploaded from each IAB-node and make centralized resource pre-allocation decisions based on this information to achieve overall network resource planning and optimize the overall network performance.
[0131] Define the state space as S h , at decision-making time T, the IAB-donor needs to obtain the following state information from the network environment: a. The average user satisfaction value of all IAB nodes in the previous centralized scheduling period T - 1 b. The average resource utilization rate of all IAB nodes in the previous centralized scheduling period T - 1 Therefore, the state space of the IAB-donor at time T is:
[0132]
[0133] Define the action space as A h , the IAB-donor delimits the schedulable sub-channel range and maximum transmit power value for the i-th IAB-node:
[0134] Starting position, sub-channel starting position Represents the number of schedulable sub-channels allocated in a certain proportion, sub-channel number According to and , a continuous sub-channel sequence can be delimited and allocated to the IAB-node for scheduling. Represents the maximum transmit power ratio value delimited for the i-th IAB-node, which is the maximum transmit power of the IAB-node in this centralized scheduling period. Through 's collaborative calculation, the action space completes the accurate mapping of the sub-channel allocation range (starting position, number) and power value, providing specific parameters for the resource scheduling of the IAB-node.
[0136] The action space of the IAB-donor is to delimit the schedulable sub-channel range and maximum transmit power value for all IAB-nodes, that is:
[0137]
[0138] Define the reward function, which can be R h Represents the reward set feedback by the environment. The reward function of the IAB-donor is formulated by comprehensively considering factors such as the network user satisfaction index, network resource utilization index, and network total rate:
[0139]
[0140] Among them, w1, w2, and w3 are weight coefficients used to adjust the relative importance of each factor in the reward function. This reward function indicates that when the IAB-donor allocates appropriate schedulable resources to the IAB-node, if it can improve the network user satisfaction, network resource utilization rate, and network total rate, it can obtain a higher reward value, thereby motivating the IAB-donor to make a better resource allocation decision.
[0141] Define P h as the state transition function, which represents the probability that the agent enters the next state after taking a certain action in a certain state.
[0142] Exemplarily, the IAB-donor allocates resource scheduling information corresponding to the next centralized scheduling period for each IAB-node according to the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period, including:
[0143] Input the user satisfaction and resource utilization rate of the IAB-node in the current centralized scheduling period into the trained first resource allocation model, and obtain the resource scheduling information corresponding to each IAB-node in the next centralized scheduling period output by the model, so as to send the resource scheduling information to the corresponding IAB-node; among them, the first resource allocation model is constructed based on the twin-delayed deep deterministic policy gradient algorithm.
[0144] Specifically, the high-level agent IAB-donor uses the twin-delayed deep deterministic policy gradient algorithm (TD3) to construct the first resource allocation model in the decision-making of resource pre-allocation in the centralized scheduling period. The TD3 algorithm can alleviate the overestimation problem by taking the minimum value of the action value through the twin Critic networks, and because of the use of the delayed policy update method, it can enhance the training problem of the agent in the complex and changeable environment, which is beneficial for the IAB-donor to learn effective strategies. At the same time, in the face of the mixed action space of discrete sub-channels and continuous power values, it can effectively handle the continuous power allocation and map the continuous values to the appropriate decision space through reasonable action mapping to achieve the optimal selection of discrete sub-channels.
[0145] Exemplarily, the first resource allocation model is constructed based on the TD3 algorithm. Refer to Figure 6 As shown, when calculating the state-action value function, the TD3 algorithm introduces two sets of networks. When defining the Critic network, the TD3 algorithm defines two networks Q1 and Q2 with the same structure. First, randomly initialize the local Actor network π θ , two local Critic networks and , while creating the corresponding target Actor network , target Critic network and , set up the experience replay pool D h and discount factor γ h and other hyperparameters. At the beginning of each round of training, reset the environment. At each time step T, the IAB-donor generates an action based on the current state , through the local Actor network , add noise n T and clip it within [-ε, ε] to obtain the actually executed action The agent executes the action to interact with the environment and obtains the next state and reward Then store the quadruple in the experience replay pool D h . To make more effective use of the data in the experience replay pool, a prioritized experience replay mechanism is introduced. When storing the quadruple in the experience replay pool D h , assign a priority to it according to the potential importance of this experience for the agent's learning. When there is enough experience in the pool, instead of using random sampling, sample a batch of joint experiences of size N h from the experience replay pool D h for training. The update of the Critic network is achieved by calculating the outputs of two target networks and
[0146]
[0147] Calculate the loss functions of the two local Critic networks respectively:
[0148]
[0149] Then, update the parameters θ1 and θ2 of the local Critic network by gradient descent method.
[0150] The TD3 algorithm proposes a delayed policy update method to reduce the update frequency of the Actor network. After the Critic network is updated multiple times, the Actor network is updated. Such an update method can ensure the relative stability of value estimation, thereby improving the algorithm performance. Therefore, every k steps, calculate the loss function of the local Actor network and update the parameters θ θ of the Actor network π by gradient ascent. The update of the target network is achieved by soft update, regularly copying parameters from the local network.
[0151] Exemplarily, refer to Figure 3 As shown, the method further includes:
[0152] Step S13, within each distributed scheduling period, the IAB-node combines local information with the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor, and allocates spectrum and power for the child nodes and UEs associated with this IAB-node, so as to achieve distributed dynamic resource optimization of the IAB-node based on local information.
[0153] Exemplarily, the local information of the IAB-node includes at least one of the following:
[0154] The configuration results of the spectrum resources and power resources of the child nodes and UEs by the IAB-node in the previous distributed scheduling period;
[0155] The channel gains between the IAB-node and the child nodes and associated UEs within the current distributed scheduling period;
[0156] The channel interference between the IAB-node and the child nodes and associated UEs in the previous distributed scheduling period; wherein, the channel interference includes at least one of co-layer interference, cross-layer interference, and self-interference generated due to the use of the same sub-channel.
[0157] Exemplarily, the IAB-node combines local information with the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor, and allocates spectrum and power for the child nodes and UEs associated with this IAB-node, including:
[0158] The IAB-node obtains the local information corresponding to the distributed scheduling period;
[0159] Input the local information and the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor into the trained second resource allocation model, and obtain the spectrum and power allocation data of each associated child node and UE output by the model; wherein, the second resource allocation model is constructed based on the deep deterministic policy gradient algorithm.
[0160] Specifically, for the IAB-node, it can collect local information. After receiving the resource scheduling information allocated by the IAB-donor for it, it can use the local information and the received resource scheduling information as input parameters, and input them into the second resource allocation model to obtain the spectrum and power allocation results of each associated child node and UE output by the model.
[0161] Specifically, refer to Figure 5As shown, the IAB-node can be used to allocate spectrum and power resources for its child nodes and UE groups within each distributed scheduling period. Define the elements of reinforcement learning for the low-level problem <S l , A l , R l , P l >. In the resource allocation scenario of the distributed scheduling period, each IAB-node acts as an independent agent, and each agent has the ability to make autonomous decisions. Based on local real-time information, it can allocate spectrum and power resources for its associated child nodes and UE groups within the schedulable resource pool divided by the IAB-donor. Each IAB-node continuously senses changes in the network state, executes corresponding resource allocation actions, and adjusts its own strategy according to the rewards feedback by the environment.
[0162] Specifically, define the space state S l , at each t time slot (i.e., the distributed scheduling period), each agent n needs to obtain the following state information from the network environment: a. The schedulable spectrum resources delimited by the IAB-donor for it b. The maximum transmit power value delimited by the IAB-donor for it c. The spectrum resources allocated by agent n to its child nodes and UEs at t - 1 d. The power resources allocated by agent n to its child nodes and UE groups at t - 1 e. The channel gain between agent n and its child nodes and associated UE groups f. The interference received by the child nodes and UE groups at t - 1 Therefore, at time t, the state space of agent n is:
[0163]
[0164] Define the action space A of agent n l To allocate sub-channels and power resources to child nodes and associated UE groups, assume the agent has c n child nodes. Then the action space of agent n is:
[0165]
[0166] Among them, represents the proportion of the number of sub-channels allocated by agent n to the i-th child node and UE group within the schedulable sub-channel range delimited by the IAB-donor for it, represents the proportion of the power value allocated by agent n to the i-th child node and UE group.
[0167] Define R lThe reward set representing environmental feedback; the relationship between the low-level agents is a fully cooperative one. Therefore, the local reward of each agent is set the same as the low-level global reward. When setting the reward, two factors are considered: the total network rate and the user QoS requirements.
[0168]
[0169] Penalty term Set as:
[0170]
[0171] Among them, the value of λ will be determined through simulation experiments.
[0172] Define P l As the state transition function, it represents the probability that the agent enters the next state after taking a certain action in a certain state.
[0173] Exemplarily, the multi-agent deep deterministic policy gradient (MADDPG) algorithm of the second resource allocation model deployed by IAB-node is constructed to solve the spectrum and power allocation problems during the distributed scheduling period, which can effectively adapt to the problem of the explosion of the action space dimension caused by the increase in the number of factor nodes and associated UEs. Compared with the deep Q-learning algorithm, it is more suitable for the characteristics of the action space.
[0174] Specifically, the low-level problem is a distributed dynamic resource optimization problem, which contains multiple agents. Each IAB-node is regarded as an independent agent. In this architecture, each agent needs to find a balance between cooperation and competition. In addition, the high-level agent will comprehensively consider the dynamic situation of each IAB-node based on the global information, thus balancing the resource competition process between them to a certain extent and indirectly providing a global perspective for the low-level agents. Therefore, we choose a distributed training method to deploy the low-level agents. As Figure 7 shown, in this process, each agent is relatively independent during the training process and makes decisions based on local information. Here, we give an overview of the training process of the low-level agents.
[0175] Each low-level agent IAB-node n randomly initializes its local Actor network Critic network and the target Actor and Critic networks At the same time, each agent constructs its own independent experience replay pool D l,n , and determines hyperparameters such as the discount factor γ l etc. At the beginning of each round of training, the environment is reset. For each time step t, each IAB agent n makes a decision based on its own current state Generate actions by the local Actor network Generate actions All agents execute actions simultaneously Interact with the environment to obtain the next state Global reward r l (t) Then, agent n stores the experience containing its own state, action, reward, and next state information in its own experience replay pool D l,n When the amount of experience in the experience pool is sufficient, randomly sample a batch of experience of size N l,n from D l for training. For each agent n, calculate the output of the local Critic network:
[0176]
[0177] Calculate the loss function of the local Critic network:
[0178]
[0179] where i represents the i-th data in the batch of experience randomly sampled from the experience replay pool. Then, train the local Critic network by minimizing the loss function and update its parameters using the gradient descent method. The Actor network updates the gradient of the policy function through the Critic network and updates the parameters of the local Actor network using the gradient ascent method. After updating the parameters of the local Actor network and Critic network, use the soft update strategy to update the parameters of the target Actor and Critic networks.
[0180] Exemplarily, the process of the resource allocation method based on hierarchical multi-agent reinforcement learning may include the following steps:
[0181]
[0182]
[0183] Exemplarily, to verify the correctness of the proposed method, compare the hierarchical resource allocation algorithm with the centralized resource allocation algorithm and distributed resource allocation algorithm based on reinforcement learning.
[0184] The simulation results compared with the centralized and distributed algorithms show that under the same scenario and channel state information conditions, the hierarchical resource allocation algorithm adopted can reach 96.24% of the rate of the centralized algorithm after convergence, but the signaling overhead of the centralized scheme is more than 3 times that of it. The rate of the compared distributed resource allocation algorithm is 15.3% worse than the proposed scheme, and the signaling overhead spent is not much different, meeting the expected goal.
[0185] The method provided by the embodiments of the present invention conforms to the characteristics of the IAB architecture. In the two-hop IAB scenario, it can perform resource pre-allocation for the IAB-donor on a long time scale, allocate a schedulable spectrum and power resource pool for all IAB-nodes, and the IAB-nodes on a short time scale further allocate spectrum and power resources for their child nodes and associated UEs on the allocated resource pool. In addition, corresponding states, actions, rewards, and penalties are designed for high-level problems and low-level problems respectively. Based on the hierarchical multi-agent algorithm framework of the TD3 algorithm and the MADDPG algorithm, it jointly solves the spectrum allocation and power control problems in the multi-hop scenario.
[0186] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0187] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0188] The embodiments of the present invention include a computer program product, which includes a computer program carried on a storage medium, and the computer program contains program codes for executing the method shown in the flowchart.
[0189] It should be noted that the storage medium shown in the embodiments of the present invention may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any storage medium other than a computer-readable storage medium, and this storage medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function.
[0191] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation to the unit itself.
[0192] In one embodiment, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps in the above method embodiments.
[0193] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A resource allocation method for a multi-hop access backhaul integrated network based on hierarchical multi-agent reinforcement learning, characterized in that: The method includes: IAB-donor obtains the user satisfaction and resource utilization of each IAB-node in the current centralized scheduling cycle; wherein the user satisfaction is used to describe the data transmission rate, and the resource utilization is used to describe the spectrum resource utilization efficiency; IAB-donor allocates resource scheduling information corresponding to the next centralized scheduling period to each IAB-node according to the user satisfaction and resource utilization of each IAB-node in the current centralized scheduling period; wherein the resource scheduling information includes schedulable sub-channels and the maximum transmission power of the sub-channels.
2. The method according to claim 1, characterized in that The centralized scheduling period includes N distributed scheduling periods; the method further includes: In each distributed scheduling period, the IAB-node combines local information with the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor to allocate spectrum and power to the child nodes and UEs associated with the IAB-node, thereby realizing distributed dynamic resource optimization of the IAB-node based on local information.
3. The method according to claim 2, characterized in that The local information of the IAB-node includes at least one of the following: The spectrum resource and power resource configuration results of the IAB-node for the sub-nodes and UEs in the previous distributed scheduling cycle; Channel gains between the IAB-node and its child nodes and associated UEs in the current distributed scheduling period; Channel interference between the IAB-node and the sub-node and the associated UE in the previous distributed scheduling period; wherein the channel interference includes at least one of the same-layer interference, cross-layer interference, and self-interference caused by using the same sub-channel.
4. The method according to claim 1, characterized in that The method further comprises: At the end of each centralized scheduling cycle, IAB-node calculates the user satisfaction and resource utilization rate during the centralized scheduling cycle and uploads them to IAB-donor; In the last frame of each centralized scheduling cycle, IAB-donor obtains the user satisfaction and resource utilization of IAB-node in the current centralized scheduling cycle.
5. The method according to claim 1 or 4, characterized in that: IAB-node calculates user satisfaction during the centralized scheduling period, including: In a distributed scheduling period, the satisfaction level of the UE group in the distributed scheduling period is determined according to the ratio of the average data transmission rate of the UE group associated with the IAB-node and the minimum data transmission rate threshold set by the user's basic service requirements; According to the satisfaction of the UE groups in multiple distributed scheduling periods, the user satisfaction of the UE group associated with the IAB-node in one centralized scheduling period is determined.
6. The method according to claims 1-4, characterized in that: IAB-node calculates resource utilization during the centralized scheduling cycle, including: The IAB-node determines the actual spectrum resource utilization rate of the IAB-node in a distributed scheduling cycle according to the ratio of the spectrum resources allocated to the IAB-node by the IAB-donor and the spectrum resources actually used by the IAB-node; According to actual spectrum resource utilization of each IAB-node in a plurality of distributed scheduling periods, the resource utilization of the IAB-node in the centralized scheduling period is determined.
7. The method according to claim 1, characterized in that The available spectrum resources of the IAB network include: M available orthogonal sub-channels; the orthogonal sub-channels include: sub-channels that can be scheduled by the primary relay node and the secondary relay node; The method further comprises: Define the signal-to-noise ratio constraints of various types of nodes in the IAB network on the sub-channel, including: configuring the signal-to-noise ratio of the primary relay node on the sub-channel according to the same-layer interference, cross-layer interference, and residual self-interference; configuring the signal-to-noise ratio of the UE group connected to the primary relay node on the sub-channel according to the same-layer interference and cross-layer interference; configuring the signal-to-noise ratio of the UE group connected to the secondary relay node on the sub-channel according to the same-layer interference and cross-layer interference; Determine the link rate between each type of node, including: between the IAB-donor and the primary relay node; the access link rate constraint from the primary relay node to the UE group connected to it; the direct backhaul link rate of the secondary relay node associated with the primary node; the access link rate from the secondary relay node to the UE group connected to it; Based on signal-to-noise ratio constraints, link rate constraints, and power allocation constraints, the total network rate within the centralized scheduling period is optimized.
8. The method according to claim 2, characterized in that: The IAB-node combines the local information with the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor to allocate spectrum and power for the child nodes and UEs associated with the IAB-node, including: IAB-node obtains local information corresponding to the distributed scheduling cycle; The local information and the resource scheduling information corresponding to the next centralized scheduling period configured by the IAB-donor are input into the trained second resource allocation model to obtain the spectrum and power allocation data of each associated sub-node and UE output by the model; wherein the second resource allocation model is constructed based on a deep deterministic policy gradient algorithm.
9. The method according to claim 1, characterized in that: The IAB-donor allocates resource scheduling information corresponding to the next centralized scheduling period to each IAB-node according to the user satisfaction and resource utilization rate of each IAB-node in the current centralized scheduling period, including: The user satisfaction and resource utilization of the IAB-node in the current centralized scheduling cycle are input into the trained first resource allocation model, and the resource scheduling information corresponding to each IAB-node in the next centralized scheduling cycle output by the model is obtained, so as to send the resource scheduling information to the corresponding IAB-node; wherein, the first resource allocation model is constructed based on the double-delay deep deterministic policy gradient algorithm.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for resource allocation of a multi-hop access and backhaul integrated network based on hierarchical multi-agent reinforcement learning as described in any one of claims 1 to 9 is implemented.