Ultra-dense Networking Intelligent Resource Optimization Method, Device, Equipment and Storage Medium

By using NOMA-based IAB network model and heterogeneous deep network in millimeter wave IAB network for reinforcement learning decisions, the resource allocation problem in large-scale scenarios is solved, efficient and dynamic spectrum and power allocation are achieved, and the performance and resource utilization of the communication network are improved.

CN120034868BActive Publication Date: 2025-06-20EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510503135.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-20
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The resource allocation problem in millimeter wave IAB network, especially in large-scale scenarios, traditional optimization methods are insufficient in efficiency and are difficult to meet the real-time and flexibility requirements of dynamic resource allocation.

Method used

The IAB network model based on NOMA is adopted, combined with heterogeneous deep networks to make reinforcement learning decisions, and the decision model is constructed through agents, states, actions and reward functions, and the channel state information is responded in real time, the maximum Q value and online decision results are generated, and the experience playback pool and policy network are updated to optimize the resource allocation strategy.

Benefits of technology

It realizes efficient and dynamic spectrum and power distribution, improves the performance and resource utilization of the next generation of communication networks, and solves the real-time and flexibility requirements in resource allocation problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034868B_ABST
    Figure CN120034868B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, equipment and storage medium for ultra-dense networking intelligent resource optimization, belonging to the field of wireless communication network resource allocation. It includes: constructing an IAB network model based on NOMA technology. Then, establishing a communication model of the IAB network and determining a constraint function on this basis. Subsequently, designing a reinforcement learning decision model, whose reward function is determined according to the constraint function. Generating decision results in real time in response to channel state information through a heterogeneous deep network and storing them in an experience replay pool for subsequent use. The policy network and the dueling network work together to improve the quality of decisions, and based on the calculated maximum Q value and the corresponding loss function, network update is achieved to obtain a target heterogeneous deep network. Finally, an optimal resource allocation strategy is formulated for the IAB network. The present application can achieve efficient and dynamic spectrum and power allocation, and improve the performance and resource utilization rate of the next-generation communication network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of wireless communication network resource allocation, and particularly relates to an ultra-dense networking intelligent resource optimization method, device, equipment, and storage medium. Background Art

[0002] With the rapid development of 5G and next-generation communication technologies, the number of access terminals in the network has increased explosively, and the user traffic demand has increased significantly. To address the network capacity bottleneck, ultra-dense networking has become a key solution, but its deployment cost is high, especially the trenching and laying of fiber-optic backhaul links and equipment installation, which lead to complex implementation. For this reason, 3GPP has proposed an Integrated Access and Backhaul (IAB) network architecture to relay traffic through wireless backhaul links, reduce deployment costs, and simplify the architecture. However, the resource competition of wireless backhaul links has exacerbated the spectrum interference problem, posing a challenge to system design.

[0003] Millimeter-wave technology is regarded as the core solution to this problem. Its frequency band provides far more spectrum resources than Sub-6 GHz, but it has high propagation loss and limited coverage, and is mainly suitable for short-distance high-speed transmission. To make up for this defect, multi-antenna technology is combined to achieve directional transmission to improve link efficiency, and Non-Orthogonal Multiple Access (NOMA) technology is introduced to further improve spectrum utilization. NOMA allows the base station to provide concurrent transmissions with differential power allocation for multiple users in the same sub-band, optimizing resource utilization. However, the combination of millimeter-wave IAB networks and NOMA still faces resource allocation problems: the backhaul and access links need to dynamically allocate spectrum sub-bands and power resources to avoid communication quality degradation, uneven resource occupancy, and multi-service demand conflicts caused by sharing conflicts. Precise sub-band allocation and power control are the keys to improving system performance, but traditional optimization methods are inefficient in large-scale scenarios.

[0004] In recent years, deep reinforcement learning has become a potential solution due to its adaptive decision-making ability. Reinforcement learning learns the optimal strategy through the interaction between the agent and the environment, and deep reinforcement learning combines deep learning to handle complex state spaces. Although classical learning algorithms are suitable for discrete action spaces, they are difficult to handle large-scale continuous action scenarios; although continuous control algorithms can handle continuous actions such as power allocation, they lack support for discrete sub-band allocation. In a mixed decision scenario, existing deep reinforcement learning algorithms have low decision-making efficiency due to mismatched action spaces, and it is difficult to meet the real-time and flexibility requirements of dynamic resource allocation in millimeter-wave IAB networks. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method, device, equipment, and storage medium for intelligent resource optimization in ultra-dense networking to address the problem of resource allocation in a hybrid action space, thereby achieving efficient and dynamic spectrum and power allocation and improving the performance and resource utilization rate of next-generation communication networks.

[0006] To solve the above technical problems, this application is implemented as follows:

[0007] In a first aspect, the embodiments of this application provide a method for intelligent resource optimization in ultra-dense networking, the method including:

[0008] Based on NOMA, establish an IAB network including a macro base station, nodes, and users; wherein, the IAB network includes a backhaul link in the millimeter-wave band and an access link applying NOMA technology, and the nodes include high-hop nodes and low-hop nodes with different distances from the macro base station;

[0009] Establish a communication model of the IAB network, and based on the communication model, determine a constraint function with the optimization goal of maximizing the throughput of the access users;

[0010] Construct a reinforcement learning decision model, including an agent, state, action, and reward function; wherein, the reward function is set based on the constraint function;

[0011] Based on the reinforcement learning decision model and the constraint function, in real-time respond to the channel state information through a heterogeneous deep network to generate the maximum Q value and an online decision result, and store them in an experience replay pool; wherein, the heterogeneous deep network includes a policy network and a dueling network;

[0012] Based on the maximum Q value and the loss function, update the policy network and the dueling network to obtain a target heterogeneous deep network;

[0013] Based on the target heterogeneous deep network, generate an optimal resource allocation strategy for the IAB network.

[0014] Preferably, the specific steps of establishing an IAB network including a macro base station, nodes, and users based on NOMA include:

[0015] Divide the nodes into high-hop nodes and low-hop nodes according to the distance from the macro base station;

[0016] Connect the high-hop nodes and the low-hop nodes using the millimeter-wave band.

[0017] Preferably, the specific steps of establishing a communication model of the IAB network and, based on the communication model, determining a constraint function with the optimization goal of maximizing the throughput of the access users include:

[0018] Define a channel parameter model; the channel parameters at least include the channel gains, signal-to-noise ratios, transmission powers, and transmission capacities of the backhaul link and the access link;

[0019] Taking the maximization of the throughput of the accessing users as the optimization goal, determine the network reach limit based on the channel parameter model;

[0020] Determine the constraint function based on the network reach limit.

[0021] Preferably, the reward function is:

[0022] ,

[0023] where r(t) represents the reward value, t represents the current time slot, represents the sum of the total rates of the users associated with the nth high-hop node under the constraint conditions, and N represents the total number of high-hop nodes, represents the sum of the total rates of the users associated with the mth low-hop node under the constraint conditions, and M represents the total number of low-hop nodes.

[0024] Preferably, based on the reinforcement learning decision model and the constraint function, the specific steps of generating the maximum Q value and the online decision result by the heterogeneous deep network in real-time response to the channel state information and storing them in the experience replay pool include:

[0025] The agent obtains the current state of the access link;

[0026] The heterogeneous deep network outputs the online action and the maximum Q value according to the current state; among them, the online action includes a continuous action set and a discrete action set;

[0027] After the online action is executed, obtain the reward value and the subsequent state of the access link;

[0028] Store the online decision result composed of the current state, the continuous action set, the discrete action set and the subsequent state in the experience replay pool.

[0029] Preferably, the specific steps of the heterogeneous deep network outputting the online action and the maximum Q value according to the current state include:

[0030] Input the current state into the policy network to output the continuous action set;

[0031] Input the current state and the continuous action set into the duel network to determine the maximum Q value;

[0032] Determine the discrete action set corresponding to the maximum Q value through the fully greedy strategy.

[0033] Preferably, the specific steps of updating the policy network and the duel network based on the maximum Q value and the loss function to obtain the target heterogeneous deep network include:

[0034] Update the policy network based on the gradient of the maximum Q value to obtain the latest policy network;

[0035] The confrontation network calculates the predicted Q value through the value function and the action optimization function;

[0036] Update the confrontation network based on the predicted Q value and the loss function to obtain the latest confrontation network, so as to comprehensively form the target heterogeneous deep network.

[0037] Compared with the prior art, the above technical solution provided by this application has at least the following beneficial effects:

[0038] This application first constructs an IAB network model including a macro base station, nodes, and users based on NOMA technology. In this network, the backhaul link uses the millimeter wave band to utilize its large bandwidth advantage, and the access link applies NOMA technology to improve the spectrum efficiency. The nodes in the network are divided into high-hop nodes and low-hop nodes to adapt to different transmission requirements. Then, a communication model of the IAB network is established, and based on this, a constraint function is determined to improve the system performance. Subsequently, a reinforcement learning decision model is designed, which consists of an agent, a state, an action, and a reward function, where the reward function is determined according to the set performance metrics. Under this framework, the maximum Q value is calculated and the online decision result is generated in real time through the heterogeneous deep network, and the result will be stored in the experience replay pool for subsequent use. The heterogeneous deep network includes a policy network and a confrontation network, which work together to improve the quality of decision-making, and based on the calculated maximum Q value and the corresponding loss function, the network is updated to obtain the optimized target heterogeneous deep network. Finally, based on the target heterogeneous deep network, an optimal resource allocation strategy for the IAB network is implemented.

[0039] In a second aspect, an ultra-dense networking intelligent resource optimization device provided by an embodiment of this application includes:

[0040] An IAB network module for establishing an IAB network including a macro base station, nodes, and users based on NOMA; wherein, the IAB network includes a backhaul link in the millimeter wave band and an access link applying NOMA technology, and the nodes include high-hop nodes and low-hop nodes with different distances from the macro base station;

[0041] A constraint function module for establishing a communication model of the IAB network and determining a constraint function with the optimization goal of maximizing the throughput of the access users based on the communication model;

[0042] A reinforcement learning module for constructing a reinforcement learning decision model, including an agent, a state, an action, and a reward function; wherein, the reward function is set based on the constraint function;

[0043] A decision-making module, configured to generate a maximum Q value and an online decision result in real time in response to channel state information through a heterogeneous deep network based on a reinforcement learning decision-making model and a constraint function, and store them in an experience replay pool; wherein, the heterogeneous deep network includes a policy network and a dueling network;

[0044] A network update module, configured to update the policy network and the dueling network based on the maximum Q value and a loss function to obtain a target heterogeneous deep network;

[0045] A resource allocation module, configured to generate an optimal resource allocation strategy for the IAB network based on the target heterogeneous deep network.

[0046] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0047] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0048] It can be understood that the beneficial effects of the technical solutions provided in the above second aspect, third aspect, and fourth aspect can refer to the relevant descriptions in the first aspect, and will not be repeated here.

[0049] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0051] Figure 1 is a schematic flowchart of a super-dense networking intelligent resource optimization method provided by some embodiments of the present application;

[0052] Figure 2 is a schematic structural diagram of a heterogeneous deep network update process shown by some embodiments of the present application;

[0053] Figure 3 is a schematic diagram of heterogeneous network algorithm iteration shown by some embodiments of the present application;

[0054] Figure 4 is a schematic diagram of the average rate obtained by an access user shown by some embodiments of the present application;

[0055] Figure 5It is a block diagram of an ultra-dense networking intelligent resource optimization device shown in some embodiments of the present application;

[0056] Figure 6 It is a block diagram of an electronic device shown in some embodiments of the present application. Specific embodiments

[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0058] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0059] Next, in conjunction with the accompanying drawings, a method for optimizing intelligent resources in ultra-dense networking provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0060] Figure 1 It is a schematic flowchart of a method for optimizing intelligent resources in ultra-dense networking shown in the first embodiment of the present application. Please refer to Figure 1 and the method includes:

[0061] Step S101: Based on NOMA, establish an IAB network including a macro base station, nodes, and users; wherein, the IAB network includes a backhaul link in the millimeter wave band and an access link applying NOMA technology, and the nodes include high-hop nodes and low-hop nodes with different distances from the macro base station; including:

[0062] Divide the nodes into high-hop nodes and low-hop nodes according to the distance from the macro base station; connect the high-hop nodes and low-hop nodes using the millimeter wave band.

[0063] In a possible implementation manner, an innovative design is made for the IAB network architecture. This network system adopts a hierarchical topology structure, consisting of a macro base station , a set of high-hop nodes and a set of low-hop nodes to form a hierarchical transmission architecture. Among them, the high-hop nodes establish a backhaul connection with the low-hop nodes through millimeter wave links to form a communication cluster, and the macro base station Then, backhaul communication is carried out with the high-hop nodes of each communication cluster through the Sub-6GHz band, thereby constructing a heterogeneous transmission system with a two-layer backhaul link.

[0064] Step S102: Establish a communication model for the IAB network, and based on the communication model, determine a constraint function with the optimization goal of maximizing the throughput of access users; including:

[0065] Define a channel parameter model; the channel parameters at least include the channel gain, signal-to-noise ratio, transmission power, and transmission capacity of the backhaul link and the access link; with the optimization goal of maximizing the throughput of access users, determine the network reach limit based on the channel parameter model; determine the constraint function based on the network reach limit.

[0066] In a possible implementation manner, in terms of spectrum resource management, the system innovatively adopts a hybrid frequency band reuse strategy. The backhaul link between the macro base station and the high-hop node and the user access link both use the Sub-6GHz band, and its available sub-band set is defined as ; while the backhaul link between the high-hop node and the low-hop node is deployed in the millimeter wave band, and the available sub-band set is . This dual-band cooperation mechanism not only retains the advantage of the wide coverage range of the Sub-6GHz band but also gives full play to the large bandwidth characteristics of the millimeter wave band, and optimizes the transmission efficiency through intelligent spectrum resource allocation. The situation of users or base stations occupying sub-bands is represented by vectors and ; when occupying the sub-channel , then its spectrum allocation vector is , otherwise .

[0067] In this embodiment, a multi-dimensional channel feature characterization system is established for the backhaul link transmission model, that is, a channel parameter model is defined.

[0068] Macro base station to high-hop node The received signal-to-noise ratio of the Sub-6GHz backhaul link is modeled as:

[0069] ,

[0070] Among them, represents the received signal-to-noise ratio from the macro base station to the high-hop node ; represents the transmission power from the macro base station to the high-hop node ; represents the transmission power from the macro base station to the high-hop node The channel gain incorporates the effects of large-scale fading such as path loss and shadow fading, as well as small-scale fading due to multipath effects; represents the sub-channel from the macro base station to the high-hop node spectrum allocation vector; represents the sum of interference caused by other communication links occupying sub-channel f; is Gaussian white noise.

[0071] High-hop node to the low-hop node The signal-to-noise ratio of the millimeter-wave backhaul link is expressed as:

[0072] ,

[0073] where represents the sub-channel from the high-hop node to the low-hop node spectrum allocation vector; represents the transmit power between the high-hop node and the low-hop node ; represents the downlink channel gain from the high-hop node to the low-hop node , which accurately characterizes the spatial selective fading characteristics of the millimeter-wave channel by fusing the line-of-sight and non-line-of-sight propagation characteristics through a probability model; represents the total interference.

[0074] In this embodiment, NOMA technology is introduced at the user access level. Set the high-hop node associated user set , the low-hop node associated user set . Through the channel gain sorting mechanism and to achieve power domain multiplexing, then the received signal-to-noise ratio models of users and are:

[0075] ,

[0076] ,

[0077] where represents the received signal-to-noise ratio between the high-hop node n and the associated user z; represents the transmit power between the high-hop node n and the associated user z; Denotes the channel gain between the high-hop node n and the associated user z; Denotes the sub-channel between the high-hop node n and the associated user z of the spectrum allocation vector; Denotes the total interference corresponding to the channel; Z denotes the total number of users associated with the high-hop node n.

[0078] Similarly, Denotes the signal-to-noise ratio between the low-hop node m and the associated user s; Denotes the transmit power between the low-hop node m and the associated user s; Denotes the channel gain between the low-hop node m and the associated user s; Denotes the sub-channel between the low-hop node m and the associated user s of the spectrum allocation vector; Denotes the total interference corresponding to the channel; is the Gaussian white noise; S denotes the total number of users associated with the low-hop node m. This model effectively characterizes the unique inter-layer interference characteristics of NOMA technology and improves the spectrum efficiency through successive interference cancellation.

[0079] High-hop node and user The access rate is calculated using the Shannon formula:

[0080] ,

[0081] where Denotes the high-hop node and user the access rate between; Denotes the high-hop node and user the sub-bandwidth used; F denotes the total number of sub-bands used by the high-hop node and user the total number of sub-bands used.

[0082] Similarly, the backhaul rates between the macro base station and the high-hop node , the high-hop node and the low-hop node , and the low-hop node and the user can be calculated respectively. and user access rate .

[0083] Therefore, the total access link rate of the high-hop node n is:

[0084] ,

[0085] Among them, represents the total access link rate of the high-hop node n, represents the access link rate between the high-hop node n and the user z.

[0086] Due to the physical limitations of hierarchical transmission, the high-hop node is subject to a constraint function for the access transmission rate as follows:

[0087] ,

[0088] Among them, represents the constrained access transmission rate of the high-hop node This constraint function accurately depicts the restrictive relationship between the upper-layer backhaul link capacity and the lower-layer transmission capacity.

[0089] Meanwhile, the low-hop node is subject to a constraint function for the transmission rate as follows:

[0090] ,

[0091] Among them, represents the constrained access transmission rate of the low-hop node ; represents the sum of the access link rates of the low-hop node .

[0092] Step S103: Construct a reinforcement learning decision model, including an agent, a state, an action, and a reward function; among them, the reward function is set based on the constraint function;

[0093] This embodiment also constructs an intelligent resource allocation framework based on multi-agent deep reinforcement learning to solve the allocation problem of high-dimensional hybrid resources including sub-band allocation and power control. Macro base stations and nodes deploy autonomous decision-making agents, and each macro base station and node serves as a decision-making unit.

[0094] Taking the access transmission of high-hop nodes as an example, the construction process of the reinforcement learning decision model is as follows: It consists of an agent, a state, an action, and a reward function. Under this model, the network continuously interacts with the environment and finally obtains the optimal allocation strategy with the goal of maximizing the reward value.

[0095] Specifically, set the number of access links of the high-hop node to be , and regard the access links as agents. Therefore, the number of agents is also . The th agent observes the gain of the th access link at this time slot and the interference of the previous time slot as the state , the states observed by all agents constitute the environmental state . Each agent's generated action includes power control and sub-band allocation. Therefore, the action generated by the -th agent is . Since the optimization objective of the constraint function in this embodiment is to maximize the throughput of the accessed users, maximizing the transmission rate of all access links in the network is considered. Therefore, the reward function is defined as:

[0096] ,

[0097] where r(t) represents the reward value; t represents the current time slot; represents the sum of the total rates of the users associated with the n-th high-hop node under the constraint conditions; N represents the total number of high-hop nodes; represents the sum of the total rates of the users associated with the m-th low-hop node under the constraint conditions; M represents the total number of low-hop nodes.

[0098] Step S104: Based on the reinforcement learning decision model and the constraint function, respond to the channel state information in real time through the heterogeneous deep network to generate the maximum Q value and the online decision result, and store them in the experience replay pool; where the heterogeneous deep network includes a policy network and a dueling network; including:

[0099] The agent obtains the current state of the access link; the heterogeneous deep network outputs the online action and the maximum Q value according to the current state; where the online action includes a continuous action set and a discrete action set; input the current state into the policy network to output the continuous action set; input the current state and the continuous action set into the dueling network to determine the maximum Q value; determine the discrete action set corresponding to the maximum Q value through the fully greedy policy. After the online action is executed, obtain the reward value and the subsequent state of the access link; store the online decision result composed of the current state, the continuous action set, the discrete action set, and the subsequent state in the experience replay pool.

[0100] This embodiment is based on multiple agents and heterogeneous two neural networks in terms of functional attributes: the policy network and the dueling network, and is trained using the reinforcement learning framework. The decision-making and training network of a complete agent consists of four deep neural networks: the online policy network, the target policy network, the online dueling network, and the target dueling network. The online network is responsible for the online decision-making process, and the target network participates in the agent training.

[0101] The online decision-making process of this embodiment, specifically, the -th agent observes the -th link state and inputs it to the online policy network. Then, the policy network outputs the continuous action that the link may adopt at this time slot denoted as , the set of continuous actions output by all agents is , that is, the power range that may be controlled. is the number of discrete actions, corresponding to the number of Sub-6GHz sub-bands. If this link is a control millimeter-wave backhaul link, then corresponds to the number of sub-bands of millimeter waves.

[0102] Further, input the set of states observed by the agent and the set of continuous actions output by the policy network into the online confrontation network. The confrontation network outputs a judgment value according to the input information. And determine the maximum value corresponding discrete action through the fully greedy strategy, that is, sub-band control. The discrete action of this process is represented as , where are the confrontation network parameters, and the set of discrete actions determined by all agents is . And select the final continuous action of this access link with the discrete action as the index. The final action is represented as .

[0103] Further, the agent interacts with the environment to obtain a reward value and the subsequent environmental state . The interaction between the agent and the environment is stored in the experience replay pool in the form of a tuple , where represents the current state, represents the set of continuous actions, represents the set of discrete actions, represents the reward value, represents the subsequent state.

[0104] Step S105: Update the policy network and the confrontation network based on the maximum Q value and the loss function to obtain the target heterogeneous deep network; including:

[0105] Update the policy network based on the gradient of the maximum Q value to obtain the latest policy network; the confrontation network calculates the predicted Q value through the value function and the action optimization function; update the confrontation network based on the predicted Q value and the loss function to obtain the latest confrontation network, so as to comprehensively form the target heterogeneous deep network.

[0106] Please refer to Figure 2 , in a possible implementation manner, the gradient of the policy network gradient based on the maximum value update is :

[0107] ,

[0108] ,

[0109] Among them, represents the parameters corresponding to the online policy network, represents the gradient of the parameters, represents the th agent, represents the policy of the policy network, then represents the Q value of the current agent making a decision based on the current policy, represents the current state, represents the continuous action set, represents the discrete action corresponding to the maximum Q value output by the zth agent, represents the discrete action space. During the update process of the decision network, is also adopted and represents the parameters corresponding to the target policy network.

[0110] The dueling network is separately composed of a value function and an action optimization function, and its predicted value is:

[0111] ,

[0112] Among them, represents the predicted value corresponding to the corresponding agent and its corresponding policy state, represents the value function, represents the action optimization function, represents the dueling network parameters, and its subscript corresponds to different functions, represents the total number of discrete actions, represents the ith discrete action in the discrete action set.

[0113] The target value can be expressed as:

[0114] ,

[0115] Among them, r is the reward value, is the discount factor, represents the function for finding the dependent variable corresponding to the maximum Q value, is the value reflected by the online dueling network for the subsequent time slot network state and the subsequent time slot continuous action , represents the continuous action set made by the target policy networks of all agents for the subsequent state. is the network parameter corresponding to the target dueling network;

[0116] Thus, the loss function is , and the update gradient For 。

[0117] Step S106: Generate the optimal resource allocation strategy for the IAB network based on the target heterogeneous deep network.

[0118] Please refer to Figure 3 , which is the graph showing the change of the proposed heterogeneous duel network multi-agent algorithm graph (HM2DP) with the iteration in this embodiment. The comparison algorithms are: heterogeneous network multi-agent algorithm (HMDP), multi-agent deterministic policy algorithm (MADDPG), and multi-agent Q-network learning algorithm (MADQN).

[0119] Please refer to Figure 4 , which is the schematic diagram of the average rate obtained by the access users with the change of users under four algorithms in this embodiment.

[0120] The ultra-dense networking intelligent resource optimization method provided by the above embodiment first constructs an IAB network model including a macro base station, nodes, and users based on NOMA technology. In this network, the backhaul link uses the millimeter wave band to utilize its large bandwidth advantage, and the access link applies NOMA technology to improve the spectrum efficiency. The nodes in the network are divided into high-hop nodes and low-hop nodes to adapt to different transmission requirements. Then, a communication model of the IAB network is established, and based on this, a constraint function is determined to improve the system performance. Subsequently, a reinforcement learning decision model is designed, which consists of an agent, a state, an action, and a reward function, where the reward function is determined according to the set performance metrics. Under this framework, the maximum Q value is calculated and the online decision result is generated in real-time by the heterogeneous deep network in response to the channel state information, and this result is stored in the experience replay pool for subsequent use. The heterogeneous deep network includes a policy network and a duel network, which work together to improve the quality of the decision-making, and based on the calculated maximum Q value and the corresponding loss function, the network is updated to obtain the optimized target heterogeneous deep network. Finally, based on the target heterogeneous deep network, the optimal resource allocation strategy for the IAB network is realized.

[0121] It should be noted that for the ultra-dense networking intelligent resource optimization method provided by the embodiments of the present application, the execution subject can be an ultra-dense networking intelligent resource optimization device, or a control module in the ultra-dense networking intelligent resource optimization device for executing the ultra-dense networking intelligent resource optimization method. In the embodiments of the present application, the case where the ultra-dense networking intelligent resource optimization device executes the ultra-dense networking intelligent resource optimization method is taken as an example to illustrate the method of the ultra-dense networking intelligent resource optimization device provided by the embodiments of the present application.

[0122] Figure 5The following is a schematic diagram of the ultra-dense networking intelligent resource optimization device shown in the second embodiment of the present application. Please refer to Figure 5 The ultra-dense networking intelligent resource optimization device 200 includes:

[0123] The IAB network module 201 is used to establish an IAB network including a macro base station, nodes, and users based on NOMA; wherein, the IAB network includes a backhaul link in the millimeter wave band and an access link applying NOMA technology, and the nodes include high-hop nodes and low-hop nodes with different distances from the macro base station; including:

[0124] The nodes are divided into high-hop nodes and low-hop nodes according to the distance from the macro base station; the high-hop nodes and low-hop nodes are connected using the millimeter wave band.

[0125] The constraint function module 202 is used to establish a communication model of the IAB network and determine a constraint function with the optimization goal of maximizing the throughput of the access users based on the communication model; including:

[0126] Define a channel parameter model; the channel parameters at least include the channel gain, signal-to-noise ratio, transmission power, and transmission capacity of the backhaul link and the access link; with the optimization goal of maximizing the throughput of the access users, determine the network reach limit based on the channel parameter model; determine the constraint function based on the network reach limit.

[0127] The reinforcement learning module 203 is used to construct a reinforcement learning decision model, including an agent, state, action, and reward function; wherein, the reward function is set based on the constraint function;

[0128] The reward function is:

[0129] ,

[0130] wherein, r(t) represents the reward value, t represents the current time slot, represents the sum of the total rates of the users associated with the nth high-hop node under the constraint conditions, N represents the total number of high-hop nodes, represents the sum of the total rates of the users associated with the mth low-hop node under the constraint conditions, M represents the total number of low-hop nodes.

[0131] The decision module 204 is used to generate the maximum Q value and the online decision result in real-time response to the channel state information based on the reinforcement learning decision model and the constraint function, and store them in the experience replay pool; wherein, the heterogeneous deep network includes a policy network and a dueling network; including:

[0132] The agent obtains the current state of the access link; the heterogeneous deep network outputs an online action and a maximum Q value according to the current state; wherein, the online action includes a continuous action set and a discrete action set; the current state is input into the policy network to output the continuous action set; the current state and the continuous action set are input into the dueling network to determine the maximum Q value; the discrete action set corresponding to the maximum Q value is determined through a fully greedy policy. After the online action is executed, a reward value and the subsequent state of the access link are obtained; the online decision result composed of the current state, the continuous action set, the discrete action set, and the subsequent state is stored in the experience replay pool.

[0133] The network update module 205 is configured to update the policy network and the dueling network based on the maximum Q value and the loss function to obtain a target heterogeneous deep network; including:

[0134] Based on the gradient of the maximum Q value, the policy network is updated to obtain the latest policy network; the dueling network calculates the predicted Q value through a value function and an action optimization function; the dueling network is updated based on the predicted Q value and the loss function to obtain the latest dueling network, so as to comprehensively form the target heterogeneous deep network.

[0135] The resource allocation module 206 is configured to generate an optimal resource allocation policy for the IAB network based on the target heterogeneous deep network.

[0136] The ultra-dense networking intelligent resource optimization device in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., which are not specifically limited in the embodiments of the present application.

[0137] The ultra-dense networking intelligent resource optimization device in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0138] The ultra-dense networking intelligent resource optimization device provided in the embodiments of the present application can achieve Figures 1 to 5For the sake of avoiding repetition, the processes implemented by the ultra-dense networking intelligent resource optimization device in the method embodiments are not described herein again.

[0139] Optionally, referring to Figure 6 , the embodiment of the present application further provides an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored on the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, it implements the processes of the above-mentioned ultra-dense networking intelligent resource optimization method embodiments and can achieve the same technical effects. For the sake of avoiding repetition, they are not described herein again.

[0140] The embodiment of the present application further provides a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements the processes of the above-mentioned ultra-dense networking intelligent resource optimization method embodiments and can achieve the same technical effects. For the sake of avoiding repetition, they are not described herein again.

[0141] Wherein, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0142] The embodiment of the present application further provides a chip. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement the processes of the above-mentioned ultra-dense networking intelligent resource optimization method embodiments and can achieve the same technical effects. For the sake of avoiding repetition, they are not described herein again.

[0143] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0144] It should be noted that in this article, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0146] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A method for optimizing ultra-dense network intelligent resources, characterized in that: include: Based on NOMA, an IAB network including a macro base station, nodes and users is established; wherein the IAB network also includes a backhaul link in the millimeter wave frequency band and an access link using the NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; Establishing a communication model of the IAB network, and based on the communication model, determining a constraint function with maximizing the throughput of access users as an optimization goal; Constructing a reinforcement learning decision model, including an agent, a state, an action and a reward function; wherein the reward function is set based on the constraint function; Based on the reinforcement learning decision model and the constraint function, the channel state information is responded to in real time through a heterogeneous deep network to generate a maximum Q value and an online decision result, and store them in an experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; Based on the maximum Q value and the loss function, updating the policy network and the duel network to obtain a target heterogeneous deep network; Based on the target heterogeneous deep network, an optimal resource allocation strategy for the IAB network is generated.

2. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: Based on NOMA, the specific steps to establish an IAB network including macro base stations, nodes and users include: Dividing the nodes into high-hop nodes and low-hop nodes according to the distances from the macro base station; The high-hop node and the low-hop node are connected using a millimeter wave frequency band.

3. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The specific steps of establishing the communication model of the IAB network and determining, based on the communication model, a constraint function with maximizing the throughput of access users as the optimization goal include: Defining a channel parameter model; the channel parameters at least include channel gain, signal-to-noise ratio, transmission power and transmission capacity of the backhaul link and the access link; Taking maximizing the throughput of access users as an optimization goal, determining the network reachability limit based on the channel parameter model; The constraint function is determined based on the network reachability limit.

4. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The reward function is: , Among them, r(t) represents the reward value, t represents the current time slot, represents the sum of the total rates of users associated with the nth high-hop node under the constraint condition, N represents the total number of high-hop nodes, It represents the sum of the total rates of users associated with the mth low-hop node under the constraint condition, and M represents the total number of low-hop nodes.

5. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The specific steps of responding to channel state information in real time through a heterogeneous deep network based on the reinforcement learning decision model and the constraint function to generate a maximum Q value and an online decision result, and storing them in an experience replay pool include: The agent obtains the current state of the access link; The heterogeneous deep network outputs an online action and the maximum Q value according to the current state; wherein the online action includes a continuous action set and a discrete action set; After the online action is executed, obtaining a reward value and a subsequent state of the access link; The online decision result consisting of the current state, the continuous action set, the discrete action set and the subsequent state is stored in the experience replay pool.

6. The ultra-dense networking intelligent resource optimization method according to claim 5, characterized in that: The heterogeneous deep network outputs an online action and a maximum Q value according to the current state, and the specific steps include: Inputting the current state into a policy network to output a set of continuous actions; Input the current state and the continuous action set into a duel network to determine the maximum Q value; The discrete action set corresponding to the maximum Q value is determined by a completely greedy strategy.

7. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The specific steps of updating the policy network and the duel network based on the maximum Q value and the loss function to obtain a target heterogeneous deep network include: Based on the gradient of the maximum Q value, updating the policy network to obtain the latest policy network; The duel network calculates the predicted Q value through the value function and the action optimization function; The duel network is updated based on the predicted Q value and the loss function to obtain a latest duel network to comprehensively form a target heterogeneous deep network.

8. An ultra-dense networking intelligent resource optimization device, used to execute the ultra-dense networking intelligent resource optimization method according to any one of claims 1 to 7, characterized in that: include: An IAB network module is used to establish an IAB network including a macro base station, nodes and users based on NOMA; wherein the IAB network includes a backhaul link in the millimeter wave frequency band and an access link using NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; A constraint function module, used to establish a communication model of the IAB network, and based on the communication model, determine a constraint function with maximizing the throughput of access users as an optimization goal; A reinforcement learning module, used to construct a reinforcement learning decision model, including an agent, a state, an action and a reward function; wherein the reward function is set based on the constraint function; A decision module, for responding to channel state information in real time through a heterogeneous deep network based on the reinforcement learning decision model and the constraint function to generate a maximum Q value and an online decision result, and storing the result in an experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; A network updating module, used for updating the policy network and the duel network based on the maximum Q value and the loss function to obtain a target heterogeneous deep network; A resource allocation module is used to generate an optimal resource allocation strategy for the IAB network based on the target heterogeneous deep network.

9. An electronic device, characterized in that: include: A memory, a processor, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, is a step of the ultra-dense networking intelligent resource optimization method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the ultra-dense networking intelligent resource optimization method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Network resource optimization method, device and equipment and readable storage medium

    CN116887291A

  • Resource allocation method and system for densely deploying NTN (Network Temporary Network) Internet of Things network

    CN118301771A