Super-dense networking intelligent resource optimization method and device, equipment and storage medium
By adopting NOMA technology and reinforcement learning decision model in millimeter wave IAB network, resource allocation problems are solved, efficient and dynamic spectrum and power allocation are achieved, and the performance and resource utilization of the communication network are improved.
Patent Information
- Application Number
- CN202510503135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The resource allocation difficulties in millimeter wave IAB networks, including the dynamic allocation of spectrum subbands and power resources, lead to decreased communication quality, uneven resource occupation and conflicts in multiple services.
The IAB network model based on NOMA is adopted, combined with millimeter wave band and multi-antenna technology, a reinforcement learning decision model is established, and the channel state information is responded in real time through heterogeneous deep networks, the maximum Q value and online decision results are generated, and resource allocation is optimized.
It realizes efficient and dynamic spectrum and power distribution, improves the performance and resource utilization of the next generation of communication networks, and solves the real-time and flexibility requirements in resource allocation problems.
Smart Images

Figure CN120034868A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of wireless communication network resource allocation, and specifically relates to an ultra-dense networking intelligent resource optimization method, device, equipment and storage medium. Background Art
[0002] With the rapid development of 5G and next-generation communication technologies, the number of access terminals in the network has exploded, and user traffic demand has increased significantly. In order to cope with the bottleneck of network capacity, ultra-dense networking has become a key solution, but its deployment cost is high, especially the trenching and equipment installation of optical fiber backhaul links, which makes implementation complicated. To this end, 3GPP proposed an Integrated Access and Backhaul (IAB) network architecture, which relays traffic through wireless backhaul links, reduces deployment costs and simplifies the architecture. However, resource competition in wireless backhaul links has exacerbated spectrum interference problems, posing a challenge to system design.
[0003] Millimeter wave technology is seen as the core solution to this problem. Its frequency band provides spectrum resources far exceeding Sub-6 GHz, but the propagation loss is high and the coverage is limited. It is mainly suitable for short-distance high-speed transmission. To make up for this defect, multi-antenna technology is combined to achieve directional transmission to improve link efficiency, and non-orthogonal multiple access (hereinafter referred to as NOMA) technology is introduced to further improve spectrum utilization. NOMA allows base stations to provide concurrent transmission with differentiated power allocation for multiple users in the same sub-band, optimizing resource utilization. However, the combination of millimeter wave IAB network and NOMA still faces the problem of resource allocation: backhaul and access links need to dynamically allocate spectrum sub-bands and power resources to avoid communication quality degradation, uneven resource occupancy and conflicts in multiple business needs caused by sharing conflicts. Accurate sub-band allocation and power control are the key to improving system performance, but traditional optimization methods are inefficient in large-scale scenarios.
[0004] In recent years, deep reinforcement learning has become a potential solution due to its adaptive decision-making capabilities. Reinforcement learning learns optimal strategies through the interaction between the agent and the environment, while deep reinforcement learning combined with deep learning can handle complex state spaces. Although classical learning algorithms are suitable for discrete action spaces, they are difficult to cope with large-scale continuous action scenarios; although continuous control algorithms can handle continuous actions such as power allocation, they lack support for discrete sub-band allocation. In mixed decision-making scenarios, existing deep reinforcement learning algorithms have low decision-making efficiency due to mismatched action spaces, and are difficult to meet the real-time and flexibility requirements of dynamic resource allocation in millimeter wave IAB networks. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide an ultra-dense networking intelligent resource optimization method, device, equipment and storage medium to address the resource allocation problem in a mixed action space, thereby achieving efficient and dynamic spectrum and power allocation, and improving the performance and resource utilization of the next generation communication network.
[0006] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a method for optimizing ultra-dense networking intelligent resources, the method comprising: Based on NOMA, an IAB network including macro base stations, nodes and users is established; the IAB network includes a backhaul link in the millimeter wave frequency band and an access link using NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; Establish a communication model of the IAB network, and based on the communication model, determine a constraint function with the optimization goal of maximizing the throughput of access users; Construct a reinforcement learning decision model, including agents, states, actions, and reward functions; the reward function is set based on the constraint function; Based on the reinforcement learning decision model and constraint function, the channel state information is responded to in real time through the heterogeneous deep network to generate the maximum Q value and online decision results, which are stored in the experience replay pool; the heterogeneous deep network includes the policy network and the duel network; Based on the maximum Q value and the loss function, the policy network and the duel network are updated to obtain the target heterogeneous deep network; Based on the target heterogeneous deep network, the optimal resource allocation strategy of the IAB network is generated.
[0007] Preferably, based on NOMA, an IAB network including macro base stations, nodes and users is established, and the specific steps include: Nodes are divided into high-hop nodes and low-hop nodes according to their distance from the macro base station; Millimeter wave frequency band is used to connect high-hop nodes and low-hop nodes.
[0008] Preferably, a communication model of the IAB network is established, and based on the communication model, a constraint function is determined with maximizing the throughput of access users as the optimization goal. The specific steps include: Define a channel parameter model; the channel parameters at least include channel gain, signal-to-noise ratio, transmission power and transmission capacity of the backhaul link and the access link; Taking maximizing the throughput of access users as the optimization goal, the network reachability limit is determined based on the channel parameter model; The constraint function is determined based on the network reachability limit.
[0009] Preferably, the reward function is: , Among them, r(t) represents the reward value, t represents the current time slot, represents the sum of the total rates of users associated with the nth high-hop node under the constraint condition, N represents the total number of high-hop nodes, It represents the sum of the total rates of users associated with the mth low-hop node under the constraint condition, and M represents the total number of low-hop nodes.
[0010] Preferably, based on the reinforcement learning decision model and constraint function, the channel state information is responded to in real time through the heterogeneous deep network to generate the maximum Q value and the online decision result, and store them in the experience replay pool. The specific steps include: The agent obtains the current status of the access link; The heterogeneous deep network outputs online actions and maximum Q values according to the current state; the online actions include continuous action sets and discrete action sets; After the online action is executed, the reward value and the subsequent status of the access link are obtained; The online decision results consisting of the current state, continuous action set, discrete action set and subsequent state are stored in the experience replay pool.
[0011] Preferably, the heterogeneous deep network outputs online actions and maximum Q values according to the current state, and the specific steps include: Input the current state into the policy network to output a set of continuous actions; Input the current state and the continuous action set into the duel network to determine the maximum Q value; The set of discrete actions corresponding to the maximum Q value is determined through a completely greedy strategy.
[0012] Preferably, based on the maximum Q value and the loss function, the strategy network and the duel network are updated to obtain the target heterogeneous deep network, and the specific steps include: Based on the gradient of the maximum Q value, update the policy network to obtain the latest policy network; The duel network calculates the predicted Q value through the value function and action optimization function; The duel network is updated based on the predicted Q value and loss function to obtain the latest duel network to comprehensively form the target heterogeneous deep network.
[0013] Compared with the prior art, the above technical solution provided by the present application includes at least the following beneficial effects: This application first builds an IAB network model including macro base stations, nodes and users based on NOMA technology. In this network, the backhaul link uses the millimeter wave frequency band to take advantage of its large bandwidth, and the access link applies NOMA technology to improve spectrum efficiency. The nodes in the network are divided into high-hop nodes and low-hop nodes to adapt to different transmission requirements. Then, the communication model of the IAB network is established, and the constraint function is determined on this basis to improve the system performance. Subsequently, a set of reinforcement learning decision models is designed, which consists of an agent, a state, an action and a reward function, where the reward function is determined according to the set performance indicators. Under this framework, the calculation of the maximum Q value and the generation of online decision results are realized by responding to channel state information in real time through a heterogeneous deep network, and the results will be stored in the experience replay pool for subsequent use. The heterogeneous deep network contains a policy network and a duel network, which work together to improve the quality of decision-making, and based on the calculated maximum Q value and the corresponding loss function, the network is updated to obtain the optimized target heterogeneous deep network. Finally, based on the target heterogeneous deep network, the optimal resource allocation strategy for the IAB network is achieved.
[0014] In a second aspect, an embodiment of the present application provides an ultra-dense networking intelligent resource optimization device, including: An IAB network module is used to establish an IAB network including macro base stations, nodes and users based on NOMA; wherein the IAB network includes a backhaul link in the millimeter wave frequency band and an access link using the NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; A constraint function module is used to establish a communication model of the IAB network and determine a constraint function with maximizing the throughput of access users as an optimization goal based on the communication model; Reinforcement learning module, used to build a reinforcement learning decision model, including agents, states, actions and reward functions; where the reward function is set based on the constraint function; A decision module is used to respond to channel state information in real time through a heterogeneous deep network based on a reinforcement learning decision model and constraint function to generate a maximum Q value and online decision results, and store them in an experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; A network update module is used to update the policy network and the duel network based on the maximum Q value and the loss function to obtain a target heterogeneous deep network; The resource allocation module is used to generate the optimal resource allocation strategy of the IAB network based on the target heterogeneous deep network.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] It can be understood that the beneficial effects of the technical solutions provided in the second, third and fourth aspects can be found in the relevant description of the first aspect, and will not be repeated here.
[0018] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 It is a flowchart of an ultra-dense networking intelligent resource optimization method provided by some embodiments of the present application; Figure 2 is a schematic diagram of the structure of a heterogeneous deep network update process shown in some embodiments of the present application; Figure 3 is a schematic diagram of heterogeneous network algorithm iteration shown in some embodiments of the present application; Figure 4 is a schematic diagram of an average rate obtained by an access user shown in some embodiments of the present application; Figure 5 is a block diagram of an ultra-dense networking intelligent resource optimization device shown in some embodiments of the present application; Figure 6 It is a block diagram of an electronic device shown in some embodiments of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0022] In the following, in conjunction with the accompanying drawings, an ultra-dense networking intelligent resource optimization method provided by an embodiment of the present application is described in detail through specific embodiments and application scenarios.
[0023] Figure 1 This is a flow chart of a method for optimizing ultra-dense network intelligent resources according to the first embodiment of the present application. Figure 1 , the method comprising: Step S101: Based on NOMA, an IAB network including a macro base station, nodes and users is established; wherein the IAB network includes a backhaul link in the millimeter wave frequency band and an access link using NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; including: According to the distance from the macro base station, the nodes are divided into high-hop nodes and low-hop nodes; the millimeter wave frequency band is used to connect the high-hop nodes and the low-hop nodes.
[0024] In a possible implementation, an innovative design is made to the IAB network architecture. The network system adopts a hierarchical topology structure consisting of macro base stations. , high jump node collection and the low-hop node set A hierarchical transmission architecture is formed. Among them, high-hop nodes establish backhaul connections with low-hop nodes through millimeter wave links to form a communication cluster, and macro base stations Backhaul communication is carried out with the high-hop nodes of each communication cluster through the Sub-6GHz frequency band, thus building a heterogeneous transmission system with a double-layer backhaul link.
[0025] Step S102: establishing a communication model of the IAB network, and based on the communication model, determining a constraint function with maximizing the throughput of access users as the optimization goal; including: A channel parameter model is defined; the channel parameters include at least the channel gain, signal-to-noise ratio, transmission power and transmission capacity of the backhaul link and the access link; the network reachability limit is determined based on the channel parameter model with maximizing the throughput of access users as the optimization goal; and the constraint function is determined based on the network reachability limit.
[0026] In a possible implementation, the system innovatively adopts a mixed frequency band reuse strategy in spectrum resource management. The backhaul link between the macro base station and the high-hop node and the user access link both use the Sub-6GHz frequency band, and the available sub-band set is defined as The backhaul link between high-hop nodes and low-hop nodes is deployed in the millimeter wave frequency band, and the available sub-band set is This dual-band coordination mechanism not only retains the advantage of the wide coverage of the Sub-6GHz band, but also gives full play to the large bandwidth characteristics of the millimeter wave band, and optimizes transmission efficiency through intelligent allocation of spectrum resources. The sub-band occupation of users or base stations is represented by vectors. and Indicates; when the subchannel is occupied , then its spectrum allocation vector ,otherwise .
[0027] This embodiment establishes a multi-dimensional channel feature characterization system for the backhaul link transmission model, that is, defines a channel parameter model.
[0028] Macro Base Station To high jump node The Sub-6GHz backhaul link receive signal-to-noise ratio is modeled as: , in, Indicates that the macro base station To high jump node The received signal-to-noise ratio; Indicates that the macro base station To high jump node The transmission power; Indicates that the macro base station To high jump node The channel gain combines the effects of large-scale fading such as path loss and shadow fading with small-scale fading due to multipath effects. Indicates that the macro base station To high jump node Subchannel Spectrum allocation vector of ; represents the sum of interference caused by other communication links occupying subchannel f; is Gaussian white noise.
[0029] High jump node To low hop node The signal-to-noise ratio of the millimeter wave backhaul link is expressed as: , in, Indicates that the node jumps from the high To the low-hop node of the sub-channel spectrum allocation vector; indicating from the high-hop node to the low-hop node transmit power between; indicating from the high-hop node to the low-hop node downlink channel gain, accurately characterizing the spatial selective fading characteristics of the millimeter-wave channel by fusing the line-of-sight propagation and non-line-of-sight propagation characteristics through a probability model; indicating the total interference.
[0030] In this embodiment, the NOMA technology is introduced at the user access level. Set the high-hop node associated user set , the low-hop node associated user set . Through the channel gain sorting mechanism and to achieve power domain multiplexing, then the received signal-to-noise ratio models of users and are: , , where, represents the received signal-to-noise ratio between the high-hop node n and the associated user z; represents the transmit power between the high-hop node n and the associated user z; represents the channel gain between the high-hop node n and the associated user z; represents the sub-channel spectrum allocation vector between the high-hop node n and the associated user z; represents the total interference corresponding to the channel; Z represents the total number of users associated with the high-hop node n.
[0031] Similarly, represents the signal-to-noise ratio between the low-hop node m and the associated user s; represents the transmit power between the low-hop node m and the associated user s; represents the channel gain between the low-hop node m and the associated user s; represents the sub-channel spectrum allocation vector between the low-hop node m and the associated user s; represents the total interference corresponding to the channel; is Gaussian white noise; S represents the total number of users associated with the low-hop node m. This model effectively characterizes the inter-layer interference characteristics unique to NOMA technology and improves spectrum efficiency through serial interference elimination.
[0032] High jump node With users The access rate is calculated using the Shannon formula: , in, Indicates high jump node With users The access rate between Indicates high jump node With users Sub-band bandwidth used; F indicates high-hop node With users Total number of sub-bands used.
[0033] Similarly, the macro base stations can be calculated separately With high jump node The return rate , high jump node With low hop node The return rate between , low hop node With users Access rate .
[0034] Therefore, the total access link rate of high-hop node n is: , in, represents the total access link rate of high-hop node n, It represents the access link rate between high-hop node n and user z.
[0035] Since hierarchical transmission is subject to physical limitations, high-hop nodes The constraint function of the access transmission rate is: , in, Indicates high jump node The access transmission rate is constrained; this constraint function accurately describes the restriction relationship between the upper layer backhaul link capacity and the lower layer transmission capacity.
[0036] At the same time, low-hop nodes The transmission rate constraint function is: , in, Indicates low hop node Constrained access transmission rate; Indicates low hop node The sum of the access link rates.
[0037] Step S103: constructing a reinforcement learning decision model, including an agent, a state, an action, and a reward function; wherein the reward function is set based on a constraint function; This embodiment also constructs an intelligent resource allocation framework based on multi-agent deep reinforcement learning to solve the allocation problem of high-dimensional mixed resources including sub-band allocation and power control. Macro base stations and nodes deploy autonomous decision-making agents, and each macro base station and node serves as a decision-making unit.
[0038] Taking high-hop node access transmission as an example, the construction process of the reinforcement learning decision model is as follows: it consists of an agent, state, action, and reward function. Under this model, the network continuously interacts with the environment and ultimately obtains the best allocation strategy with the goal of maximizing the reward value.
[0039] Specifically, the number of high-hop node access links is set to , the access link is regarded as an intelligent agent, so the number of intelligent agents is also . No. The agent observes The gain of the access link at this time slot and the interference of the previous time slot are taken as the state , the states observed by all agents constitute the environment state Each agent generates actions including power control and sub-band allocation, so the The action generated by the agent is Since the optimization goal of the constraint function of this embodiment is to maximize the throughput of access users, the transmission rate of all access links in the network is maximized. Therefore, the reward function is defined as: , Among them, r(t) represents the reward value; t represents the current time slot; It represents the sum of the total rates of users associated with the nth high-hop node under the constraint condition; N represents the total number of high-hop nodes; It represents the sum of the total rates of users associated with the mth low-hop node under the constraint condition; M represents the total number of low-hop nodes.
[0040] Step S104: Based on the reinforcement learning decision model and constraint function, the channel state information is responded to in real time through the heterogeneous deep network to generate the maximum Q value and the online decision result, and store them in the experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; including: The agent obtains the current state of the access link; the heterogeneous deep network outputs online actions and maximum Q values according to the current state; the online actions include continuous action sets and discrete action sets; the current state is input into the policy network to output the continuous action set; the current state and the continuous action set are input into the duel network to determine the maximum Q value; the discrete action set corresponding to the maximum Q value is determined through a completely greedy strategy. After the online action is executed, the reward value and the subsequent state of the access link are obtained; the online decision result consisting of the current state, continuous action set, discrete action set and subsequent state is stored in the experience replay pool.
[0041] This embodiment is based on multi-agents, and heterogeneously constructs two neural networks in terms of functional attributes: the policy network and the duel network, and uses a reinforcement learning framework for training. A complete agent's decision-making and training network consists of four deep neural networks: the online policy network, the target policy network, the online duel network, and the target duel network. The online network is responsible for the online decision-making process, and the target network participates in the agent training.
[0042] In the online decision-making process of this embodiment, specifically, The agent observes The link state is input into the online strategy network, and then the strategy network outputs the continuous actions that may be taken by the link at this time slot Expressed as , the set of continuous actions output by all agents is , that is, the power range that can be controlled. is the number of discrete actions, corresponding to the number of Sub-6GHz sub-bands. If the link is a control millimeter wave backhaul link, then Corresponding to the number of millimeter wave sub-bands.
[0043] Furthermore, the set of states observed by the agent and the set of continuous actions output by the strategy network are input into the online duel network, and the duel outputs the judgment based on the input information. And determine the maximum value through a completely greedy strategy The discrete action corresponding to the value is sub-band control. Expressed as ,in, To compete with the network parameters, the set of discrete actions determined by all agents is . And use the discrete action as an index to select the final continuous action of the access link. The final action Expressed as .
[0044] Furthermore, the agent interacts with the environment to obtain a reward value and subsequent environmental status The interaction between the agent and the environment is expressed as a tuple The form is stored in the experience replay pool, where Indicates the current state. represents a set of continuous actions, represents a set of discrete actions, Represents the reward value, Indicates the subsequent status.
[0045] Step S105: Based on the maximum Q value and the loss function, the policy network and the duel network are updated to obtain a target heterogeneous deep network; including: Based on the gradient of the maximum Q value, the policy network is updated to obtain the latest policy network; the duel network calculates the predicted Q value through the value function and the action optimization function; based on the predicted Q value and the loss function, the duel network is updated to obtain the latest duel network to comprehensively form the target heterogeneous deep network.
[0046] See also Figure 2 In one possible implementation, the policy network gradient is based on the maximum The updated gradient of the value for: , , in, represents the parameters corresponding to the online policy network, represents the gradient of the parameter, Indicates An intelligent agent, represents the policy of the policy network, It represents the Q value of the decision made by the current agent based on the current strategy. Indicates the current state. represents a set of continuous actions, represents the discrete action corresponding to the maximum Q value output by the z-th agent, Represents the discrete action space. In the updating process of the decision network, Represents the parameters corresponding to the target policy network.
[0047] The duel network consists of a value function and an action optimization function, and its prediction The values are: , in, Represents the prediction of the corresponding agent and its corresponding policy state value, represents the value function, represents the action optimization function, Represents the duel network parameters, and its subscripts correspond to different functions. represents the total number of discrete actions, Represents the i-th discrete action in the discrete action set.
[0048] Target value It can be expressed as: , Among them, r is the reward value, is the discount factor, It means to find the dependent variable function corresponding to the maximum Q value, Is the online duel network for the subsequent time slot network status and subsequent time slots for continuous action The reaction value, Represents the set of continuous actions taken by the target policy network of all agents for subsequent states. The network parameters corresponding to the target duel network; Therefore, the loss function is , update the gradient for .
[0049] Step S106: Generate an optimal resource allocation strategy for the IAB network based on the target heterogeneous deep network.
[0050] See also Figure 3 ,This embodiment explores the change of the proposed heterogeneous duel network multi-agent algorithm graph (HM2DP) with the iteration situation. ,The compared algorithms are: heterogeneous network multi-agent algorithm (HMDP), multi-agent deterministic strategy algorithm (MADDPG), multi-agent Q network learning algorithm (MADQN).
[0051] See also Figure 4 , which is a schematic diagram of the average rate obtained by access users as the number of users changes under the four algorithms in this embodiment.
[0052] The ultra-dense networking intelligent resource optimization method provided in the above embodiment firstly builds an IAB network model including macro base stations, nodes and users based on NOMA technology. In this network, the backhaul link adopts the millimeter wave frequency band to take advantage of its large bandwidth advantage, and the access link applies NOMA technology to improve spectrum efficiency. The nodes in the network are divided into high-hop nodes and low-hop nodes to adapt to different transmission requirements. Then, a communication model of the IAB network is established, and the constraint function is determined on this basis to improve the system performance. Subsequently, a set of reinforcement learning decision models is designed, which consists of an intelligent agent, a state, an action and a reward function, wherein the reward function is determined according to the set performance index. Under this framework, the calculation of the maximum Q value and the generation of online decision results are realized by responding to the channel state information in real time through the heterogeneous deep network, and the results will be stored in the experience replay pool for subsequent use. The heterogeneous deep network includes a policy network and a duel network, which work together to improve the quality of decision-making, and based on the calculated maximum Q value and the corresponding loss function, the network is updated to obtain the optimized target heterogeneous deep network. Finally, based on the target heterogeneous deep network, the optimal resource allocation strategy is formulated for the IAB network.
[0053] It should be noted that the ultra-dense networking intelligent resource optimization method provided in the embodiment of the present application can be executed by an ultra-dense networking intelligent resource optimization device, or a control module in the ultra-dense networking intelligent resource optimization device for executing the loading ultra-dense networking intelligent resource optimization method. In the embodiment of the present application, the method of the ultra-dense networking intelligent resource optimization device provided in the embodiment of the present application is explained by taking the ultra-dense networking intelligent resource optimization device executing the loading ultra-dense networking intelligent resource optimization method as an example.
[0054] Figure 5 is a schematic diagram of an ultra-dense networking intelligent resource optimization device according to the second embodiment of the present application, please refer to Figure 5 The ultra-dense networking intelligent resource optimization device 200 includes: The IAB network module 201 is used to establish an IAB network including a macro base station, nodes and users based on NOMA; wherein the IAB network includes a backhaul link in the millimeter wave frequency band and an access link using the NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; and includes: According to the distance from the macro base station, the nodes are divided into high-hop nodes and low-hop nodes; the millimeter wave frequency band is used to connect the high-hop nodes and the low-hop nodes.
[0055] The constraint function module 202 is used to establish a communication model of the IAB network and determine a constraint function with maximizing the throughput of access users as an optimization goal based on the communication model; including: A channel parameter model is defined; the channel parameters include at least the channel gain, signal-to-noise ratio, transmission power and transmission capacity of the backhaul link and the access link; the network reachability limit is determined based on the channel parameter model with maximizing the throughput of access users as the optimization goal; and the constraint function is determined based on the network reachability limit.
[0056] A reinforcement learning module 203 is used to construct a reinforcement learning decision model, including an agent, a state, an action, and a reward function; wherein the reward function is set based on a constraint function; The reward function is: , Among them, r(t) represents the reward value, t represents the current time slot, represents the sum of the total rates of users associated with the nth high-hop node under the constraint condition, N represents the total number of high-hop nodes, It represents the sum of the total rates of users associated with the mth low-hop node under the constraint condition, and M represents the total number of low-hop nodes.
[0057] The decision module 204 is used to respond to the channel state information in real time through the heterogeneous deep network based on the reinforcement learning decision model and the constraint function to generate the maximum Q value and the online decision result, and store them in the experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; including: The agent obtains the current state of the access link; the heterogeneous deep network outputs online actions and maximum Q values according to the current state; the online actions include continuous action sets and discrete action sets; the current state is input into the policy network to output the continuous action set; the current state and the continuous action set are input into the duel network to determine the maximum Q value; the discrete action set corresponding to the maximum Q value is determined through a completely greedy strategy. After the online action is executed, the reward value and the subsequent state of the access link are obtained; the online decision result consisting of the current state, continuous action set, discrete action set and subsequent state is stored in the experience replay pool.
[0058] The network update module 205 is used to update the policy network and the duel network based on the maximum Q value and the loss function to obtain the target heterogeneous deep network; including: Based on the gradient of the maximum Q value, the policy network is updated to obtain the latest policy network; the duel network calculates the predicted Q value through the value function and the action optimization function; based on the predicted Q value and the loss function, the duel network is updated to obtain the latest duel network to comprehensively form the target heterogeneous deep network.
[0059] The resource allocation module 206 is used to generate an optimal resource allocation strategy for the IAB network based on the target heterogeneous deep network.
[0060] The ultra-dense networking intelligent resource optimization device in the embodiment of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0061] The ultra-dense networking intelligent resource optimization device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0062] The ultra-dense networking intelligent resource optimization device provided in the embodiment of the present application can achieve Figures 1 to 5 The various processes implemented by the ultra-dense networking intelligent resource optimization device in the method embodiment will not be described here to avoid repetition.
[0063] Optionally, see Figure 6 The embodiment of the present application also provides an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, each process of the above-mentioned ultra-dense networking intelligent resource optimization method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0064] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned ultra-dense networking intelligent resource optimization method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0065] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0066] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned ultra-dense networking intelligent resource optimization method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0067] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0068] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0069] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0070] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for optimizing ultra-dense network intelligent resources, characterized in that: include: Based on NOMA, an IAB network including a macro base station, nodes and users is established; wherein the IAB network also includes a backhaul link in the millimeter wave frequency band and an access link using the NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; Establishing a communication model of the IAB network, and based on the communication model, determining a constraint function with maximizing the throughput of access users as an optimization goal; Constructing a reinforcement learning decision model, including an agent, a state, an action and a reward function; wherein the reward function is set based on the constraint function; Based on the reinforcement learning decision model and the constraint function, the channel state information is responded to in real time through a heterogeneous deep network to generate a maximum Q value and an online decision result, and store them in an experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; Based on the maximum Q value and the loss function, updating the policy network and the duel network to obtain a target heterogeneous deep network; Based on the target heterogeneous deep network, an optimal resource allocation strategy for the IAB network is generated.
2. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: Based on NOMA, the specific steps to establish an IAB network including macro base stations, nodes and users include: Dividing the nodes into high-hop nodes and low-hop nodes according to the distances from the macro base station; The high-hop node and the low-hop node are connected using a millimeter wave frequency band.
3. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The specific steps of establishing the communication model of the IAB network and determining, based on the communication model, a constraint function with maximizing the throughput of access users as the optimization goal include: Defining a channel parameter model; the channel parameters at least include channel gain, signal-to-noise ratio, transmission power and transmission capacity of the backhaul link and the access link; Taking maximizing the throughput of access users as an optimization goal, determining the network reachability limit based on the channel parameter model; The constraint function is determined based on the network reachability limit.
4. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The reward function is: , Among them, r(t) represents the reward value, t represents the current time slot, represents the sum of the total rates of users associated with the nth high-hop node under the constraint condition, N represents the total number of high-hop nodes, It represents the sum of the total rates of users associated with the mth low-hop node under the constraint condition, and M represents the total number of low-hop nodes.
5. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The specific steps of responding to channel state information in real time through a heterogeneous deep network based on the reinforcement learning decision model and the constraint function to generate a maximum Q value and an online decision result, and storing them in an experience replay pool include: The agent obtains the current state of the access link; The heterogeneous deep network outputs an online action and the maximum Q value according to the current state; wherein the online action includes a continuous action set and a discrete action set; After the online action is executed, obtaining a reward value and a subsequent state of the access link; The online decision result consisting of the current state, the continuous action set, the discrete action set and the subsequent state is stored in the experience replay pool.
6. The ultra-dense networking intelligent resource optimization method according to claim 5, characterized in that: The heterogeneous deep network outputs an online action and a maximum Q value according to the current state, and the specific steps include: Inputting the current state into a policy network to output a set of continuous actions; Input the current state and the continuous action set into a duel network to determine the maximum Q value; The discrete action set corresponding to the maximum Q value is determined by a completely greedy strategy.
7. The ultra-dense networking intelligent resource optimization method according to claim 1, characterized in that: The specific steps of updating the policy network and the duel network based on the maximum Q value and the loss function to obtain a target heterogeneous deep network include: Based on the gradient of the maximum Q value, updating the policy network to obtain the latest policy network; The duel network calculates the predicted Q value through the value function and the action optimization function; The duel network is updated based on the predicted Q value and the loss function to obtain a latest duel network to comprehensively form a target heterogeneous deep network.
8. An ultra-dense networking intelligent resource optimization device, used to execute the ultra-dense networking intelligent resource optimization method according to any one of claims 1 to 7, characterized in that: include: An IAB network module is used to establish an IAB network including a macro base station, nodes and users based on NOMA; wherein the IAB network includes a backhaul link in the millimeter wave frequency band and an access link using NOMA technology, and the nodes include high-hop nodes and low-hop nodes at different distances from the macro base station; A constraint function module, used to establish a communication model of the IAB network, and based on the communication model, determine a constraint function with maximizing the throughput of access users as an optimization goal; A reinforcement learning module, used to construct a reinforcement learning decision model, including an agent, a state, an action and a reward function; wherein the reward function is set based on the constraint function; A decision module, for responding to channel state information in real time through a heterogeneous deep network based on the reinforcement learning decision model and the constraint function to generate a maximum Q value and an online decision result, and storing the result in an experience replay pool; wherein the heterogeneous deep network includes a policy network and a duel network; A network updating module, used for updating the policy network and the duel network based on the maximum Q value and the loss function to obtain a target heterogeneous deep network; A resource allocation module is used to generate an optimal resource allocation strategy for the IAB network based on the target heterogeneous deep network.
9. An electronic device, characterized in that: include: A memory, a processor, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, is a step of the ultra-dense networking intelligent resource optimization method as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the ultra-dense networking intelligent resource optimization method as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Network resource optimization method, device and equipment and readable storage medium
CN116887291A
Resource allocation method and system for densely deploying NTN (Network Temporary Network) Internet of Things network
CN118301771A
Deep learning-based user clustering in millimeter wave non-orthogonal multiple acccess communications
US20240072923A1
Full-duplex non-orthogonal multiple access-based transmit power control device employing deep reinforcement learning
US20240422693A1
Cited By
Super-dense networking intelligent communication reliable transmission and resource optimization method and device
CN121985344A
Methods and devices for reliable transmission and resource optimization in ultra-dense network intelligent communication
CN121985344B