Resource allocation method and device in network, electronic device and storage medium

By constructing user utility functions and agent models to optimize resource allocation, the problems of poor user performance and resource waste in dense heterogeneous networks are solved, fairness among users and joint optimization of network energy and spectrum efficiency are achieved, and resource utilization is improved.

CN118741706BActive Publication Date: 2025-09-19CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410989394.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-09-19
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

Dense heterogeneous networks suffer from problems such as poor user performance experience, low resource utilization, and resource waste, especially energy consumption and interference caused by the dense deployment of small base stations.

Method used

By constructing a user utility function and combining the weights of spectrum efficiency and energy efficiency, the allocation of resource blocks and transmission power are optimized. A proxy model is used for user association and resource allocation to ensure that the user rate meets the minimum requirements. The system utility function is optimized through a linear weighted method to achieve fairness in resource allocation among users and joint optimization of energy and spectrum efficiency.

Benefits of technology

Under the premise of ensuring the minimum rate requirements of users, fairness in resource allocation among users is achieved, while the energy efficiency and spectrum efficiency of the network are improved, and energy consumption and interference are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118741706B_ABST
    Figure CN118741706B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a resource allocation method and apparatus, an electronic device, and a storage medium in a network. Compared with related technologies, the embodiments of the present application achieve fairness in resource allocation among users by allowing users to freely associate with base stations. This is combined with a resource allocation method that jointly optimizes energy efficiency and spectrum efficiency. On the basis of jointly optimizing network energy efficiency and spectrum efficiency, fairness among users is also taken into account.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a resource allocation method and device in a network, an electronic device, and a storage medium. Background Art

[0002] With the booming development of mobile internet services and the rapid adoption of smart devices, the massive access of devices and the massive data transmission are driving explosive growth in the capacity demand for wireless communication networks. Dense heterogeneous networks can effectively improve network coverage and throughput, but the dense deployment of small base stations can lead to significant energy consumption and interference. Furthermore, differences in channel quality, coverage range, and transmit power between different types of base stations lead to increasingly serious resource allocation inequities, further resulting in poor user experience, low resource utilization, and even waste. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for allocating resources in a network, the main purpose of which is to solve the problems of poor user performance experience, low resource utilization, and even resource waste.

[0004] According to a first aspect of the present disclosure, a method for allocating resources in a network is provided, comprising:

[0005] Acquire a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; the user is associated with one base station;

[0006] Allocate the different resource blocks occupied by each base station to different users with related relationships, and construct a utility function to characterize each user based on the downlink rate of each user;

[0007] Calculating a function to be optimized based on the sum of the utility functions of each user and a first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and a second weight; wherein the spectrum efficiency and energy efficiency are calculated based on the network model; and the first weight and the second weight are calculated in advance;

[0008] Each of the macro base station and each of the small base station is used as an agent to repeatedly allocate resource blocks to associated users and determine the transmit power on different resource blocks;

[0009] When the transmission power corresponding to each user meets or exceeds the minimum rate requirement, the weights in the function to be optimized are obtained according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value, thereby obtaining the target optimization function.

[0010] Optionally, the allocating different resource blocks occupied by each base station to different associated users, and constructing a utility function representing each user according to the downlink rate of each user further includes:

[0011] Calculate the received power between the user and the associated base station and the co-channel interference power and noise interference power between the user and the non-associated base station;

[0012] Calculating a signal-to-interference-plus-noise ratio (SINR) based on the received power, the co-channel interference power, and the noise interference power, and calculating a downlink rate of the user based on the SINR;

[0013] The utility function of the user is calculated according to the downlink rate and a preset fairness factor.

[0014] Optionally, before calculating the function to be optimized based on the sum of the utility functions of each user and the first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight, the method further includes:

[0015] Counting the downlink rate of each user to obtain the system capacity of the network model;

[0016] Counting the power consumption of the base stations in the network model to obtain the total power consumption of the network model;

[0017] Counting the bandwidth of each user to obtain the total bandwidth of the network model;

[0018] Calculating the energy efficiency of the network model according to the system capacity and the total power consumption;

[0019] The spectrum efficiency of the network model is calculated according to the system capacity and the total bandwidth.

[0020] Optionally, the calculating the function to be optimized based on the sum of the utility functions of each user and the first weight and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight further includes:

[0021] Based on a linear weighted method, the sum of the utility functions of the respective users and the system power consumption are combined into the function to be optimized.

[0022] Optionally, when the transmit power corresponding to each user satisfies a requirement of being greater than or equal to a minimum rate, obtaining each weight in the function to be optimized according to a weighted square of a distance between each target value corresponding to the target state and its maximum target value, and obtaining the target optimization function further includes:

[0023] When the transmission power corresponding to each of the users meets or exceeds the minimum rate requirement, determining the reward value to be a first preset value;

[0024] Calculating a first priority parameter and a second priority parameter according to the reward value;

[0025] Calculating a reward function for each base station according to the first priority parameter and the second priority parameter;

[0026] Update the action value function according to the preset learning rate and the preset discount rate;

[0027] Repeat the above calculation process until the number of calculations reaches the preset calculation threshold; and obtain the target optimization function.

[0028] According to a second aspect of the present disclosure, a resource allocation device in a network is provided, characterized by comprising:

[0029] An acquisition unit is configured to acquire a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; and the user is associated with one base station;

[0030] An allocation unit is used to allocate different resource blocks occupied by each base station to different users with associated relationships, and construct a utility function representing each user according to the downlink rate of each user;

[0031] a first calculation unit, configured to calculate a function to be optimized based on the sum of the utility functions of each user and a first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and a second weight; wherein the spectrum efficiency and energy efficiency are calculated based on the network model; and the first weight and the second weight are calculated in advance;

[0032] a determining unit, configured to use each of the macro base station and each of the small base stations as agents, repeatedly allocating resource blocks to associated users, and determining transmit power on different resource blocks;

[0033] The second calculation unit is used to obtain the weights in the function to be optimized according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value when the transmission power corresponding to each user meets the requirement of being greater than or equal to the minimum rate, so as to obtain the target optimization function.

[0034] Optionally, the allocation unit is further configured to:

[0035] Calculate the received power between the user and the associated base station and the co-channel interference power and noise interference power between the user and the non-associated base station;

[0036] Calculating a signal-to-interference-plus-noise ratio (SINR) based on the received power, the co-channel interference power, and the noise interference power, and calculating a downlink rate of the user based on the SINR;

[0037] The utility function of the user is calculated according to the downlink rate and a preset fairness factor.

[0038] Optionally, the device further includes:

[0039] a statistical unit, configured to calculate, by the first calculating unit, the function to be optimized based on the sum of the utility functions of each user and the first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight, a statistical downlink rate of each user to obtain the system capacity of the network model;

[0040] The statistical unit is further configured to count the power consumption of the base stations in the network model to obtain the total power consumption of the network model;

[0041] The statistics unit is further configured to count the bandwidth of each user to obtain the total bandwidth of the network model;

[0042] a third calculating unit, configured to calculate the energy efficiency of the network model according to the system capacity and the total power consumption;

[0043] A fourth calculation unit is configured to calculate the spectrum efficiency of the network model according to the system capacity and the total bandwidth.

[0044] Optionally, the first computing unit is further configured to:

[0045] Based on a linear weighted method, the sum of the utility functions of the respective users and the system power consumption are combined into the function to be optimized.

[0046] Optionally, the second computing unit is further configured to:

[0047] When the transmission power corresponding to each of the users meets or exceeds the minimum rate requirement, determining the reward value to be a first preset value;

[0048] Calculating a first priority parameter and a second priority parameter according to the reward value;

[0049] Calculating a reward function for each base station according to the first priority parameter and the second priority parameter;

[0050] Update the action value function according to the preset learning rate and the preset discount rate;

[0051] Repeat the above calculation process until the number of calculations reaches the preset calculation threshold; and obtain the target optimization function.

[0052] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0053] at least one processor; and

[0054] a memory communicatively connected to the at least one processor; wherein,

[0055] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.

[0056] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.

[0057] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.

[0058] The resource allocation method, device, electronic device and storage medium in the network provided by the present disclosure have the following main technical solutions: obtaining a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; the user is associated with a base station; different resource blocks occupied by each base station are allocated to different users with an associated relationship, and a utility function characterizing each user is constructed according to the downlink rate of each user; the function to be optimized is calculated according to the sum of the utility functions of each user and a first weight and the sum of the power consumption of spectrum efficiency and energy efficiency and a second weight; wherein the spectrum efficiency and energy efficiency are calculated according to the network model; the first weight and the second weight are calculated in advance; each macro base station and each small base station are used as agents to repeatedly allocate resource blocks to the associated users and determine the transmission power on different resource blocks; when the transmission power corresponding to each user meets the requirement of being greater than or equal to the minimum rate, the weights in the function to be optimized are obtained according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value, thereby obtaining the target optimization function. Compared with related technologies, the embodiments of the present application achieve fairness in resource allocation among users by allowing users to freely associate with base stations, and combine this with a resource allocation method that jointly optimizes energy efficiency and spectrum efficiency. On the basis of jointly optimizing network energy efficiency and spectrum efficiency, fairness among users is also taken into account.

[0059] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0061] Figure 1 A schematic diagram of a flow chart of a resource allocation method in a network provided by an embodiment of the present disclosure;

[0062] Figure 2 A schematic diagram of a flow chart of a resource allocation method in a network provided by an embodiment of the present disclosure;

[0063] Figure 3 A schematic diagram of a flow chart of a resource allocation method in a network provided by an embodiment of the present disclosure;

[0064] Figure 4 A schematic diagram of a flow chart of a resource allocation method in a network provided by an embodiment of the present disclosure;

[0065] Figure 5 A schematic diagram of the structure of a resource allocation device in a network provided by an embodiment of the present disclosure;

[0066] Figure 6 A schematic diagram of the structure of a resource allocation device in a network provided by an embodiment of the present disclosure;

[0067] Figure 7 A schematic block diagram of an exemplary electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0068] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0069] The following describes a method, apparatus, electronic device, and storage medium for allocating resources in a network according to embodiments of the present disclosure with reference to the accompanying drawings.

[0070] Figure 1 A flowchart of a resource allocation method in a network provided by an embodiment of the present disclosure is provided.

[0071] like Figure 1 As shown, the method comprises the following steps:

[0072] Step 101, obtain a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; the user is associated with a base station.

[0073] A macro base station is located at the center of the network, providing basic services to users. Its coverage area is relatively large. N small base stations and M users are randomly distributed in the network. Each user can only associate with a macro base station or one small base station, but the macro base station and each small base station can be associated with multiple users.

[0074] Step 102: Allocate different resource blocks occupied by each base station to different users with associated relationships, and construct a utility function representing each user according to the downlink rate of each user.

[0075] The total number of base stations is N+1, and the base station set is represented by BS={B0,B1,...,B n ,...,B N}, where B0 represents a macro base station, B1, B2, ..., B n ,...,B N Represents a small base station. The set of M users is represented as UE={U1,U2,...,U m ,...,U M}, U m Represents any user. The system is based on OFDMA technology. The spectrum is divided into K orthogonal resource blocks. The bandwidth of each resource block is B. The resource block set is represented by R = {R1, R2, ..., R k ,...,R K}, R k Represents any resource block. Due to limited spectrum resources in the network, the present invention assumes that all resource blocks are shared between base stations. Therefore, co-channel interference will occur between users reusing the same resource blocks. To reduce co-channel interference, the present embodiment further stipulates that each resource block occupied by each base station can only be allocated to a different user associated with it, but each user can be allocated multiple resource blocks.

[0076] In some embodiments, the utility function represents a function of the user's receiving rate, which is associated with the resource block associated with the user and the interference received. In practical applications, it can be calculated based on the device performance and interference conditions of the resource block. This embodiment of the present application is not limited to this.

[0077] Step 103: Calculate the function to be optimized based on the sum of the utility function of each user and the first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight; wherein the spectrum efficiency and energy efficiency are calculated based on the network model; and the first weight and the second weight are calculated in advance.

[0078] The sum of the utility functions of all users is calculated based on the utility function of each user, which represents the total utility of the entire network.

[0079] According to the first weight calculated in advance, the sum of the user utility functions is multiplied by the first weight to obtain the weighted sum of the utility functions.

[0080] The sum of the power consumption of spectrum efficiency and energy efficiency is calculated. In some embodiments, this sum can be calculated based on a network model and a specific algorithm and is used to measure the resource utilization efficiency and power consumption of the network. Based on a pre-calculated second weight, the sum of the power consumption of spectrum efficiency and energy efficiency is multiplied by the second weight to obtain a weighted sum of power consumption.

[0081] Finally, the sum of the weighted utility functions is combined with the sum of the weighted power consumption to obtain the final function to be optimized. This function comprehensively considers both the utility and power consumption of the network and can be used to guide the decision-making process of resource allocation and power optimization.

[0082] In step 104, each of the macro base station and each of the small base stations is used as an agent to repeatedly allocate resource blocks to associated users and determine the transmit power on different resource blocks.

[0083] Select each base station in the network as an agent and define the agent's actions, states, and rewards as follows:

[0084] The agent's actions determine the user association and resource allocation strategy. Therefore, the agent needs to select associated users and allocate resource blocks and transmit power on these resource blocks for each associated user. The transmit power of macro and small base stations on resource blocks is discretized into H levels and X levels, respectively. Each agent's action set consists of a user association action set, a resource block allocation action set, and a transmit power allocation action set, expressed as:

[0085]

[0086] Among them, A0 represents the action set with macro base station B0 as the agent, A n Indicates that the small base station B n is the set of actions of the agent.

[0087] Step 105, when the transmission power corresponding to each user meets or exceeds the minimum rate requirement, the weights in the function to be optimized are obtained according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value, thereby obtaining the target optimization function.

[0088] The macro base station and each small base station are used as agents respectively. Each agent repeatedly performs the actions of selecting associated users and allocating resource blocks and transmission power on the resource blocks to the associated users. After each agent completes the action, if the rates of all users in the network meet the minimum rate requirement, the function to be optimized is used as the reward function of each agent, the state is set to 1, and the weights in the function to be optimized are obtained according to the weighted square of the distance between each target value and its maximum target value under the state-action. Otherwise, both the reward and the state are set to 0 until the optimal action strategy is obtained.

[0089] In some embodiments, when calculating the utility function, the impact of other base stations on the transmission utility between the user and the associated base station needs to be considered. Figure 2 , Figure 2 A flowchart of a resource allocation method in a network provided by an embodiment of the present disclosure includes:

[0090] Step 201 : Calculate the received power between the user and the associated base station and the co-channel interference power and noise interference power between the user and the non-associated base stations.

[0091] Represents the user association variable, if base station B n Associated User U m ,but Otherwise Represents the resource block allocation variable. If user U m Occupied resource block R k ,but otherwise Indicates B n In R k The transmit power on .

[0092] If user U m Associated base station B n And occupies resource block R k , then U m The received power is:

[0093]

[0094] in, Indicates B n In R k Up to U m The channel gain includes path fading and shadow fading.

[0095] Since all base stations in the network reuse all resource blocks, if U m Associated base station B n And occupies R k , then Um In R k The co-channel interference power from other base stations is:

[0096]

[0097] In addition to the co-channel interference from other base stations, U m In R k The power is also affected by noise interference:

[0098]

[0099] in, is the noise power spectral density.

[0100] Step 202: Calculate the signal-to-interference-plus-noise ratio (SINR) based on the received power, co-channel interference power, and noise interference power, and calculate the downlink rate of the user based on the SINR.

[0101] According to the definition of SINR, U m Associated base station B n When R k The SINR on is:

[0102]

[0103] U m Multiple resource blocks can be allocated, and the achieved downlink rate is:

[0104]

[0105] Step 203: Calculate the user's utility function based on the downlink rate and a preset fairness factor.

[0106] In order to make resource allocation among users more fair, the present invention uses a utility function that characterizes user fairness, namely the β-utility function:

[0107]

[0108] Where β is the fairness factor, β∈[0,∞). There are three special cases for the value of β: if β=0, that is, user fairness is not considered, and the β-utility function represents the user rate; if β=1, formula (6) represents proportional fairness, and the β-utility function is the logarithm of the user rate; if β=∞, absolute fairness between users can be achieved. Therefore, as the fairness factor increases, user fairness increases, which means that resources in the network are distributed in a more fair manner, but at the expense of network capacity. Different fairness factor values ​​can be selected for different optimization objectives. For example, if only network capacity is optimized, the fairness factor is set to zero, that is, β=0; if the fairness of users in the network is considered, β>0 is selected, and different fairness factors are set according to the requirements for network performance and user fairness. Therefore, the size of the β value represents the size of user fairness. The relationship between network performance and user fairness can be adjusted by adjusting the fairness factor to achieve a relative balance between network performance and user fairness.

[0109] See also Figure 3 , Figure 3 A schematic diagram of a flow chart of a resource allocation method in a network provided in an embodiment of the present disclosure includes:

[0110] Step 301: Count the downlink rate of each user to obtain the system capacity of the network model.

[0111] The system capacity is the sum of the rates of all users, expressed as:

[0112]

[0113] Step 302: Count the power consumption of the base stations in the network model to obtain the total power consumption of the network model.

[0114] The total power consumption of each base station consists of two parts: total transmit power and static circuit power. The total power consumption of the system, that is, the power consumption of all base stations, is expressed as:

[0115]

[0116] Among them, P MC and P SC are the static circuit power of macro base stations and small base stations respectively.

[0117] Step 303: Count the bandwidth of each user to obtain the total bandwidth of the network model.

[0118] The total bandwidth consumed in the system is the sum of the bandwidth occupied by all users and can be expressed as:

[0119]

[0120] Among them, if but Represents R k Occupied, if but That is R k Not occupied.

[0121] Step 304: Calculate the energy efficiency of the network model based on the system capacity and the total power consumption.

[0122] The energy efficiency of the system is:

[0123]

[0124] Step 305: Calculate the spectrum efficiency of the network model according to the system capacity and the total bandwidth.

[0125] The spectrum efficiency is:

[0126]

[0127] In some embodiments, the calculating the function to be optimized based on the sum of the utility functions of each user and the first weight and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight further includes:

[0128] Based on a linear weighted method, the sum of the utility functions of the respective users and the system power consumption are combined into the function to be optimized.

[0129] The present invention uses the β-utility function to characterize user fairness, optimizes the system spectrum efficiency by maximizing the sum of all users' β-utility functions, and optimizes the system energy efficiency by minimizing system power consumption. Under the premise of ensuring the minimum rate requirement of users, in order to achieve a balance between system energy efficiency and spectrum efficiency and user fairness, the present invention uses a linear weighted method to combine the sum of all users' β-utility functions and system power consumption into a utility function, which is characterized as Where W SE and W EE is a priority parameter, and the user association, spectrum allocation, and power allocation problems in the network are modeled as optimization problems with this utility function as the optimization objective.

[0130] The optimization problem is characterized as:

[0131]

[0132] Among them, C1 and C2 represent the user association and spectrum allocation decision variables in the optimization problem respectively; C3 means that each user can only be associated with one base station; C4 means that all users are guaranteed to be associated; C5 means that each user can be allocated multiple resource blocks; C6 means that different resource blocks are allocated between users associated with the same base station; C7, C8 and C9 mean that the base station's transmit power on the resource block is positive and the total transmit power cannot exceed its maximum total transmit power, where P MT and P ST They are the maximum total transmit power of macro base stations and small base stations respectively; C10 means that the rate of each user cannot be lower than the minimum rate, which guarantees the user's minimum rate requirement.

[0133] In some embodiments, during optimization, it is necessary to balance network energy efficiency, spectrum efficiency, and user fairness; see Figure 4 , Figure 4 A schematic diagram of a flow chart of a resource allocation method in a network provided in an embodiment of the present disclosure includes:

[0134] Step 401: When the transmission power corresponding to each user meets or exceeds the minimum rate requirement, a reward value is determined to be a first preset value.

[0135] This paper proposes a fairness-oriented user association and resource allocation method in dense heterogeneous networks to solve the above-mentioned optimization problem. For the downlink user association, spectrum allocation, and power allocation problems in the dense heterogeneous network under study, each base station in the network is selected as an agent, and the agent's actions, states, and rewards are defined. The details are as follows:

[0136] The agent's actions determine the user association and resource allocation strategy. Therefore, the agent needs to select associated users and allocate resource blocks and transmit power on these resource blocks for each associated user. The transmit power of macro and small base stations on resource blocks is discretized into H levels and X levels, respectively. Each agent's action set consists of a user association action set, a resource block allocation action set, and a transmit power allocation action set, expressed as:

[0137]

[0138] Among them, A0 represents the action set with macro base station B0 as the agent, A n Indicates that the small base station B n is the set of actions of the agent.

[0139] The method proposed in this invention jointly optimizes system energy efficiency and spectrum efficiency while taking into account user fairness under the premise of ensuring the minimum rate requirement of each user. Therefore, in order to ensure the minimum rate requirement of each user, the state of the agent is defined according to whether the user association and resource allocation actions it performs meet the minimum rate requirement of each user. If the user association and resource allocation actions performed by the agent meet the minimum rate requirement of each user, the new state s' n =1, otherwise s n '=0.

[0140] Step 402: Calculate a first priority parameter and a second priority parameter according to the reward value.

[0141] In the learning process of agents, the reward function plays an important role. To be consistent with the optimization goal defined by formula (12), the present invention sets the reward function of each agent to be:

[0142]

[0143] The reward function is defined according to the optimization goal of the system and the minimum rate requirement of each user, which makes the agent's reward reasonable and efficient.

[0144] Through the above definition of agent state and reward, it can be seen that each agent obtains the same state and reward after performing an action, that is, all agents share the same state and reward. The state of each agent indicates whether the rates of all users in the network meet the minimum rate requirement after all agents perform user association and resource allocation actions. After the agent performs user association and resource allocation actions, if the rates of all users meet the minimum rate requirement, the reward given to the agent is the optimization objective of the optimization problem of the present invention, otherwise the reward is 0. The optimization objective of the problem of the present invention is to maximize the sum of the β-utility functions of all users and minimize the system power consumption. All agents share the same state and reward. Therefore, all agents continuously interact and learn to jointly explore the optimal action strategy to achieve the joint optimization of system energy efficiency and spectrum efficiency and ensure user fairness.

[0145] For W EE and W SE To avoid the influence of subjective decision on the value of fixed weight, the present invention uses the weighted square of the distance d between each target value and its maximum target value under the state-action as the evaluation criterion:

[0146]

[0147] Since the sum of the user's β-utility function and the system power consumption are not in the same dimension, in order to reflect the fairness of different objectives, the objective values ​​are first normalized:

[0148]

[0149] Among them, the minimum target value is mapped to 0 and the maximum target value is mapped to 1. Therefore, to calculate the weight, the weight coefficient under this state-action can be objectively obtained by solving the following optimization problem:

[0150]

[0151] st W SE +W EE =1 (19)

[0152] Step 403: Calculate the reward function of each base station according to the first priority parameter and the second priority parameter.

[0153] Step 404 : Update the action-value function according to a preset learning rate and a preset discount rate.

[0154] Each base station acts as an agent to associate users and allocate resource blocks and transmit power on resource blocks for its associated users. It collects user association and resource allocation information and channel information in the network through interaction with the system environment, calculates the rate of its associated users, and obtains corresponding rewards based on whether the rates of all users in the network meet the minimum rate requirements. In the learning process, the agent attempts to select strategy π through user association and resource allocation actions. n (s n ) to maximize its discounted reward, all agents hope to find their own optimal Q value. n is the state-action value function of the agent The update formula is:

[0155]

[0156] According to formula (20), after performing an action, the agent will update its Q table by the Q value of the current state, the reward, and the maximum Q value of the next state.

[0157] The agent’s optimal state-action value function is defined as:

[0158]

[0159] In state s n Under this condition, the corresponding optimal action selection strategy for:

[0160]

[0161] Step 405: Repeat the above calculation process until the number of calculations reaches a preset calculation threshold; and obtain the target optimization function.

[0162] If the rate per user Both meet the minimum rate requirement R min , then for each agent: state s n '=1、Calculate W SE and W EE ,award Update Q n (s n ,a n ), update status s n =s n ', repeat steps step3-step6 until the maximum number of learning times is completed, otherwise for each agent: state s n '=0、Reward R n (s n ,a n )=0, update Q n (s n ,a n ), update status s n =s n ', repeat step 5-step 6 until the rate of each user is Both meet the minimum rate requirement R min .

[0163] In order to achieve the joint optimization of network energy efficiency and spectrum efficiency while taking into account user fairness, the present invention proposes a fairness-oriented user association and resource allocation method in a dense heterogeneous network. In order to quantify the results of the above method, that is, the energy efficiency, spectrum efficiency and user fairness achieved by the method, the definitions of energy efficiency and spectrum efficiency are respectively as formulas (10) and (11). For user fairness, the present invention uses Jain's index to measure it. The larger the value, the greater the user fairness, which is defined as:

[0164]

[0165] The user association and resource allocation method of the present invention is distributed, that is, the macro base station and each small base station are defined as an agent, and each agent obtains the same state and reward after performing an action. Its advantage is that all agents share the same state and reward, and through continuous interactive learning, jointly explore the optimal action strategy to achieve joint optimization of system energy efficiency and spectrum efficiency and ensure user fairness.

[0166] Corresponding to the above-mentioned resource allocation method in a network, the present invention further provides a resource allocation device in a network. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment and will not be repeated in the present invention.

[0167] Figure 5 A schematic diagram of a resource allocation device in a network provided by an embodiment of the present disclosure is shown in FIG. Figure 5 Shown, including:

[0168] The acquisition unit 51 is configured to acquire a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; and the user is associated with one base station;

[0169] an allocation unit 52, configured to allocate different resource blocks occupied by each base station to different users having an associated relationship, and construct a utility function representing each user according to the downlink rate of each user;

[0170] A first calculation unit 53 is configured to calculate a function to be optimized based on the sum of the utility functions of each user and a first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and a second weight; wherein the spectrum efficiency and energy efficiency are calculated based on the network model; and the first weight and the second weight are calculated in advance;

[0171] a determining unit 54, configured to use each of the macro base station and each of the small base stations as agents, repeatedly allocating resource blocks to associated users, and determining transmit power on different resource blocks;

[0172] The second calculation unit 55 is used to obtain the weights in the function to be optimized according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value when the transmission power corresponding to each user meets the requirement of being greater than or equal to the minimum rate, so as to obtain the target optimization function.

[0173] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the allocation unit 52 is further configured to:

[0174] Calculate the received power between the user and the associated base station and the co-channel interference power and noise interference power between the user and the non-associated base station;

[0175] Calculating a signal-to-interference-plus-noise ratio (SINR) based on the received power, the co-channel interference power, and the noise interference power, and calculating a downlink rate of the user based on the SINR;

[0176] The utility function of the user is calculated according to the downlink rate and a preset fairness factor.

[0177] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the device also includes:

[0178] a statistical unit 56 configured to calculate, before the first calculating unit 53 calculates the function to be optimized based on the sum of the utility functions of each user and the first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight, statistically calculate the downlink rate of each user to obtain the system capacity of the network model;

[0179] The statistical unit 57 is further configured to count the power consumption of the base stations in the network model to obtain the total power consumption of the network model;

[0180] The statistics unit 58 is further configured to count the bandwidth of each user to obtain the total bandwidth of the network model;

[0181] A third calculation unit 59 is configured to calculate the energy efficiency of the network model according to the system capacity and the total power consumption;

[0182] The fourth calculation unit 510 is configured to calculate the spectrum efficiency of the network model according to the system capacity and the total bandwidth.

[0183] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the first calculation unit 53 is further configured to:

[0184] Based on a linear weighted method, the sum of the utility functions of the respective users and the system power consumption are combined into the function to be optimized.

[0185] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the second calculation unit 55 is further configured to:

[0186] When the transmission power corresponding to each of the users meets or exceeds the minimum rate requirement, determining the reward value to be a first preset value;

[0187] Calculating a first priority parameter and a second priority parameter according to the reward value;

[0188] Calculating a reward function for each base station according to the first priority parameter and the second priority parameter;

[0189] Update the action value function according to the preset learning rate and the preset discount rate;

[0190] Repeat the above calculation process until the number of calculations reaches the preset calculation threshold; and obtain the target optimization function.

[0191] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principles are the same, which is no longer limited in the embodiment of the present disclosure.

[0192] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0193] Figure 7 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0194] like Figure 7 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.

[0195] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0196] The computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the resource allocation method in a network. For example, in some embodiments, the resource allocation method in a network can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the aforementioned resource allocation method in the network in any other appropriate manner (for example, by means of firmware).

[0197] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0198] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0199] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0200] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0201] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0202] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0203] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0204] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0205] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A resource allocation method in a network, characterized in that: include: Acquire a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; the user is associated with one base station; Allocate the different resource blocks occupied by each base station to different users with related relationships, and construct a utility function to characterize each user based on the downlink rate of each user; Calculating a function to be optimized based on the sum of the utility functions of each user and a first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and a second weight; wherein the spectrum efficiency and energy efficiency are calculated based on the network model; and the first weight and the second weight are calculated in advance; Each of the macro base station and each of the small base station is used as an agent to repeatedly allocate resource blocks to associated users and determine the transmit power on different resource blocks; When the transmit power corresponding to each of the users satisfies a requirement of being greater than or equal to a minimum rate, obtaining each weight in the function to be optimized according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value, thereby obtaining a target optimization function; Wherein, when the transmit power corresponding to each of the users meets or exceeds the minimum rate requirement, the weights in the function to be optimized are obtained according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value, and the target optimization function is obtained, comprising: When the transmission power corresponding to each of the users meets or exceeds the minimum rate requirement, determining the reward value to be a first preset value; Calculating a first priority parameter and a second priority parameter according to the reward value; Calculating a reward function for each base station according to the first priority parameter and the second priority parameter; Update the action value function according to the preset learning rate and the preset discount rate; Repeat the above calculation process until the number of calculations reaches the preset calculation threshold; and obtain the target optimization function.

2. The method according to claim 1, characterized in that The method of respectively allocating different resource blocks occupied by each base station to different users having an associated relationship and constructing a utility function representing each user according to the downlink rate of each user includes: Calculate the received power between the user and the associated base station and the co-channel interference power and noise interference power between the user and the non-associated base station; Calculating a signal-to-interference-plus-noise ratio (SINR) based on the received power, the co-channel interference power, and the noise interference power, and calculating a downlink rate of the user based on the SINR; The utility function of the user is calculated according to the downlink rate and a preset fairness factor.

3. The method according to claim 2, characterized in that Before calculating the function to be optimized based on the sum of the utility functions of each user and the first weight and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight, the method includes: Counting the downlink rate of each user to obtain the system capacity of the network model; Counting the power consumption of the base stations in the network model to obtain the total power consumption of the network model; Counting the bandwidth of each user to obtain the total bandwidth of the network model; Calculating the energy efficiency of the network model according to the system capacity and the total power consumption; The spectrum efficiency of the network model is calculated according to the system capacity and the total bandwidth.

4. The method according to any one of claims 1 to 3, characterized in that The function to be optimized is calculated based on the sum of the utility functions of each user and the first weight and the sum of the power consumption of spectrum efficiency and energy efficiency and the second weight, including: Based on a linear weighted method, the sum of the utility functions of the respective users and the system power consumption are combined into the function to be optimized.

5. A resource allocation device in a network, characterized in that: include: An acquisition unit is configured to acquire a network model; wherein the network model includes at least one base station and at least one user; the base stations include a macro base station and at least one small base station; and the user is associated with one base station; An allocation unit is used to allocate different resource blocks occupied by each base station to different users with associated relationships, and construct a utility function representing each user according to the downlink rate of each user; a first calculation unit, configured to calculate a function to be optimized based on the sum of the utility functions of each user and a first weight, and the sum of the power consumption of spectrum efficiency and energy efficiency and a second weight; wherein the spectrum efficiency and energy efficiency are calculated based on the network model; and the first weight and the second weight are calculated in advance; a determining unit, configured to use each of the macro base station and each of the small base stations as agents, repeatedly allocating resource blocks to associated users, and determining transmit power on different resource blocks; A second calculation unit is configured to obtain, when the transmit power corresponding to each user satisfies a minimum rate requirement or greater than or equal to the minimum rate requirement, each weight in the function to be optimized according to the weighted square of the distance between each target value corresponding to the target state and its maximum target value, thereby obtaining a target optimization function; The second computing unit is further configured to: When the transmission power corresponding to each of the users meets or exceeds the minimum rate requirement, determining the reward value to be a first preset value; Calculating a first priority parameter and a second priority parameter according to the reward value; Calculating a reward function for each base station according to the first priority parameter and the second priority parameter; Update the action value function according to the preset learning rate and the preset discount rate; Repeat the above calculation process until the number of calculations reaches the preset calculation threshold; and obtain the target optimization function.

6. The device according to claim 5, characterized in that The allocation unit is further configured to: Calculate the received power between the user and the associated base station and the co-channel interference power and noise interference power between the user and the non-associated base station; Calculating a signal-to-interference-plus-noise ratio (SINR) based on the received power, the co-channel interference power, and the noise interference power, and calculating a downlink rate of the user based on the SINR; The utility function of the user is calculated according to the downlink rate and a preset fairness factor.

7. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

9. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Heterogeneous network resource allocation method based on reinforcement learning

    CN112351433A

  • Satellite-ground collaborative network slice resource allocation method based on reinforcement learning

    CN117858256A