Method and apparatus for resource joint allocation in heterogeneous network with hybrid access

By using reinforcement learning methods to optimize the allocation of communication and computing resources in heterogeneous networks with mixed human-machine-thing access, the problem of poor network performance and high cost caused by poor resource allocation is solved. This approach minimizes the total system overhead and meets QoS requirements, thereby improving network performance and cost-effectiveness.

CN116567721BActive Publication Date: 2025-11-21BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211660667.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-11-21
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

In heterogeneous networks with mixed human-machine-thing access, existing technologies have failed to effectively address the problem of poor network performance and high service costs caused by inadequate resource allocation, especially in mobile edge computing scenarios assisted by drones. How to reasonably allocate limited communication and computing resources to meet the QoS requirements of different devices is a key challenge.

Method used

A joint communication-computation resource allocation method based on reinforcement learning is adopted. By defining the agent's state set, action set, and reward function, the state set is traversed to update the reward, and the maximum reward value is found, thereby optimizing resource allocation, minimizing the total system overhead, and ensuring the QoS requirements of each device.

Benefits of technology

In complex interference/interconnection environments, optimizing channel allocation, power allocation, and computing resource allocation minimizes total system overhead, improves network performance, reduces service costs, and adapts to heterogeneous networks with mixed human-machine-thing access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116567721B_ABST
    Figure CN116567721B_ABST
Patent Text Reader

Abstract

The application discloses a resource joint allocation method and device in a human-machine-object hybrid access heterogeneous network. The method comprises the following steps: determining constraint conditions, decision variables and optimization targets of devices in the human-machine-object hybrid access heterogeneous network; defining a state set, an action set and a reward function of an intelligent agent in the devices based on the determined constraint conditions, decision variables and optimization targets; traversing the state set of the intelligent agent based on the defined state set, action set and reward function, updating the income of the current state of the intelligent agent according to the reward value of the current state of the intelligent agent and the income estimation of the next state until the maximum reward value is found; and obtaining a resource allocation scheme that optimizes the optimization target based on the maximum reward value. The application solves the technical problems of poor network performance and high service cost caused by poor resource allocation in the network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of communication, in particular to a resource joint allocation method and device in a heterogeneous network of human-machine-object mixed access. BACKGROUND

[0002] With the emergence of the Internet of Things, future communication devices are not only mobile communication of human-type devices, but also communication interconnection between many kinds of machines and objects. Various types of sensors, vehicles, smart furniture, smart appliances, etc. can all be used as communication objects in the Internet of Things, and have generated diversified application scenarios such as smart home, smart transportation, smart shopping, smart health, etc. Therefore, coexistence of cellular networks and the Internet of Things will be a typical communication scenario in future communications, and must support communication requirements under human-machine-object mixed access networks. When human-type devices and object-type devices communicate, the communication rates, latency, and reliability of different devices and different services corresponding to different services quality (QoS) requirements are different. For example, virtual reality (VR) / augmented reality (AR) / mixed reality (MR) requires low latency and high reliability from the communication network, while some real-time data uploading sensors require stable transmission rate support. Therefore, it is necessary to support data communication of various types of devices under human-machine-object mixed access.

[0003] The massive data generated by the above devices urgently need to be processed, which requires higher performance information infrastructure. In traditional communication networks, information needs to be transmitted to remote cloud data centers for processing. With the increase of real-time information processing tasks, the defects of traditional communication networks such as high latency and backhaul overload become more and more obvious. Recently, the concept of deploying distributed servers / computing nodes at the edge of the wireless network, i.e. mobile edge computing (MEC), has been proposed. However, traditional edge nodes are usually deployed on the ground and are relatively fixed in position, making it difficult to meet the needs of fluctuations in requests and user mobility through static deployment. If the edge nodes are deployed in a super-dense manner, the cost of infrastructure construction will be greatly increased, and during off-peak hours, idle edge nodes will cause resource waste. Therefore, mobile air edge nodes provide a new solution. Unmanned aerial vehicles (UAVs) have the characteristics of low cost, high mobility, and high flexibility, which provide the possibility of covering hotspots and offloading traffic. In hover mode, UAVs can serve as stable air platforms to perform tasks, and if they carry MEC servers, they can be expected to solve the problem of effective deployment of MEC. As mobile edge nodes, UAVs simplify the cumbersome deployment of fixed edge nodes; their hovering stability and line-of-sight transmission characteristics can provide users with reliable low-latency communication links. In a multi-UAV assisted MEC network, edge nodes can be closer to active users, thereby providing higher service quality and lower latency.

[0004] However, UAV-aided MEC services also face many challenges. The UAV not only needs a large amount of energy to maintain its own flight. It also needs to provide part of the energy to the communication and computing units on the UAV to provide reliable data transmission and processing services. Since the size of the battery on the UAV is limited, how to reasonably allocate limited communication and computing resources to users while guaranteeing the different QoS requirements of human-machine-objects is particularly important. Communication resources mainly include channel allocation, power control, interference control, etc. It is necessary to study a reasonable communication-computing resource joint allocation algorithm to balance network performance and service cost.

[0005] At present, no effective solution has been proposed for the above problems. SUMMARY

[0006] Embodiments of the present application provide a resource joint allocation method and device in a human-machine-object hybrid access heterogeneous network, to at least solve the technical problems of poor network performance and high service cost caused by poor resource allocation in the network.

[0007] According to an aspect of embodiments of the present application, a resource joint allocation method in a human-machine-object hybrid access heterogeneous network is provided, comprising: determining constraint conditions, decision variables and optimization objectives of optimizing devices in the human-machine-object hybrid access heterogeneous network; defining a state set, an action set and a reward function of an agent in the devices based on the determined constraint conditions, decision variables and optimization objectives; based on the defined state set, action set and reward function, traversing the state set of the agent, updating the revenue of the current state of the agent according to the reward value of the current state of the agent and the revenue estimate of the next state until the maximum reward value is found; based on the maximum reward value, obtaining a resource allocation scheme that optimizes the optimization objectives.

[0008] According to another aspect of embodiments of the present application, a resource joint allocation device in a human-machine-object hybrid access heterogeneous network is also provided, comprising: a determination module configured to determine constraint conditions, decision variables and optimization objectives of optimizing devices in the human-machine-object hybrid access heterogeneous network; a definition module configured to define a state set, an action set and a reward function of an agent in the devices based on the determined constraint conditions, decision variables and optimization objectives; an update module configured to, based on the defined state set, action set and reward function, traverse the state set of the agent, update the revenue of the current state of the agent according to the reward value of the current state of the agent and the revenue estimate of the next state until the maximum reward value is found; and an allocation module configured to, based on the maximum reward value, obtain a resource allocation scheme that optimizes the optimization objectives.

[0009] According to a further aspect of the embodiments of the present application, a human-machine-thing hybrid access heterogeneous network is also provided, comprising the following devices: a macro base station, a small base station, a UAV, a human type device, an Internet of Things device, and a mobile edge computing server, wherein the mobile edge computing server comprises the resource joint allocation apparatus in the human-machine-thing hybrid access heterogeneous network as described above, and the intelligent agent comprises the UAV and the small base station.

[0010] In the embodiments of the present application, based on the defined state set, action set and reward function, the state set of the intelligent agent is traversed, the reward of the current state of the intelligent agent is updated according to the reward value of the current state of the intelligent agent and the reward estimation of the next state, until the maximum reward value is found, thereby solving the technical problems of poor network performance and high service cost caused by poor resource allocation in the network. BRIEF DESCRIPTION OF DRAWINGS

[0011] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their descriptions serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0012] Figure 1 is a flowchart of a resource joint allocation method in a human-machine-thing hybrid access heterogeneous network according to an embodiment of the present application;

[0013] Figure 2 is a flowchart of another resource joint allocation method in a human-machine-thing hybrid access heterogeneous network according to an embodiment of the present application;

[0014] Figure 3 is a structural schematic diagram of a human-machine-thing hybrid access heterogeneous network according to an embodiment of the present application;

[0015] Figure 4 is a comparison schematic diagram of the number of Internet of Things devices and the average energy consumption of SBS according to an embodiment of the present application;

[0016] Figure 5 is a comparison schematic diagram of the computing resource of a UAV and the average time delay of an Internet of Things device according to an embodiment of the present application;

[0017] Figure 6 is a comparison schematic diagram of the number of human type devices and the average time delay of an Internet of Things device according to an embodiment of the present application. DETAILED DESCRIPTION

[0018] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should be within the scope of protection of the present application.

[0019] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0020] Embodiment 1

[0021] According to the embodiments of the present application, a resource joint allocation method in a human-machine-object hybrid access heterogeneous network is provided, as shown in the formula (1), the method comprises the steps of: Figure 1

[0022] Step S102, determining the constraint conditions, decision variables and optimization objectives of the devices in the human-machine-object hybrid access heterogeneous network;

[0023] The devices in the heterogeneous network include the human type devices, the Internet of Things devices, the small base stations, the macro base stations and the unmanned aerial vehicles; the intelligent agents include the small base stations and the unmanned aerial vehicles.

[0024] First, the heterogeneous network is initialized. For example, the channel allocation, the transmission power allocation and the CPU cycle number allocation are initialized, and the prior information for determining the optimization objectives and the constraint conditions is calculated according to the initial values obtained by the initialization; a revenue table is created for each intelligent agent, which represents the revenue corresponding to the state and action of the corresponding intelligent agent, wherein the revenue table can be continuously updated according to the reward value of the current state of the corresponding intelligent agent and the maximum revenue in the next state.

[0025] ​Then, the constraint conditions are determined. For example, the constraint condition of the quality of service requirement of the human type device in the heterogeneous network is modeled as a minimum transmission rate constraint; the constraint condition of the quality of service requirement of the Internet of Things device in the heterogeneous network is modeled as a maximum transmission power constraint; the constraint condition of the quality of service requirement of the small base station (SBS) in the heterogeneous network is modeled as a user association number constraint; and the constraint condition of the quality of service requirement of the unmanned aerial vehicle (UAV) in the heterogeneous network is modeled as a user association number constraint, a maximum allowed power consumption constraint, and a computing capability constraint.

[0026] Then, the optimization target is determined. For example, the signal-to-interference-and-noise ratio of a user when the user uses a subchannel in the agent is calculated, and the data transmission rate of the user when transmitting data to the agent on the subchannel is calculated based on the signal-to-interference-and-noise ratio; the transmission delay and the transmission energy consumption of the user are calculated based on the data transmission rate; the computing delay required by the task and the computing energy consumption required by the task are calculated; the overhead of the user is calculated based on the transmission delay, the transmission energy consumption, the computing delay required by the task, and the computing energy consumption required by the task; and the overhead of each user is minimized to determine the optimization target; wherein the user includes a human type device and / or an Internet of Things device.

[0027] Step S104, based on the determined constraint conditions, decision variables, and optimization target, defining a state set, an action set, and a reward function of an agent in the device;

[0028] For example, the state set is constructed based on the input data size R required by the user to complete the task and the CPU cycle number D required by the user to complete the task in the decision variable; the action set is constructed based on the channel allocation matrix Γ, the power allocation p, and the CPU cycle number allocation in the decision variable; the reward function is constructed based on the minimum transmission rate constraint the maximum transmission delay constraint the maximum power consumption constraint and the computing capability constraint and the optimization target.

[0029] Step S106, based on the defined state set, action set, and reward function, traversing the state set of the agent, updating the reward of the current state of the agent according to the reward value of the current state of the agent and the reward estimation of the next state until the maximum reward value is found.

[0030] For example, the following steps are executed in a loop until the maximum reward value is found:

[0031] 1) After the agent selects an action based on the current state, the reward value of the current state of the agent is calculated according to the feedback of the current network environment;

[0032] For example, after the agent randomly selects an action based on the current state or selects an action with the maximum return according to a learning strategy, the agent shares the selected action, the current state, and the corresponding return to the heterogeneous network, wherein the corresponding return is determined based on the current state and the selected action; and a reward value of the current state of the agent is calculated according to the feedback of resource usage of all devices in the heterogeneous network, so as to perform distributed cooperation with other devices in the heterogeneous network.

[0033] 2) Obtain a return estimate of a next state of the agent, and update the return of the current state according to the return estimate and the reward value of the current state;

[0034] For example, a next state of the agent is obtained; an action with the maximum reward in the next state is selected to obtain a return estimate of the next state; and the return of the current state is updated according to the return estimate of the next state.

[0035] 3) From the current state to a next state, and taking the next state as the current state.

[0036] In step S108, a resource allocation scheme that optimizes the optimization target is obtained based on the maximum reward value.

[0037] The scenarios to which the resource allocation algorithm of the prior art is directed can be roughly divided into the following categories:

[0038] I. Only considering coexistence of things and machines. That is, the unmanned aerial vehicle assisted ground network only considers Internet of Things devices, and does not consider coexistence of human type devices and thing type devices. For example, in a disaster scenario, unmanned aerial vehicle assisted communication cannot be used in a ground network.

[0039] II. Only considering coexistence of people and machines. That is, the unmanned aerial vehicle assisted ground network only considers human type devices, and does not consider coexistence of human type devices and thing type devices. For example, unmanned aerial vehicle assisted cellular network communication.

[0040] III. Only considering coexistence of people and things. That is, only the communication of the ground network in the coexistence of human type devices and thing type devices is considered, and there is no unmanned aerial vehicle assisted communication. For example, H2H / M2M coexistence scenario.

[0041] These types of scenarios do not truly realize the coexistence of the three types of devices, and only single researches the resource allocation method under the coexistence of one or two types of devices. The network scenarios of type one and type two are too ideal, and few interference factors are considered, which is inconsistent with the reality. In the scenario of type three, when the number of user devices increases sharply or the task volume is very large, it is difficult for the ground network to support user demand. In addition, many existing technologies only consider communication resource management without allocating computing resources, which will also increase energy consumption and computing cost in the network.

[0042] The application provides a communication-computing resource joint allocation method based on reinforcement learning in a human-machine-object hybrid access heterogeneous network. In a network scenario where human type devices, object type devices, unmanned aerial vehicles and ground base stations coexist, communication resources and computing resources are jointly optimized. The application aims to minimize the total overhead (energy consumption and task completion delay weighting) by optimizing channel allocation, power allocation and computing resource allocation for users (human / object devices) in a complex interference / interconnection environment, and to ensure the QoS requirements of each device.

[0043] In the communication-computing resource joint allocation method provided by the application, the uplink of the human-machine-object hybrid access network is considered. The network includes human type devices, object type devices, unmanned aerial vehicles and ground base stations. The object type devices are delay-sensitive devices. Users can be connected to ground base stations or unmanned aerial vehicles, but each user can only select to access one base station. Users accessing the same cell base station do not have inter-cell interference, and users accessing different cells but working in the same channel have adjacent cell interference. The maximum transmit power of human type devices and object type devices and the maximum computing resource constraints of unmanned aerial vehicles and ground base stations are different.

[0044] In the communication-computing resource joint allocation method provided by the application, the total energy consumption minimization problem is modeled as a multi-agent collaborative control problem. The ground base stations and unmanned aerial vehicles are regarded as agents. The actions, states and rewards of the agents are defined according to the constraint conditions, decision variables and optimization objectives of the optimization problem. After selecting actions, the agents share information and return the current reward value according to the network environment state. The agents cooperate in a distributed manner. In the process of interacting with the wireless network environment, the communication-computing resource allocation is continuously iteratively optimized according to the changes of the network environment, and the optimal solution of the optimization objective is found.

[0045] Embodiment 2

[0046] According to the embodiments of the application, a communication-computing resource joint allocation method based on reinforcement learning is provided, as shown in Figure 2 The method comprises the following steps:

[0047] Step S202, system initialization.

[0048] The channel allocation, transmit power allocation and CPU cycle number allocation are initialized, and the channel gain, SINR and other prior information are calculated according to the initial values; the SBS and UAV are regarded as agents, and a Q table is created for each agent and initialized; the environment state S is initialized t .

[0049] In step S204, the state set, action set and reward function are designed for the agent.

[0050] Each agent makes an action selection according to the current state, and establishes a system total cost optimization model. Specifically, the constraint conditions of transmission power and transmission delay and the QoS demand are formulated, wherein the QoS demand is modeled as follows according to the characteristics of different devices: the human type device usually has a high requirement for transmission rate, and thus the QoS demand thereof is modeled as a minimum transmission rate constraint; the Internet of Things device is limited in volume and has small energy, and in order to reduce the energy consumption thereof, the QoS demand thereof is modeled as a maximum transmission power constraint; when too many users access the same SBS, resource allocation will be insufficient, which reduces the user experience, and thus the QoS demand of the SBS is modeled as a user association quantity constraint; the UAV has small battery and limited resources carried, and has certain energy consumption in flight, and thus the QoS demand thereof is modeled as a user association quantity constraint, a maximum allowed power consumption constraint and a computing capability constraint.

[0051] In step S206, the reward value is calculated.

[0052] Based on the above optimization model, each agent randomly selects an action or selects an action with the maximum Q value according to a learning strategy, and then shares information, calculates the current reward value according to the current environment state, and then moves to the next state S t+1 , updates the Q table; and then iterates and circulates according to the process for training.

[0053] The state set is defined as follows:

[0054] S t ={R,D}

[0055] The R includes the input data size required by the user to complete a task, and the D includes the CPU cycle number required by the user to complete a task. The environment state changes over time.

[0056] A greedy learning strategy is adopted when selecting an action, and a greedy factor ε is introduced. When a random number x is less than ε, a random action is selected, and when x is greater than ε, the current optimal action is selected, and the action set is defined as:

[0057] A t ={Γ,p,f}

[0058] where Γ represents the channel allocation matrix, p ∈ {p1, p2,..., P max} represents the power allocation, f ∈ {f1, f2,..., F max} represents the CPU cycle number allocation, P max represents the maximum transmit power, F max represents the maximum computing power.

[0059] Different BSs are indexed by j ∈ {1, 2}, and BS1 and BS2 represent the UAV and the SBS, respectively. A Boolean variable is defined to represent whether user i is associated with BSj, if represents the association, represents the non-association. The number of BSs associated with a user is limited to 1, i.e., given subchannel assignment variable is used to represent whether subchannel x is allocated to user i, represents the allocation, represents the non-allocation.

[0060] The signal-to-interference-plus-noise ratio (SINR) of user i using subchannel x in BSj is:

[0061]

[0062] where is the channel gain of user i to BSj on subchannel x, and Pi represents the transmit power of user i, represents the association of user i with BSj, i represents the user, j represents the base station, i' represents other users except user i, j' represents other base stations except base station j, represents the subchannel allocation of user i', P i′ represents the transmit power of user i', represents the channel gain of user i' to BSj on subchannel x, represents the subchannel allocation of MUEd, Pd represents the transmit power of MUEd, represents the channel gain of MUEd to the macro base station on subchannel x, d represents the MUE, and D represents the set of all MUEs. The second term in the denominator represents the inter-cell interference of other cell co-channel users, and the third term represents the interference of MUEs served by the MBS, σ 2 is the additive white Gaussian noise power.

[0063] The data transmission rate of user i to BSj on subchannel x is:

[0064]

[0065] Where B represents the sub-channel bandwidth, This represents the signal-to-interference-plus-noise ratio (SIR) when user i uses subchannel x in BSj.

[0066] Without considering local user computation, tasks can only be uploaded to the SBS for computation or unloaded to a drone for execution. When user i uploads a task to SBSj for computation via subchannel x, the transmission delay is:

[0067]

[0068] Where Ri represents the size of the input data that user i needs to complete the task. This represents the data transmission rate of user i to BSj on subchannel x.

[0069] Transmission energy consumption is:

[0070]

[0071] Where Pi represents the transmit power of user i.

[0072] The latency required for base station computing tasks is:

[0073]

[0074] Among them, C j f is the CPU cycles required for a base station to process 1 bit of data. j The computing resources allocated to the base station, Di represents the number of CPU cycles required for user i to complete the task.

[0075] The energy consumption for task computation is:

[0076]

[0077] Among them, K j f represents the base station jCPU capacitance coefficient. j Di represents the computing resources allocated to the base station, Di represents the number of CPU cycles required for user i to complete the task, and Cj represents the number of CPU cycles required for the base station to process 1 bit of data.

[0078] Therefore, when all tasks are uploaded to SBS for processing, the overhead is:

[0079]

[0080] Where ω is the weighting factor for calculating latency and energy consumption. This represents the transmission energy consumption of user i uploading data to SBS. This represents the computational energy consumption of SBS for calculating user i's task. denotes the transmission latency of user i uploading to the SBS, denotes the computation latency of the SBS computing the task of user i;

[0081] When offloading the task to the UAV processing, the overhead is:

[0082]

[0083] wherein, denotes the transmission energy consumption of user i uploading to the UAV, denotes the computation energy consumption of the UAV computing the task of user i, denotes the transmission latency of user i uploading to the UAV, denotes the computation latency of the UAV computing the task of user i.

[0084] Therefore, for the ith user, its overhead is denoted as:

[0085]

[0086] wherein, denotes the overhead of user i uploading the task to the SBS processing.

[0087] So far, the optimization model of the present application is summarized as follows:

[0088]

[0089] s.t.C1:

[0090] C2:

[0091] C3:

[0092] C4:

[0093] C5:

[0094] C6:

[0095] C7:

[0096] C8:

[0097] C9:

[0098] C10:

[0099] where Γ, p and f represent the subchannel allocation, the transmit power allocation and the CPU cycle number allocation strategy respectively, S1 and S2 represent the maximum number of associated users of the UAV and the SBS respectively, represents the maximum power consumption of the UAV, and η n represents the minimum transmission rate of the human-type device; C1 represents the subchannel allocation constraint; C2 and C3 represent the association factor constraint; C4, C5 and C6 represent the user association, the computing capacity and the maximum power consumption constraint of the QoS requirement of the UAV respectively; C7 represents the QoS requirement constraint of the human-type device; C8 represents the maximum transmit power constraint of the user; C9 represents the maximum tolerable transmission delay constraint of the Internet of Things device; C10 represents the QoS requirement constraint of the SBS; x represents a subchannel, X represents a set of subchannels, i represents a user, I represents a set of users, S1 represents the maximum number of associated users of the UAV, fu represents the computing resource allocated by the UAV, ai represents whether the user i is associated with the UAV, N represents a set of human-type devices, Pm represents the transmit power of the Internet of Things device, Pn represents the transmit power of the human-type device, M represents a set of Internet of Things devices, and S2 represents the maximum number of associated users of the SBS.

[0100] In order to ensure the QoS requirements of various types of devices in the training process of each agent, the QoS constraint is added to the design of the reward function in the present application:

[0101] First, the QoS constraint of the human-type device, i.e., the minimum transmission rate constraint, is considered, and the reward and punishment mechanism designed by the present application is as follows:

[0102]

[0103] For the Internet of Things device, not only the QoS constraint needs to be considered, but also the maximum transmission delay constraint of the delay-sensitive device needs to be met, and the reward and punishment mechanism designed by the present application is as follows:

[0104]

[0105] where Rm represents the input data size required by the Internet of Things device m to complete a task, represents the maximum transmission delay that can be tolerated by the Internet of Things device m.

[0106] The QoS constraint of the UAV is more, and only the maximum power consumption constraint and the computing capacity constraint are considered in the reward function, and the reward and punishment mechanism designed by the present application is as follows:

[0107]

[0108] where a i represents whether the user i is associated with the UAV, and Pi represents the transmit power of the user i, represents the maximum power consumption of the UAV, fu represents the computing resource allocated to the UAV, represents the maximum computing capacity of the UAV.

[0109] The user association QoS constraint of the UAV and the SBS is not embodied in the reward function.

[0110] In addition, because the reward mechanism of the reinforcement learning is to maximize the probability expectation value of the cumulative sum of the scalar reward signals received by the agent, and the optimization target of the present application is to minimize the total system overhead, it is necessary to perform reciprocal processing on the optimization target, and introduce three harmonic coefficients C1, C2 and C3. In summary, the present application designs the reward function as:

[0111]

[0112] wherein C1, C2 and C3 are positive real numbers, used to balance the reward and punishment system, and Zt is the total overhead calculated. All agents use the same reward function.

[0113] In step S208, based on the maximum reward value, a resource allocation scheme that optimizes the optimization target is obtained.

[0114] The resource allocation method proposed in the embodiments of the present application is based on the Q-learning theory in reinforcement learning. In the same human-machine-object hybrid access network system scenario proposed in the present application, other machine learning methods can be used to achieve the same effect. For example, when performing communication-computing resource joint allocation, if the deep reinforcement learning theory is adopted and also based on the multi-agent distributed framework, this method can become an alternative scheme of the present application.

[0115] The resource joint allocation method provided by the embodiments of the present application minimizes the total system overhead to find the optimal communication resource and computing resource allocation scheme, and guarantees the different QoS requirements of various types of devices. In addition, the resource joint allocation method provided by the embodiments of the present application is based on the distributed multi-agent learning framework, which greatly relieves the load of the base station, not only improves the training convergence speed, but also improves the system performance.

[0116] Embodiment 3

[0117] According to the embodiments of the present application, a human-machine-object hybrid access heterogeneous network system is provided, as shown in Figure 3 which includes a macro base station (MBS), a small base station (SBS), a UAV, a human type device, an Internet of Things device, an MBS service user (MUE), an MEC server, and X orthogonal subchannels with a total bandwidth W and a subchannel bandwidth B.

[0118] Wherein, the Internet of Things devices are delay-sensitive devices; the SBS can establish a communication link with the user, allocate a channel, transmit power and computing resources for the user; the UAV is highly fixed, used for assisting the SBS communication, as an edge node, equipped with a MEC server, and can also establish a communication link with the user, allocate a channel, transmit power and computing resources for the user; each user can be connected with the SBS or the UAV, but can only select one base station for connection.

[0119] In the embodiment, the SBS is represented by , the UAV is represented by , the Internet of Things devices are represented by , and are all delay-sensitive nodes, the human-type devices are represented by , and the number of MUEs is represented by D, which does not participate in resource allocation and only acts as a fixed interference.

[0120] The SBS can establish a communication link with the user, allocate a channel, transmit power and computing resources for the user.

[0121] The UAV is highly fixed, used for assisting the SBS communication, as an edge node, equipped with a MEC server, and can also establish a communication link with the user, allocate a channel, transmit power and computing resources for the user.

[0122] Each user can be connected with the SBS or the UAV, but can only select one base station for connection. For any user The users are randomly divided into K clusters, I1, I2... IK. K The sets are mutually exclusive and exhaustive. The users accessing the same cell base station do not have inter-cell interference, and the users accessing different cells but working in the same channel have adjacent cell interference.

[0123] In this network, it is assumed that each user i has a task to be completed. For the Internet of Things device m, R m represents the size of the input data, D m represents the number of CPU cycles required for calculating the task, represents the maximum communication delay that can be tolerated by the task. For the human-type device n,

[0124] In the heterogeneous network provided by the embodiment, the joint resource allocation can be performed by the following method:

[0125] Firstly, the system is initialized, the channel allocation, transmit power allocation and CPU cycle number allocation are initialized, and the channel gain, SINR and other prior information are calculated according to the initial values; the SBS and the UAV are regarded as agents, and a Q table is created for each agent and initialized; the environment state is initialized.

[0126] Next, the state / action / reward function is designed for the agent, each agent makes action selection according to the current state, and a system total cost optimization model is established. Specifically, the constraint conditions of transmission power and transmission delay and QoS requirements are formulated, wherein the QoS requirements are modeled as follows according to the characteristics of different devices: the human type device usually has a high requirement for transmission rate, and therefore the QoS requirement of the human type device is modeled as a minimum transmission rate constraint; the Internet of Things device is limited in volume and has small energy, and in order to reduce the energy consumption of the Internet of Things device, the QoS requirement of the Internet of Things device is modeled as a maximum transmission power constraint; when too many users access the same SBS, resource allocation is insufficient, and user experience is reduced, and therefore the QoS requirement of the SBS is modeled as a user association quantity constraint; the UAV has small battery and limited resources, and has certain energy consumption in flight, and therefore the QoS requirement of the UAV is modeled as a user association quantity constraint, a maximum allowed power consumption constraint and a computing capability constraint.

[0127] Finally, based on the above optimization model, each agent randomly selects an action or selects an action with the maximum Q value according to a learning strategy, then information sharing is performed, a current reward value is calculated according to a current environment state, then the system is transferred to a next state, and a Q table is updated; then the training is performed through iteration and circulation according to the process.

[0128] The heterogeneous network in the embodiment can perform resource allocation by using the methods in embodiments 1 and 2, and therefore, how the heterogeneous network performs resource allocation is not described herein again.

[0129] The heterogeneous network system for human-machine-object mixed access provided by the application has rich types of devices, and various interference / interconnection relationships are considered, and truly realizes coexistence of human, machine and object.

[0130] Simulation experiment

[0131] The application is set as follows: the MBS coverage radius is 500 m, the SBS coverage radius is 125 m, the UAV height is 50 m, the coverage radius is 100 m, there is 1 MBS, 5 SBSs, 4 UAVs, and 0-25 Internet of Things devices and 0-20 human type devices randomly scattered in the MBS coverage range, 5 MUEs are set in the environment to act as fixed interference. The UAV selects the four SBSs with the densest users for assistance. The total system bandwidth is 10 MHz, there are 30 subchannels in total, and the channel gain is randomly generated according to the distance between the user and the base station. The maximum transmission power of the human type device is 32 dBm, the maximum transmission power of the object type device is 24 dBm, the maximum tolerable transmission delay of the Internet of Things device is 10 ms, and the minimum transmission rate of the human type device is 10 Mbps.

[0132] The channel allocation, power allocation, and CPU cycle number allocation algorithm proposed in the present application is named as: communication computing resource joint allocation algorithm for unmanned aerial vehicle assisted human and internet of things coexisting ground network (UAC-CCRA) according to its characteristics, and is compared with the following three resource allocation algorithms: (1) communication resource allocation algorithm for unmanned aerial vehicle assisted human and internet of things coexisting ground network (UAC-CRA), which only does not consider the allocation of computing resources on the algorithm provided in the present application; (2) communication computing resource joint allocation algorithm for unmanned aerial vehicle assisted internet of things (UAI-CCRA), which does not participate in resource allocation on the algorithm provided in the present application, and only acts as interference; (3) communication computing resource allocation algorithm for unmanned aerial vehicle assisted human and internet of things coexisting ground network (Non-UC-CCRA), which only does not join the unmanned aerial vehicle assistance on the algorithm provided in the present application.

[0133] The algorithm provided in the present application and the three comparative algorithms are sequentially named as schemes 1, 2, 3, and 4.

[0134] Referring to Figure 4 , the change of SBS average energy consumption with the increase of the number of internet of things devices accessing the system is introduced. As can be seen from the figure, with the increase of the number of internet of things devices, the SBS average energy consumption continuously rises. This is because when the number of internet of things devices is 0, the SBS only needs to allocate resources to human type devices, and the resources are sufficient, and each user can get the optimal allocation scheme, so the SBS average energy consumption is very small. And because of scheme 3, the system only provides resources for the internet of things, so when the number of internet of things devices is 0, the SBS average energy consumption is 0. With the increase of the number of users, the resources in the system are gradually tight, and it is not possible to make each user achieve the optimal allocation scheme, and it will also increase the co-channel interference and reduce the SINR, so the energy consumption gradually rises, and the rising rate becomes faster and faster.

[0135] As can be seen from the figure, the SBS average energy consumption of scheme 3 is the smallest, because the SBS and the unmanned aerial vehicle only provide resources for the internet of things devices, and the number of users served is less than that of other schemes, and the resources are more sufficient, so the performance of the SBS is better. The performance of scheme 4 is the worst, because there is no assistance of the unmanned aerial vehicle in the system, and the number of devices that each SBS needs to serve is more than that of other schemes, thereby reducing the performance of the SBS. By comparing the curves of scheme 2 and scheme 1, when the number of devices accessing is small, the difference between the two is not large, but with the increase of the number of devices, the energy consumption difference between the two becomes larger and larger, thereby reflecting the advantage of optimizing the allocation of computing resources.

[0136] Referring to Figure 5, introduce the change of the average time delay of the Internet of Things device with the increase of the computing resources of the unmanned aerial vehicle. As can be seen from the figure, with the increase of the computing resources of the unmanned aerial vehicle, the average time delay of the Internet of Things device of the other schemes except scheme 4 decreases, and the decreasing amplitude becomes smaller and smaller, and gradually becomes gentle. This is because there is no unmanned aerial vehicle assistance in the communication system of scheme 4, so the change of the computing resources of the unmanned aerial vehicle will not affect its performance. When the computing resources of the unmanned aerial vehicle are very small, that is, the available resources are very few, the time delay of the task uploaded to the unmanned aerial vehicle increases, so the average time delay of the Internet of Things device is high. With the increasing of the computing resources of the unmanned aerial vehicle, the computing time delay of the unmanned aerial vehicle also decreases, and when most of the computing resources allocated to the task are optimal, the influence of the increase of the computing resources of the unmanned aerial vehicle on the computing time delay is also smaller and smaller, so the decreasing amplitude of the average time delay of the Internet of Things device becomes smaller and smaller. Since scheme 3 does not share resources with the human type device, scheme 3 is still optimal in terms of performance for the Internet of Things device. The performance of scheme 1 is the best, which also reflects the advantage of increasing the allocation of computing resources.

[0137] With reference to Figure 6 , introduce the change of the average time delay of the Internet of Things device with the increase of the number of human type devices. As can be seen from the figure, with the increase of the number of human type devices, the average time delay of the Internet of Things device continuously increases. This is because when the number of human type devices in the system increases, not only the same channel interference increases, the received end SINR decreases, the transmission time delay of the Internet of Things increases, but also the computing resources allocated to the task of the Internet of Things device by the base station decrease because of the occupation of the computing resources, the computing time delay of the task increases, thereby reducing the QoS of the Internet of Things device. For scheme 3, although the human type device will not occupy the resources with the Internet of Things device, the interference to the Internet of Things device will be greater, so the resource allocation of the human type device will also reduce the system performance.

[0138] Based on the above analysis, it can be proved by the simulation results that the communication-computing resource joint allocation method scheme based on reinforcement learning in the human-computer-thing hybrid access heterogeneous network proposed in the application is feasible, and the overall performance is good under the guarantee of the QoS requirements of different devices.

[0139] The above application embodiment serial numbers are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0140] The integrated units in the above embodiments, if implemented in the form of software function units and sold or used as independent products, can be stored in the above computer-readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions to make one or more computer devices (which can be personal computers, servers or network devices, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application.

[0141] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0142] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Of course, the above device embodiment is only illustrative, and for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.

[0143] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiment according to actual needs.

[0144] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0145] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method for resource joint allocation in a hybrid access heterogeneous network, characterized in that, The method comprises: determining constraint conditions, decision variables and optimization objectives of devices in a heterogeneous network of human-machine-object hybrid access; defining a state set, an action set and a reward function of an agent in the devices based on the determined constraint conditions, decision variables and optimization objectives; traversing the state set of the agent, updating a return of a current state of the agent according to a reward value of the current state and a return estimation of a next state until a maximum reward value is found, based on the defined state set, action set and reward function; obtaining a resource allocation scheme that optimizes the optimization objectives based on the maximum reward value; wherein the reward function is obtained by: wherein C1, C2 and C3 are positive real numbers to balance the reward and punishment system, Zt is the calculated total cost, is a constraint based on the minimum transmission rate in the constraints, is a constraint for the maximum transmission latency, is a constraint for the computational power.

2. The method of claim 1, wherein, determining constraint conditions of devices in a heterogeneous network of human-machine-object hybrid access, comprising at least one of: modeling a constraint condition of a quality of service requirement of a human type device in the heterogeneous network as a minimum transmission rate constraint; modeling a constraint condition of a quality of service requirement of an Internet of Things device in the heterogeneous network as a maximum transmission power constraint; modeling a constraint condition of a quality of service requirement of a small base station (SBS) in the heterogeneous network as a user association number constraint; modeling a constraint condition of a quality of service requirement of an unmanned aerial vehicle (UAV) in the heterogeneous network as a user association number constraint, a maximum allowed power consumption constraint, and a computing capability constraint; wherein the devices in the heterogeneous network include the human type device, the Internet of Things device, the small base station, a macro base station and the unmanned aerial vehicle; and the agent includes the small base station and the unmanned aerial vehicle.

3. The method of claim 1, wherein, defining a state set, an action set and a reward function of an agent in the devices based on the determined constraint conditions, decision variables and optimization objectives, comprising: constructing the state set based on input data size R and CPU cycle number D of a task required to be completed by a user in the decision variables; constructing the action set based on a channel allocation matrix Γ, a power allocation p, and a CPU cycle number allocation in the decision variables; a constraint of minimum transmission rate among the constraints a constraint of maximum transmission latency a constraint of maximum power consumption and a constraint of computing capability and the optimization target to construct the reward function.

4. The method of claim 3, wherein, the optimization objectives are determined by: calculating a signal-to-interference-and-noise ratio of a user using a subchannel in the agent, and calculating a data transmission rate of the user transmitting data to the agent on the subchannel based on the signal-to-interference-and-noise ratio; calculating a transmission delay and a transmission energy consumption of the user based on the data transmission rate; calculating a computing delay required by the task and a computing energy consumption required by the task; calculating an overhead of the user based on the transmission delay, the transmission energy consumption, the computing delay required by the task and the computing energy consumption required by the task; minimizing the overhead of each user to determine the optimization objectives; wherein the user includes a human type device and / or an Internet of Things device.

5. The method of claim 1, wherein, traversing the state set of the agent, updating a return of a current state of the agent according to a reward value of the current state and a return estimation of a next state until a maximum reward value is found, comprises: repeatedly performing the following steps until the maximum reward value is found: After the agent selects an action based on a current state, a reward value of the current state of the agent is calculated according to feedback of a current network environment; An income estimate of a next state of the agent is obtained, and an income of the current state is updated according to the income estimate and the reward value of the current state; The current state enters a next state, and the next state is taken as a current state.

6. The method of claim 5, wherein, After the agent selects an action based on a current state, a reward value of the current state of the agent is calculated according to feedback of a current network environment, including: After the agent randomly selects an action based on a current state or selects an action with a maximum income according to a learning strategy, the agent shares the selected action, the current state and a corresponding income to the heterogeneous network, wherein the corresponding income is determined based on the current state and the selected action; A reward value of the current state of the agent is calculated according to feedback of resource usage of all devices in the heterogeneous network, so as to perform distributed cooperation with other devices in the heterogeneous network.

7. The method of claim 5, wherein, An income estimate of a next state of the agent is obtained, and an income of the current state is updated according to the income estimate, including: A next state of the agent is obtained; An action with a maximum reward in the next state is selected, so as to obtain an income estimate of the next state; An income of the current state is updated according to the income estimate of the next state.

8. The method according to any one of claims 1 to 7, characterized in that, Before determining a constraint condition, a decision variable and an optimization target of optimizing a device in the human-machine-object hybrid access heterogeneous network, the method further includes: Channel allocation, transmit power allocation and CPU cycle number allocation are initialized, and prior information for determining the optimization target and the constraint condition is calculated according to initial values obtained by the initialization; A reward table of each agent is created, the reward table indicating an income corresponding to a state and an action of a corresponding agent, wherein the reward table can be continuously updated according to a reward value of a current state of the agent and a maximum income in a next state.

9. A resource co-allocation device in a heterogeneous network with hybrid human-machine-object access, characterized in that, including: A determination module configured to determine a constraint condition, a decision variable and an optimization target of optimizing a device in the human-machine-object hybrid access heterogeneous network; A definition module configured to define a state set, an action set and a reward function of an agent in the device based on the determined constraint condition, the decision variable and the optimization target; An update module configured to traverse the state set of the agent based on the defined state set, the action set and the reward function, and update an income of a current state of the agent according to a reward value of the current state of the agent and an income estimate of a next state until a maximum reward value is found; An allocation module configured to obtain a resource allocation scheme that optimizes the optimization target based on the maximum reward value; The reward function is obtained by: wherein C1, C2 and C3 are positive real numbers to balance the reward and punishment system, Zt is the calculated total cost, is a constraint based on the minimum transmission rate in the constraints, is a constraint for the maximum transmission latency, is a constraint for the computational power.

10. A heterogeneous network of human-machine mixed access, characterized in that, including the following devices: a macro base station, a small base station, a drone, a human type device, an Internet of Things device, and a mobile edge computing server, wherein the mobile edge computing server includes the apparatus of claim 9, and the agent includes the drone and the small base station.