Network Resource Allocation Method Integrating Direct Transmission Communication of Fusion Terminals and Multi-Access Edge Computing
By adopting a network resource allocation method based on deep reinforcement learning in intelligent inspection of substations, the problem of optimized allocation of communication resources under D2D spectrum multiplexing and interference conditions under 5G technology is solved, and efficient resource allocation and throughput improvement is achieved.
Patent Information
- Application Number
- CN202211334660.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Under 5G technology, in intelligent inspection of substations, the optimization allocation of communication resources for direct terminal communication and multi-access edge computing is especially the challenges under D2D spectrum multiplexing and interference conditions.
The network resource allocation method based on deep reinforcement learning is adopted to integrate terminal direct communication and multi-level edge offloading. By iteratively selecting the devices for direct communication of terminals, the pre-trained DDQN framework is used to optimize the resource allocation strategy, including selecting the offload location, spectrum resources and power resources.
It maximizes resource allocation benefits under link interference and power constraints, improves network throughput, reduces computing delay, and improves the utilization rate of communication resources.
Smart Images

Figure CN115866787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of power grid system data transmission and equipment inspection, and particularly relates to a distributed network resource allocation method based on deep reinforcement learning for integrating terminal direct communication and multi-access edge computing. Background Art
[0002] The high-reliability, large-connection and low-latency characteristics of the 5th Generation Mobile Communication Technology (5G) will empower the rapid development of the power industry. With the development of smart grids, it is of great practical significance to use 5G and machine learning algorithms to achieve intelligent and efficient inspection of substations. Among them, key technologies such as Device to Device (D2D) and Mobile edge computing (MEC) can effectively improve the ability of 5G to serve smart grids, but it is necessary to solve the problem of optimal allocation of communication resources under D2D spectrum reuse and interference conditions.
[0003] MEC can provide cloud computing capabilities in the radio access network close to terminal devices. Applications and services run at the edge of the mobile network, reducing service latency and congestion in other parts of the mobile core network. Wenhe Li et al. proposed an intelligent control method for live working robots based on cloud computing and edge computing. By setting typical scenarios for the operation of live working robots in substations, it was verified by examples that the proposed intelligent control method can meet the computing power requirements of substation tasks. D. Han et al. used unmanned aerial vehicles as edge nodes to assist in task offloading and relaying of Internet of Things devices, and obtained the maximum system security capacity by jointly optimizing the position of the unmanned aerial vehicle, the task offloading rate and the offloading user allocation. They proposed to train the intelligent inspection task allocation mechanism based on deep reinforcement learning to reduce the latency and energy consumption of task offloading. However, the above MEC policy offloading only optimizes the offloading latency and energy consumption in terms of offloading location and computing power, without considering the allocation and optimization of communication resources. In view of the complexity of the substation transmission environment, the diversity of data and its offloading methods, it is necessary to study efficient wireless resource allocation and scheduling mechanisms to meet the requirements of non-interfering, stable and reliable transmission during data offloading.
[0004] The new cognitive-based D2D network formed by the combination of D2D communication and wireless network obtains proximity gain and channel reuse gain through spectrum resource reuse, thereby improving the data transmission efficiency of the 5G communication network and meeting the needs of concurrent device access. Different from the MEC system, the distributed computing D2D network has more complex topology management requirements and requires an efficient resource scheduling strategy. Aiming at the problem of collaborative D2D communication resource optimization in the uplink cellular network, under the condition of energy consumption constraint, the spectrum and power resources are optimally allocated with the goal of maximizing the network average throughput. Emna Fakhfakh et al. proposed a D2D mode selection scheme based on a new standard, which maximizes the system throughput and cellular traffic offloading efficiency by introducing noise parameters related to resource allocation. In addition, the researchers used game theory to analyze the resource allocation problem of maximizing the energy efficiency of cognitive D2D networks, and achieved the balance of energy efficiency and spectral efficiency under the constraint of user communication interference threshold. There is also research on the mode and resource allocation of D2D users accessing the cellular network based on evolutionary theory, which maximizes the total D2D user data rate. The above resource optimization methods can obtain the optimal solution when the data volume is small. When the number of system resources is large, the algorithm solving complexity increases, and deep reinforcement learning shows good performance in solving resource optimization problems. Under the condition of channel interference, using D2D technology to achieve spectrum reuse, complete the reliable transmission of inspection equipment data and the efficient utilization of MEC network resources are key problems to be solved urgently and also the focus of this embodiment. Summary of the Invention
[0005] The purpose of the present invention is to provide a network resource allocation method based on deep reinforcement learning for integrating terminal direct communication and multi-level edge offloading to solve at least one of the technical problems existing in the above background technology.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] The present invention provides a network resource allocation method for integrating terminal direct communication and multi-access edge computing, including:
[0008] Iteratively select devices for terminal direct communication;
[0009] According to the selected devices, use the resource allocation strategy of the pre-trained deep reinforcement learning framework based on DDQN to select the offloading location, spectrum resources, and power resources;
[0010] Among them, the training of the resource allocation model includes:
[0011] Randomly initialize the strategy, start the environment simulator, and generate devices for direct terminal communication, inspection targets, and base stations integrating multi-access edge computing; initialize the Q-network and Target Q-network to generate initial weights;
[0012] Iteratively select devices for direct terminal communication, select the offloading location according to the resource optimization strategy, and determine the power and spectrum to be transmitted;
[0013] The environment simulator selects an action from the Q-network according to ε-greedy, enters a new state, calculates the network throughput and energy consumption according to the current spectrum occupancy, generates a reward according to the set reward function and calculates the new Q value, and saves the calculated network throughput, energy consumption, and updated Q value in Experience Replay;
[0014] Sample data from Experience Replay for network training; update the weights of the Target Q-network every once in a while until the LOSS function converges to obtain the finally trained resource allocation model.
[0015] Preferably, the resource optimization strategy includes comprehensively considering the requirements of throughput, energy consumption, and computing delay metrics, and establishing a resource optimization allocation model based on the maximization of the comprehensive benefit function as:
[0016]
[0017] s.t.C1:
[0018] C2:
[0019] Among them, M inspection devices in the substation cooperate to complete the inspection work, and Μ∈{1,2,…M} represents the set composed of M inspection devices, corresponding to M transmission links from the devices to the base station; C m represents the capacity of the transmission link from the m-th device to the base station; represents the capacity of the i-th receiver of the m-th device; τ m represents the computing delay; α k,m represents the channel reuse coefficient. When the transmission link from the k-th device to the device reuses the spectrum of the transmission link from the m-th device to the base station, then α k,m =1, otherwise α k,m =0; χ [j,m′][m,i] =1 indicates that the m'-th receiver of the j-th inspection device and the i-th receiver of the m-th inspection device use the same spectrum resource, otherwise χ [j,m′][m,i] =0; P m represents the transmission power of the m-th transmission link from the device to the base station; Denotes the transmission power consumption of the k-th device-to-device transmission link; Denotes the power consumption when the m-th device task is offloaded; P′ denotes the circuit power consumption of the device-to-device transmission link; Denotes the maximum transmit power that the m-th device can provide; Denotes the transmission power of the transmission link from the m-th device to the m'-th device; Is the D2D link transmission power from the j-th inspection device to the m'-th inspection device; Denotes the peak interference power that the channel can tolerate; Denotes the interference power gain of the D2D link from the j-th inspection device to the m'-th inspection device..
[0020] Preferably, under the condition of known transmit power p m and noise power σ 2 , the signal-to-interference-plus-noise ratio γ m of the transmission link from the m-th device to the base station is related to the spectrum resource allocation of the device-to-device transmission link:
[0021]
[0022] where κ = {1, 2,..., K = M·(M - 1) / 2} represents all possible link sets; P m and are the transmission powers of the m-th device-to-base station (D2B) link and the k-th device-to-device (D2D) transmission link respectively, h m is the power gain of the channel corresponding to the transmission link from the m-th device to the base station, h k denotes the interference power gain of the k-th D2D transmission link; when the k-th D2D link reuses the spectrum of the m-th D2B link, then α k,m = 1, otherwise α k,m = 0.
[0023] According to the signal-to-interference-plus-noise ratio expression, the transmission link capacity C m of the m-th device to the base station is:
[0024] C m = w·log2(1 + γ m )
[0025] where w is the sub-channel bandwidth.
[0026] Preferably, for the i-th receiver of the m-th inspection device, its signal-to-interference-plus-noise ratio is expressed as:
[0027]
[0028] where, is the transmission power of the i-th receiver of the m-th inspection device, g m,i is the power gain of the i-th receiver of the m-th inspection device; is the noise power in the received signal, ρ is the interference power of the device-to-base station transmission link multiplexing the same resource block, ρ D is the total interference power of all device-to-device transmission links sharing the same resource block;
[0029]
[0030] where represents the spectrum reuse factor, means that the n-th device-to-base station transmission link and the i-th receiver of the m-th inspection device share the same spectrum, otherwise is the interference power gain of the n-th device-to-base station transmission link; P n is the transmission power of the device-to-base station transmission link;
[0031]
[0032] where is the transmission power of the transmission link from the j-th inspection device to the m'-th inspection device; is the D2D link interference power gain from the j-th inspection device to the m'-th inspection device. χ [j,m'],[m,i] = 1 means that the m'-th receiver of the j-th inspection device and the i-th receiver of the m-th inspection device use the same spectrum resource, otherwise χ [j,m'],[m,i] = 0.
[0033] Finally, the capacity of the i-th receiver of the m-th inspection device is:
[0034]
[0035] Preferably, under the condition of satisfying the rate and delay constraints of the device-to-base station transmission link, considering the device-to-device transmission link and the device-to-base station transmission link comprehensively, the network throughput is:
[0036]
[0037] Then the total power consumption E of the system is:
[0038]
[0039] where, τ m is the calculation delay, is the calculation power consumption, then when the task is calculated locally, the processing delay is:
[0040]
[0041] where u m is the local computation data volume; ξ m is the computational complexity of the D2D device, that is, the number of CPU cycles required to process 1 bit of data; f m represents the CPU frequency of the device.
[0042] According to the local computation volume and the CPU parameters of the inspection device, the power consumption during device task offloading can be calculated as:
[0043] where κ m represents the switched-capacitor factor, η m is the coefficient factor.
[0044] Preferably, since the power consumption will affect the network throughput, an equilibrium adjustment needs to be made between the power consumption and the throughput in the reward function, and the calculation delay condition is used as a penalty to reduce the impact on the reward. The reward function is:
[0045]
[0046] Let represent the energy efficiency of the system, then the system reward function can be simplified as:
[0047]
[0048] After normalizing the profit, it is:
[0049]
[0050] Among them, the equilibrium factor λ is used to balance the power consumption and the throughput to obtain a weighted benefit function.
[0051] Preferably, the state space S of the DDQN is defined as S = V t ×C t ×G t ×H t-1 ; where V t = {v1, v2} represents the offloading location, v1 represents local offloading, and v2 represents offloading to the integrated MEC server; C t = {c1, c2,..., c g} represents the information set of g sub-channels, c g = 0 indicates that the current sub-channel is not occupied, and c g = x indicates that the sub-channel is repeatedly occupied x times at the current moment; G t = {g1, g2,..., g v}\ represents a set of \(v\) link power gains; the interference signal strength \(H\) received in the previous time slot t-1 , representing the local observation result in each sub-channel.
[0052] Preferably, the action selection of DDQN includes offloading location, spectrum, and power information; define the action \(A = \{a_1, A_2, a_3\}\), where \(a_1\in\{0, 1\}\), \(a_1 = 0\) represents selecting local offloading, and \(a_1 = 1\) represents selecting offloading to the integrated MEC server; \(A_2\) represents the channel selection vector, which is the set of allocated sub-channels; \(a_1\in\{p_1,...p i ,...p l},\(a_1 = p i represents the \(i\)-th allocated power as \(p i , and \(l\) is the number of sub-channels; after the agent selects an action, it interacts with the environment to generate a reward and update the state.
[0053] Preferably, the Loss function adopts the mean squared error function:
[0054]
[0055] Advantages of the present invention: Through MEC, multi-level offloading is achieved. At the same time, by using D2D communication technology, the reuse and distributed scheduling of communication resources are realized. A system benefit function considering indicators such as joint network throughput, power consumption, and computing delay is established. The problem of maximizing benefits under link interference and power constraints is proposed, and optimal offloading selection and resource allocation are achieved; a deep reinforcement learning framework based on DDQN is used to realize the joint optimization of 5G resource block allocation and computing offloading, maximizing network throughput and reducing computing delay.
[0056] Advantages of additional aspects of the present invention will be more clearly presented in the following description part, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 Schematic diagram of the system model of the D2D-assisted MEC network described in the embodiments of the present invention.
[0059] Figure 2 Schematic diagram of the influence of the number of inspection devices on the benefit function described in the embodiments of the present invention.
[0060] Figure 3Schematic diagram of the impact of the number of inspection devices described in the embodiments of the present invention on the system throughput.
[0061] Figure 4 Schematic diagram of the impact of the number of subcarriers described in the embodiments of the present invention on the system throughput.
[0062] Figure 5 Schematic diagram of the impact of the number of inspection devices and subcarriers described in the embodiments of the present invention on the MEC offloading probability.
[0063] Figure 6 Schematic diagram of the change of the reward function with the episode described in the embodiments of the present invention. Detailed implementation manners
[0064] The following details the implementation manners of the present invention. The examples of the implementation manners are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The implementation manners described through the drawings are exemplary and are only used to explain the present invention and cannot be construed as a limitation to the present invention.
[0065] Those skilled in the art of this technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.
[0066] It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0067] Those skilled in the art of this technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used here may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements and / or their groups.
[0068] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0069] For ease of understanding the present invention, the following further explains the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0070] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0071] Embodiment
[0072] In this embodiment, aiming at the problems existing in the application of 5G technology in substation inspection equipment, considering the respective characteristics of MEC and D2D technologies, a D2D-assisted MEC network offloading algorithm is proposed. Multilevel offloading is achieved through MEC, and at the same time, D2D communication technology is used to realize the reuse and distributed scheduling of communication resources. In order to achieve the optimal offloading selection and resource allocation, a system benefit function combining indicators such as joint network throughput, power consumption, and computing delay is established, and a benefit maximization problem under link interference and power constraints is proposed. Finally, a deep reinforcement learning framework based on DDQN is adopted to realize the joint optimization of 5G resource block allocation and computing offloading, maximize the network throughput, and reduce the computing delay as much as possible.
[0073] The resource allocation model of the D2D-based MEC system is as Figure 1 shown. Considering that there is a substation within the range of a base station with an integrated MEC server, M inspection devices in the substation cooperate to complete the inspection work, corresponding to M D2B links. The device set and the device-to-base station (D2B) link set are defined as The inspection devices can obtain the location information of other inspection devices through D2D communication. The inspection devices collect sensing data and can choose to process it on the local device or offload it to the base station (D2B link). The interference at the base station is more controllable and less uplink resource is used. Therefore, it is assumed that each message has a group of receivers for processing and can communicate with other devices respectively (the total number of receivers does not exceed M). The uplink spectrum of the D2B link is reused with the D2D link.
[0074] The allocation of radio resources is divided into two dimensions: time domain and frequency domain. In the time domain, the resource allocation is mainly for each Transmission Time Interval (TTI). In the frequency domain, the total bandwidth is divided into several equal-bandwidth sub-channels, and sub-channel allocation is required. A single TTI and a single sub-channel form a system resource block (RB), which is the smallest radio resource unit required for device data transmission. Therefore, the interference to the D2B link comes from background noise and the D2D link signals sharing the same sub-band.
[0075] Under the condition of known transmit power and noise power σ 2 the signal-to-interference-plus-noise ratio γ m of the m-th D2B link is closely related to the spectrum resource allocation of the D2D link and can be expressed as:
[0076]
[0077] where κ = {1, 2, …, M·(M - 1) / 2} represents all possible link sets; P m and represent the transmission powers of the m-th D2B link and the k-th D2D link respectively; h m is the power gain corresponding to the m-th D2B channel, and h k represents the interference power gain of the k-th D2D link; α k,m represents the channel reuse coefficient. When the k-th D2D link reuses the spectrum of the m-th D2B link, then α k,m = 1, otherwise α k,m = 0.
[0078] According to the signal-to-interference-plus-noise ratio expression, the capacity C m of the m-th D2B link is:
[0079] C m = w·log2(1 + γ m ) (2)
[0080] where w is the sub-channel bandwidth.
[0081] Similarly, for the i-th receiver of the m-th inspection device, its signal-to-interference-plus-noise ratio is expressed as:
[0082]
[0083] In Equation (3), is the transmission power of the i-th receiver of the m-th inspection device, and g m,i is the power gain of the i-th receiver of the m-th inspection device; is the received noise power, ρ is the interference power of the D2B link multiplexing the same RB, ρ D is the total interference power of all D2D links sharing the same RB.
[0084] ρ in Equation (3) is as shown in (4):
[0085]
[0086] where represents the spectrum reuse factor, means that the i-th receiver of the n-th D2B link and the m-th inspection device share the same spectrum, otherwise is the interference power gain of the n-th D2B link; Pn is the transmission power of the D2B link.
[0087] ρ in Equation (3) D is as shown in (5):
[0088]
[0089] where is the transmission power of the D2D link from the j-th inspection device to the m'-th inspection device; is the interference power gain of the D2D link from the j-th inspection device to the m'-th inspection device; χ [j,m′][m,i] also represents the spectrum reuse factor, χ [j,m′][m,i] = 1 means that the m'-th receiver of the j-th inspection device and the i-th receiver of the m-th inspection device use the same spectrum resource, otherwise χ [j,m′][m,i] = 0.
[0090] Finally, the capacity of the i-th receiver of the m-th inspection device is:
[0091]
[0092] According to the network mathematical model, D2D technology improves the resource utilization rate through spectrum reuse, but it also inevitably brings link interference. Therefore, under the conditions of meeting the rate and delay constraints of the D2B link, the quality of the D2D link should be improved as much as possible. Considering both D2D and D2B links comprehensively, the throughput of the D2D-assisted MEC network can be expressed as:
[0093]
[0094] In this embodiment, the system energy consumption and calculation model are as follows:
[0095] Most of the inspection devices are power - limited. Therefore, it is necessary to consider the power consumption of MEC task calculation and offloading. Since the integrated MEC server is deployed in the network management center and is active, the power consumption limit of the MEC server can be ignored. This embodiment focuses on calculating the power consumption of the inspection device.
[0096] Define the circuit power consumption of the D2D device as P′, and the transmission power consumption of the D2D device as The calculated power consumption is Then the total power consumption E of the MEC system is directly related to the various power consumptions of the D2D inspection device, expressed as:
[0097]
[0098] Calculating the delay is another key indicator of task processing, which is closely related to the computing resources of the inspection device or server. The more computing resources there are, the smaller the processing delay. Task offloading is mainly divided into two levels: local offloading of the D2D device and offloading to the integrated MEC server. Compared with local offloading, the integrated server has stronger power supply and computing capabilities. When optimizing the algorithm in this embodiment, the computing delay and power consumption generated by local offloading are mainly considered.
[0099] Define τ m as the computing delay, as the computing power consumption. Then when the task is calculated locally, the processing delay is:
[0100]
[0101] where u m is the local computing data volume; ξ m is the computing complexity of the D2D device, that is, the number of CPU cycles required to process 1 bit of data; f m represents the CPU frequency of the D2D inspection device.
[0102] According to the local computing volume and the CPU parameters of the inspection device, the power consumption when the device task is offloaded can be calculated as
[0103] In the formula, κ m represents the switching - capacitor factor, and η m is the coefficient factor.
[0104] In this embodiment, the resource - optimized allocation model is specifically as follows:
[0105] Considering the limitation of the battery capacity of the D2D inspection device, its transmission power cannot be infinitely large. Therefore, the transmission power satisfies the following constraints:
[0106]
[0107] Among them is the maximum transmit power that the m-th D2D inspection device can provide.
[0108] In addition, since the interference power will affect the D2D device and cause transmission interruption, affecting the communication quality, the constraint of the interference power also needs to be satisfied during the resource allocation process:
[0109]
[0110] Among them represents the peak interference power that the channel can tolerate, and the definitions of other variables are the same as those in Equation (5).
[0111] From the perspective of intelligent inspection data transmission and task execution, the resource optimization algorithm in this embodiment needs to improve the throughput of the D2D-assisted MEC network under basic constraints such as power and interference, and ensure the minimum calculation delay of sensing data through a reasonable task offloading and resource allocation algorithm. Therefore, considering the index requirements such as throughput, energy consumption, and delay, a resource optimization allocation model based on the maximization of the comprehensive benefit function is established as follows:
[0112]
[0113] s.t.C1:
[0114] C2:
[0115] Since the value ranges of indexes such as capacity, power consumption, and delay are different and there are differences in measurement, normalization processing is required during the optimization solution process. After normalizing the original data x, the result The specific method is:
[0116]
[0117] where x max is the data maximum value, and x mid is half of the data maximum value.
[0118] The above optimization problem is a mixed integer non-linear programming problem, and multiple optimization variables are coupled with each other. It is difficult to use traditional convex optimization schemes to solve even under all statistical distributions. In addition, the relationship between the observed values and the optimal resource allocation solution is often implicit and difficult to establish using analytical methods. Therefore, an offloading decision optimization and resource allocation algorithm based on the Deep Reinforcement Learning (DRL) framework is proposed, which uses the implicit relationship between the observed values and the optimal resource allocation to realize the online interaction between the state and the system.
[0119] The DDQN algorithm is as follows:
[0120] Deep reinforcement learning combines the decision-making ability of reinforcement learning and the powerful data analysis ability of deep neural networks
[16] , which can solve the dimensionality explosion problem brought about by the Q-Learning algorithm when the state space is large. The updated mathematical expression is:
[0121] Among them, R t+1 is the reward, s t+1 is the next state, a is the selected action, γ is the decay factor for R, are the parameters of the Q network.
[0122] The DQN algorithm selects the maximum value during the update , and this max operation causes the value function to be overestimated. Therefore, a double network can be used to select actions and evaluate the current state value
[10] , that is, the DDQN algorithm. The algorithm update process is as follows:
[0123]
[0124] Among them, θ t and are the parameters of the Q network and the Target Q network respectively. DDQN selects actions in a greedy manner from the Q network and evaluates the Q value in the Target Q network.
[0125] In this embodiment, the optimization strategy based on DDQN is specifically:
[0126] The D2B link has strict requirements for latency and reliability. In DDQN, these constraints are directly represented as the reward function. The goal of the resource management scheme proposed in this embodiment is to ensure that the latency constraint of the D2B link is satisfied while minimizing the interference of the D2D link to the D2B link.
[0127] Since power consumption will affect the network throughput, a balanced adjustment of power consumption and throughput needs to be made in the reward function. The calculation delay condition is used as a penalty to reduce the impact on the reward. The reward function can be expressed as:
[0128]
[0129] Let represent the energy efficiency of the system, then the system reward function can be simplified to the following form:
[0130]
[0131] After normalization, it is:
[0132]
[0133] It can be seen that the reward function is similar but not exactly the same as the benefit function. We use a balance factor λ to balance power consumption and throughput to obtain a weighted benefit function. Similarly, the data needs to be normalized, and the processing rule is the same as formula (14).
[0134] The observations related to resource allocation are channel and interference information. Define the state space S of the DDQN as S = V t ×C t ×G t ×H t-1 , where:
[0135] 1) V t = {v1, v2} represents the offloading location, where v1 represents local offloading and v2 represents offloading to the integrated MEC server; 2) C t = {c1, c2, …… c g} represents the information set of g sub-channels. c g = 0 indicates that the current sub-channel is not occupied, and c g = x indicates that the sub-channel is repeatedly occupied x times at the current moment; 3) G t = {g1, g2, ……, g v} represents the set of v link power gains; 4) The interference signal strength H received in the previous time slot t-1 , representing the local observation result in each sub-channel, also includes information shared by neighbors, such as the channel index selected by neighbors in the previous time slot.
[0136] The action selection of the DDQN includes offloading location, spectrum, and power information. Define the action A = {a1, A2, a3}, where a1 ∈ {0, 1}, a1 = 0 indicates selecting local offloading, and a1 = 1 indicates selecting offloading to the integrated MEC server. A2 represents the channel selection vector, which is the set of sub-channels allocated. a1 ∈ {p1, … p i , … p l}, a1 = p i indicates that the i-th allocated power is p i , and l is the number of sub-channels. After the agent selects an action, it interacts with the environment to generate a reward and update the state. The DDQN algorithm is practiced according to the set reward function, state space, and action. The specific environment settings and parameters are introduced below.
[0137] The resource optimization allocation method provided in this embodiment is as follows:
[0138] It is divided into two stages, the training stage and the testing stage. Training and testing data are generated through the interaction between the environment simulator and the agent, which are used to optimize the Q-network and the Target Q-network. At the beginning stage, each training sample includes s t 、s t+1 、a t and r t , generating an experience pool Experience Replay. The action selection adopts ε-greedy, randomly selecting actions with a probability of 10%, and selecting the action with the largest Q value with a probability of 90%.
[0139] The environment simulator includes D2D devices and an integrated MEC server and their channels, where the positions of the D2D devices are randomly generated. By selecting the spectrum and power of the D2D link, the simulator can provide s t+1 and R t to the agent. In each iteration of the training stage, 50 data are sampled from the Experience Replay, which can suppress the temporal correlation of the generated data. Then, actions are selected through the Q-network, evaluated using the Target Q-network, and the weights are updated. The Loss function uses the mean squared error function:
[0140]
[0141] The initialization of the spectrum and power selection strategy for each D2D link is random, and the utility function is iteratively calculated using the Q-network. In the testing stage, actions in the D2D link are selected according to the trained network and evaluated accordingly.
[0142] The main steps of resource optimization allocation include the following parts:
[0143] 1) System modeling: There are a total of M D2D devices, inspection targets, and a base station with an integrated MEC server.
[0144] 2) Parameter definition: Define parameters such as channels, fading, and noise (specific values are shown in Table 1), and define system resource parameters or variables (optimization objectives).
[0145] 3) Index calculation: According to the model and parameters, calculate the signal-to-interference-plus-noise ratio of the m-th D2B; the D2B link capacity; calculate the signal-to-interference-plus-noise ratio and capacity of the i-th receiver of the m-th D2D device; network throughput and power consumption and other indicators.
[0146] 4) Algorithm description: Then, the dual-network DQN is used for channel and power allocation, and the specific process is summarized as follows:
[0147]
[0148]
[0149] In this embodiment, simulations and analyses of the above resource allocation method are provided:
[0150] The simulation configuration is as follows: The simulation is based on the TensorFlow 1.0 framework in Python. Considering a substation environment of 500m × 500m, M D2D inspection devices are randomly generated, and a base station with an integrated MEC server is located 2 km away from the center of the substation. The channel adopts the Rice model, and the simulation parameters are shown in Table 1.
[0151] The network model uses a BP neural network, including an input layer, three hidden layers, and an output layer. The number of neurons in the three hidden layers is 64, 128, and 128 respectively, and the activation function is the ReLU function.
[0152] Table 1 D2D-assisted MEC network parameters
[0153]
[0154] The DDQN algorithm proposed in this embodiment is compared with the MEC-U algorithm and the random algorithm Random. Among them, the MEC-U algorithm means that tasks are all offloaded to the integrated MEC server, and the rest is the same as the algorithm in this embodiment; the Random algorithm means that communication resources and offloading locations are randomly selected. The results are as Figure 2 shown, indicating that the algorithm adopted in this embodiment has good performance.
[0155] When the number of inspection devices is small, the amount of resources provided by the system can meet the communication requirements, and at this time, the system benefit functions brought by the two algorithms are close. As the number of inspection devices increases, the increasing communication requirements lead to a shortage of resources, and the reuse of spectrum resources results in a smaller utility function. However, compared with the MEC-U algorithm and the Random algorithm, the DDQN framework proposed in this embodiment optimizes the resource allocation strategy to reduce channel interference by deeply exploring the implicit relationship between interference and allocation strategies, and at the same time reduces the calculation delay through offloading decisions, keeping the benefit function at a relatively high level. The data simulation results show that the DDQN algorithm proposed in this embodiment has a certain degree of reliability and effectiveness.
[0156] Figure 3It shows the relationship between the system throughput and the number of inspection devices, and compares it with MEC-U and Random selection. The results show that the system throughput first increases and then decreases with the increase in the number of inspection devices. This is because when the number of inspection devices is small, the system resources are not fully utilized, and the amount of data to be transmitted in the network is limited. The system throughput decreases after the number of inspection devices reaches a certain amount because the limited network resources lead to an increase in channel interference. The offloading and resource allocation strategies obtained by the DDQN algorithm are significantly better than those of MEC-U and Random allocation. By reasonably selecting the offloading location and efficient scheduling strategies, it can better combat channel interference and has excellent performance.
[0157] Subsequently, this embodiment studied the variation of the system throughput with the number of subcarriers and compared it with the Random and AFSA algorithms. The results are as Figure 4 shown. The more subcarriers there are, the greater the system throughput. This is because when the resources are sufficient, the interference between channels is less, and more data chooses to be offloaded and calculated on the integrated MEC server, resulting in an increase in the system throughput. The resource allocation strategy of the DDQN algorithm proposed in this embodiment is superior to those of the AFSA algorithm and the Random algorithm, with less channel interference and greater throughput in the system.
[0158] This embodiment studied the influence of the number of inspection devices and the number of subcarriers d on the offloading strategy, and the results are as Figure 5 shown. The probability of choosing to offload on the integrated MEC server increases with the increase in the number of subcarriers and decreases with the increase in the number of inspection devices. This is because when the resources are relatively loose compared to the communication requirements, choosing to offload on the integrated MEC server can reduce the calculation delay. When the resources are relatively scarce, the interference between systems increases, and tasks more often choose local offloading to reduce interference and ensure the reliability and effectiveness of the system.
[0159] Figure 6 The training results of the rewards of different algorithms were studied. It can be seen that the DDQN algorithm can select actions in the training set, thereby improving the reward, can explore the implicit relationship between resource allocation and reward, has a higher reward than random allocation, and demonstrates good performance.
[0160] During the intelligent inspection process of electric power, sinking resources through MEC can relieve the pressure on the core network and provide fast computing services. In response to the need for interconnection between inspection devices in this embodiment, MEC and D2D technologies are combined to establish a D2D-assisted MEC network. In order to reduce the interference between different links, a 5G resource optimization problem with throughput, power consumption, and calculation delay as indicators is established, and it is solved through the DDQN framework and the effectiveness of the algorithm is verified by simulation. In subsequent work, we will perform game calculations on the data transmission of different inspection devices and optimize the driving trajectories.
[0161] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device to perform a series of operation steps on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a process Figure 1 a process or multiple processes and / or blocks Figure 1 steps of the functions specified in a block or multiple blocks.
[0163] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions disclosed in the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.
Claims
1. A network resource allocation method integrating terminal direct communication and multi-access edge computing, characterized in that, Including: Iteratively select the devices for direct terminal communication; According to the selected devices, use the resource allocation strategy of the pre-trained deep reinforcement learning framework based on DDQN to select the offloading location and spectrum resources; Among them, the training of the resource allocation strategy includes: Randomly initialize the strategy, start the environment simulator, generate the devices for direct terminal communication, the inspection targets, and the base stations integrating multi-access edge computing; initialize the Q-network and the Target Q-network, and generate the initial weights; Iteratively select the devices for direct terminal communication, select the offloading location according to the resource optimization strategy, and determine the power and spectrum to be transmitted; The environment simulator selects actions from the Q-network according to the ε-greedy method, enters a new state, calculates the network throughput and energy consumption according to the current spectrum occupancy, generates a reward according to the set reward function and calculates the new Q value, and saves the calculated network throughput, energy consumption, and the updated Q value in the Experience Replay; Sample data from Experience Replay for network training; update the weights of the Target Q-network every once in a while until the LOSS function converges to obtain the finally trained resource allocation model; the Loss function uses the mean squared error function: Among them, the resource optimization strategy includes comprehensively considering the requirements of throughput, energy consumption, and delay metrics, and establishing a resource optimization allocation model based on the maximization of the comprehensive benefit function as: Among them, M inspection devices in the substation cooperate to complete the inspection work, and there are M transmission links from the devices to the base station; C m represents the capacity of the transmission link from the m-th device to the base station; represents the capacity of the i-th receiver of the m-th device; τ m represents the calculation delay; α k,m represents the channel reuse factor. When the transmission link from the k-th device to the device reuses the spectrum of the transmission link from the m-th device to the base station, then α k,m = 1, otherwise α k,m = 0; χ [j,m′],[m.i] = 1 indicates that the m'-th receiver of the j-th inspection device and the i-th receiver of the m-th inspection device use the same spectrum resource, otherwise χ [j,m′],[m.i] = 0; P m represents the transmission power of the m-th transmission link from the device to the base station; represents the transmission power consumption of the k-th device-to-device transmission link; represents the power consumption when the m-th device unloads tasks; P' represents the circuit power consumption of the device-to-device transmission link; represents the maximum transmit power that the m-th device can provide; represents the transmission power of the transmission link from the m-th device to the m'-th device; is the D2D link transmission power from the j-th inspection device to the m'-th inspection device; represents the peak interference power that the channel can tolerate; represents the interference power gain of the D2D link from the j-th inspection device to the m'-th inspection device; Since the power consumption will affect the network throughput, an equilibrium adjustment needs to be made to the power consumption and throughput in the reward function, and the delay condition is calculated as a penalty to reduce the impact on the reward. The reward function is: Let represent the energy efficiency of the system, then the system reward function can be simplified as: After normalization, it is: Among them, the equilibrium factor λ is used to balance the power consumption and throughput to obtain the weighted benefit function; Define the state space S of DDQN as S = V t × C t × G t × H t-1 ; where Vt = {v1, v2} represents the offloading location, v1 represents local offloading, and v2 represents offloading to the integrated MEC server; C t = {c1, c2,... c g} represents the information set of g sub-channels, c g = 0 indicates that the current sub-channel is not occupied, and c g = x indicates that the sub-channel is repeatedly occupied x times at the current moment; G t = {g1, g2,..., g v} represents the set of v link power gains; H is the intensity of the interference signal received in the previous time slot t-1 , representing the local observation result in each sub-channel; The action selection of DDQN includes offloading location, spectrum, and power information; define action A = {a1, A2, a3}, where a1 ∈ {0, 1}, a1 = 0 indicates selecting local offloading, and a1 = 1 indicates selecting integrated MEC server offloading; A2 represents the channel selection vector, which is the set of allocated sub-channels; a1 ∈ {p1,…p i ,…p l}, a1 = p i indicates that the allocated power is pi, and l is the number of sub-channels; after the agent selects an action, it interacts with the environment to generate a reward and update the state.
2. The network resource allocation method integrating terminal direct communication and multi-access edge computing according to claim 1, characterized in that, Under the condition of known transmit power and noise power σ 2 the signal-to-interference-plus-noise ratio γ m of the m-th device to base station transmission link is related to the spectrum resource allocation of the device-to-device transmission link: where κ = {1, 2, …, K = M·(M - 1) / 2} represents all possible link sets; h m is the power gain of the transmission link channel from the m-th device to the base station, and h k represents the interference power gain of the k-th device-to-device transmission link; According to the signal-to-interference-plus-noise ratio expression, the transmission link capacity C of the m-th device to the base station m is as follows: C m = w·log2(1 + γ m ) Where w is the sub-channel bandwidth.
3. The network resource allocation method integrating terminal direct communication and multi-access edge computing according to claim 2, characterized in that, For the i-th receiver of the m-th inspection device, its signal-to-interference-plus-noise ratio is expressed as: Among them, is the transmission power of the i-th receiver of the m-th inspection device, g m,i is the power gain of the i-th receiver of the m-th inspection device; is the received noise power, ρ is the interference power of the device using the same resource block to the base station transmission link, ρ D is the total interference power of all device-to-device transmission links sharing the same resource block; Among them represents the spectrum reuse factor indicates that the transmission link of the nth device to the base station and the ith receiver of the mth inspection device share the same spectrum, otherwise is the interference power gain of the transmission link of the nth device to the base station; P n is the transmission power of the transmission link from the device to the base station Among them is the transmission power of the transmission link from the j-th inspection device to the m'-th inspection device; is the D2D link interference power gain from the j-th inspection device to the m'-th inspection device; Finally, the capacity of the i-th receiver of the m-th inspection device is:
4. The network resource allocation method integrating terminal direct transmission communication and multi-access edge computing according to claim 3, characterized in that, Under the conditions of satisfying the rate and delay constraints of the device-to-base station transmission link, comprehensively considering the device-to-device transmission link and the device-to-base station transmission link, the network throughput is: Then the total power consumption E of the system is: where τ m is for calculating the time delay, is for calculating the power consumption. Then when the task is calculated locally, the processing time delay is: where u m is the amount of locally computed data; ξ m is the computational complexity of the D2D device, i.e., the number of CPU cycles required to process 1 bit of data; f m represents the CPU frequency of the device; Based on the local computing volume and the CPU parameters of the inspection device, the power consumption during device task offloading can be calculated as follows: Among them, κ m represents the switched-capacitor factor, and η m is the coefficient factor.
Citation Information
Patent Citations
MEC task unloading and resource allocation method based on deep reinforcement learning
CN113612843A
A D2D-MEC offloading method based on deep reinforcement learning, and a computer program product
CN114938381A