A handover offloading method and device based on deep reinforcement learning in a D2D environment

By optimizing the offloading action using a branch-deep Q network in a D2D environment and combining it with a base station handover mechanism, the complexity of computational offloading in multi-user scenarios is solved, thereby improving offloading efficiency and accuracy and adapting to complex IoT environments.

CN115103338BActive Publication Date: 2026-03-31UNIV OF SCI & TECH BEIJING +3
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2026-03-31

Smart Images

  • Figure CN115103338B_ABST
    Figure CN115103338B_ABST
Patent Text Reader

Abstract

The application discloses a handover unloading method and device based on deep reinforcement learning in a D2D environment, which comprises the following steps: an edge computing unloading environment capable of simultaneously realizing D2D unloading among users and task handover among base stations is established, and relevant environment parameters are initialized; based on the unloading environment, the composition of unloading actions is determined, and the relationship between the unloading actions and the unloading environment interaction is established; according to the relationship between the unloading actions and the unloading environment interaction, the unloading cost is calculated, the reinforcement learning four-tuple is constructed, and the reward and penalty items are set; based on the constructed reinforcement learning four-tuple and the reward and penalty items, the deep reinforcement learning algorithm is adopted to realize the handover unloading in the D2D environment and minimize the weighted sum of all user time delays and energy consumptions. And aiming at the problems of multiple users, multiple base stations and difficult-to-discretize representation of unloading actions, the deep Q network is structurally improved, high-dimensional discrete actions are decomposed into lower-dimensional action combinations, the problem is greatly simplified, and the method is beneficial to implementation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of wireless network communication, edge computing offloading, and reinforcement learning, and particularly to a handover offloading method and apparatus based on deep reinforcement learning in a D2D environment. Background Technology

[0002] In recent years, with the increasing shift of cloud functions to the network edge and the explosive growth of internet data, cloud computing has become increasingly inadequate for solving data processing problems. To address this explosive growth in data volume, Mobile Edge Computing (MEC) has emerged. MEC refers to deploying computing and storage resources at the edge of mobile networks to provide an IT service environment and cloud computing capabilities, thereby offering users ultra-low latency and high bandwidth network service solutions. Today, MEC can also be interpreted as "multi-access-point edge computing." In MEC, compute offloading is a crucial technology. Compute offloading refers to the technology where terminal devices delegate some or all of their computing tasks to the cloud computing environment to address the shortcomings of mobile devices in terms of resource storage, computing performance, and energy efficiency.

[0003] End-to-end (D2D) communication technology refers to the technology that allows adjacent devices to communicate directly without going through a base station. It allows nearby terminals to interconnect under the control of a base station (such as Bluetooth, Wi-Fi, etc.) to achieve data communication, thereby forming a data communication network.

[0004] Given the current state of wireless communication characterized by the widespread adoption of smart devices and the surge in computationally intensive tasks, research on computation offloading in D2D environments is of great significance. Furthermore, with the increasing complexity of network structures and rising user demands, the offloading decisions requiring optimization are becoming increasingly complex. The edge computing offloading optimization problem often becomes an NP-hard problem, and traditional mathematical optimization methods struggle to solve such NP-hard problems.

[0005] Furthermore, most current research on computational offloading techniques focuses on optimizing offloading requirements between users and servers, with limited research on offloading in D2D environments. Current research on D2D edge environments concentrates on edge caching, with limited research on computational offloading. Moreover, existing research on computational offloading in D2D edge environments largely relies on traditional mathematical optimization methods, which are insufficient for handling multi-user scenarios and NP-hard problems.

[0006] To address the NP-hard problems that traditional methods struggle with in D2D environments, deep reinforcement learning is needed to solve these optimization problems. Generally, unloading tasks are difficult to decompose, often requiring discrete action deep reinforcement learning methods for optimization. However, in D2D environments with task handover, each unloading action has numerous sub-actions. Traditional discrete deep reinforcement learning methods, i.e., deep Q-networks, often perform well in handling "0-1" unloading problems. But in complex environments with numerous sub-actions, where the target of each action is any device in the environment, the action dimensionality explodes. Another solution is to serialize discrete actions, which solves the dimensionality problem but introduces inaccurate optimization results. Furthermore, since some tasks are inherently indivisible, this optimization method deviates from the true purpose of unloading. Therefore, improvements are needed in the overall deep reinforcement learning network structure, action selection, and attribute settings of existing methods. Summary of the Invention

[0007] This invention provides a handover and offloading method and apparatus based on deep reinforcement learning in a D2D environment to solve the technical problem that in a multi-base station and multi-user environment, the total number of task offloading actions increases geometrically with the number of users, and the large number of users may lead to base station congestion.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] On one hand, the present invention provides a handover and offloading method based on deep reinforcement learning in a D2D environment, the handover and offloading method based on deep reinforcement learning in a D2D environment comprising:

[0010] Establish an edge computing offloading environment that can simultaneously realize D2D offloading between users and task handover between base stations, initialize the random location relationship between users and base stations, establish the channel state relationship between users and users and between base stations, and determine the D2D offloading and base station handover offloading mechanism.

[0011] Based on the established edge computing offloading environment, the composition of offloading actions is determined, and the relationship between offloading actions and the offloading environment is established. The establishment of the relationship between offloading actions and the offloading environment includes: determining the offloading relationship between users, between users and base stations, and between base stations based on the offloading actions.

[0012] Based on the relationship between the unloading action and the unloading environment, the unloading cost is calculated. The unloading cost is used as the target to be optimized, and a reinforcement learning quadruple is constructed. Rewards and penalties are set according to the unloading cost.

[0013] Based on the constructed reinforcement learning quadruple and reward and penalty terms, a preset deep reinforcement learning algorithm is used to realize handover offloading in a D2D environment and minimize the weighted sum of latency and energy consumption for all users.

[0014] Furthermore, the offloading environment is a circular space with a radius of R, and there are M users and K base stations in the offloading environment, with each user and base station evenly distributed in the circular space;

[0015] Channel conditions include distance, bandwidth, channel transmission power, and noise power spectral density; where the distance between each device is d; the transmission bandwidth and transmission power of each interconnection relationship in the offloading environment are assumed to be constant; and the transmission bandwidth between users is B. d2d The transmission bandwidth between the user and the base station is B. mobile The transmission bandwidth between base stations is B. base The transmission power between users is P. d2d The transmission power between the user and the base station is P. mobile The transmission power between base stations is P base The noise power spectral density in the environment is n0.

[0016] The computing power of each device is set as follows: the user's local computing power is f. local The base station's computing power is f base .

[0017] Furthermore, the D2D offloading and base station handover offloading mechanism includes:

[0018] User offloading actions include local offloading, D2D offloading, and base station offloading. In local offloading, the user uses its own computing resources to perform computational tasks it generates. In D2D offloading, the communication distance is limited by a threshold value, d. th If the threshold is exceeded, D2D uninstallation cannot be performed; if the target user of D2D uninstallation has already performed local uninstallation, D2D uninstallation will fail. When D2D uninstallation cannot be performed or D2D uninstallation fails, it needs to be transferred back to the source user for local uninstallation as a penalty for incorrectly selecting the target user.

[0019] The handover process of a base station includes handing over unloading tasks to other base stations to reduce the unloading pressure on a single base station. Each base station can only hand over one task at a time to prevent excessive load on other base stations. After the unloading process is completed, each base station allocates computing resources equally to the tasks it has finally received.

[0020] Furthermore, the process of generating the unloading action, and the unloading relationships between users and base stations, and the handover relationships between base stations determined based on the unloading action, include:

[0021] Based on the uninstallation environment settings, the total number of tasks is M, and the tasks for each user i are represented as R. i ={D i ,L i}, where D i L represents the task data size. i The number of CPU cycles required for the task;

[0022] The total unloading action of the entire unloading environment is represented by A, and its dimension is equal to the total number of devices, i.e., M+K, that is, A=[A0,…,A M-1 ,…,A M+K-1 ], and user i's uninstallation action is A. i ∈[A0,…,A M-1 The handover action of base station p is A. p ∈[A M ,…,A M+K-1 ];

[0023] Each A i The following are uninstallation sub-actions:

[0024] 1)A i,local This indicates whether user i has performed a local uninstallation;

[0025] 2) It indicates whether user i performs a D2D uninstallation targeting user j;

[0026] 3) It indicates whether user i performed an incorrect D2D uninstallation action, that is, whether user i uninstalled to a user j that had already performed a local uninstallation;

[0027] 4) It indicates whether user i has performed base station unloading with base station p as the target.

[0028] All of the above sub-actions take values ​​of 0 or 1, and the sum of these sub-actions is 1; A i,local =1 indicates that user i uninstalled locally, A i,local =0 indicates that user i did not uninstall locally; This indicates that user i performed a D2D uninstallation targeting user j. This indicates that user i has not uninstalled D2D; This indicates that user i performed an incorrect D2D uninstallation action. This indicates that user i did not perform the erroneous D2D uninstallation action; This indicates that user i performed a base station offloading operation with base station p as the target. This indicates that user i has not performed base station offloading;

[0029] Each Ap There is a unique handover action. It indicates whether base station p will hand over its task to base station q, A p The value can be 0 or 1; This indicates that base station p is handing over its task to base station q. This indicates that base station p will not hand over its tasks to base station q.

[0030] Furthermore, the calculation process for the unloading overhead includes:

[0031] The total unloading delay T for each user i is calculated using the following formula. i for:

[0032]

[0033] in, This represents the transmission rate between user i and user j. This represents the transmission rate between user i and base station p. z represents the transmission rate between base station q and base station p. p z represents the final number of tasks assigned to base station p after the offloading operation is completed. q f represents the final number of tasks assigned to base station q after the offloading operation is completed. m It is local computing power. The value is 0 or 1. When the value is 1, it means that base station p will hand over its task to base station q. When the value is 0, it means that base station p will not hand over its task to base station q.

[0034] The total energy consumption E for each user's unloading is calculated using the following formula. i for:

[0035]

[0036] Among them, P com The computing power corresponding to each gigahertz of computing power;

[0037] The total uninstallation cost C for user i is calculated using the following formula. i for:

[0038] C i =η1T i +η2E i

[0039] Where η1 and η2 are the time delay and energy consumption coefficients, respectively, both of which take values ​​in the range of (0,1), and η1+η2=1.

[0040] Furthermore, the reinforcement learning quadruple is<s,a,reward,s′> Where s is the current state, a is the total unload action, reward is the reward, and s′ is the state after the action interacts with the current environment. It is the weighted sum of the total user unload cost obtained after the action interacts with the environment. When the next action interacts with the environment, s′ becomes the state s in this case.

[0041] The reward is set as the opposite of the total offload cost; there is also a penalty in the reward, which is implicitly set in the cost of each user; if a user performs an illegal action during the offload process, the cost for this action is set to a maximum value, and the corresponding reward is the negative of that maximum value. The illegal action refers to the handover action that occurs when the current base station has not received any tasks; the optimization goal is to maximize the reward.

[0042] Furthermore, a pre-defined deep reinforcement learning algorithm is employed to achieve handover and offloading in a D2D environment, minimizing the weighted sum of latency and energy consumption for all users, including:

[0043] Construct a branched, dual-depth Q-network for dueling and initialize the neural network parameters;

[0044] Train and update the Q-network, then save the model.

[0045] Furthermore, the branched dual-depth duel Q network maps each dimension's action to a set of output nodes, with each output node corresponding to each sub-action of that dimension. Finally, the action is aggregated again through an action selection layer into an M+K dimensional action. When selecting actions, an ε-greedy strategy is used, and ε decays as the training cycle increases to enhance the exploration effect.

[0046] Both the target network and the evaluation network use the same branch-depth Q-network structure.

[0047] Furthermore, the training method for the branched dual-depth duel Q-network is as follows:

[0048] After selecting an action through a neural network, the action is interacted with the environment. If the action is completed, the environment is reset; otherwise, it is not reset. Regardless of whether the action is completed or not, the resulting quadruple is stored in the memory pool for subsequent training.

[0049] The quadruplets required to train and update the network are extracted from the memory pool. When calculating the temporal difference error, the temporal difference error is calculated separately for each dimension. That is, the loss function is the mean squared expected value of the temporal difference error of each dimension's action. Then, gradient descent is applied to the loss function to update and evaluate the network parameters. The parameters of the target network are updated every certain number of steps until training is complete.

[0050] On the other hand, the present invention also provides a handover and unloading device based on deep reinforcement learning in a D2D environment, the handover and unloading device based on deep reinforcement learning in a D2D environment comprising:

[0051] The offloading environment construction module is used to establish an edge computing offloading environment that can simultaneously realize D2D offloading between users and task handover between base stations, initialize the random location relationship between users and base stations, establish the channel state relationship between users and users and between base stations, and determine the D2D offloading and base station handover offloading mechanism.

[0052] The unloading action and device relationship determination module is used to determine the composition of unloading actions and establish the interaction relationship between unloading actions and the unloading environment based on the edge computing unloading environment established by the unloading environment construction module. The establishment of the interaction relationship between unloading actions and the unloading environment includes: determining the unloading relationship between users, between users and base stations, and between base stations based on the unloading actions.

[0053] The reinforcement learning quadruple construction module is used to determine the relationship between the unloading action determined by the module and the unloading environment based on the unloading action and device relationship, calculate the unloading cost, use the unloading cost as the target to be optimized, construct reinforcement learning quadruples, and set reward and penalty items based on the unloading cost.

[0054] The reinforcement learning module is used to implement handover and offloading in a D2D environment based on the reinforcement learning quadruple constructed by the reinforcement learning quadruple construction module and the reward and penalty terms, using a preset deep reinforcement learning algorithm, and minimizing the weighted sum of latency and energy consumption for all users.

[0055] In another aspect, the present invention also provides an electronic device comprising a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described method.

[0056] In another aspect, the present invention also provides a computer-readable storage medium storing at least one instruction that is loaded and executed by a processor to implement the above-described method.

[0057] The beneficial effects of the technical solution provided by this invention include at least the following:

[0058] 1. This invention proposes a D2D handover offloading method based on deep reinforcement learning. Current edge offloading methods are mostly limited to user-to-base station structures, rarely considering offloading requirements in D2D environments, and even when they do, optimization is limited to environments with a small number of devices. This method considers the offloading needs in the context of the ever-increasing number of D2D devices and incorporates the concept of base station handover, making it better suited to the increasingly complex IoT environment.

[0059] 2. The concept of base station handover in this invention makes the base station also a sending device for offloading tasks, allowing it to participate in the offloading process itself. Traditional methods do not consider base station handover, which may result in one base station receiving a large number of tasks while other base stations remain relatively idle, causing unnecessary congestion. Base station handover can alleviate this congestion problem.

[0060] 3. Because all devices are involved in the unloading process, the total action space of this invention grows exponentially compared to existing methods, making it a large-scale discrete action problem. Currently, common reinforcement learning methods for solving large-scale discrete action problems involve making the unloading action continuous and using policy gradient methods. However, for tasks that are inherently inseparable, the optimization results of continuous action processing deviate from the original intention of discrete action unloading. Therefore, this invention combines a branch-deep Q-network to solve the difficult problem of solving such large-scale discrete action problems.

[0061] 4. To solve the problem of unloading large-scale discrete actions, this invention designs an action mapping relationship, which maps complex high-dimensional unloading actions to a set of smaller-dimensional actions that can represent all interconnections in the current environment through a transformation relationship similar to a hash function. The unloading method provided by this invention can be applied to environments where devices have more unloading paths or more complex interconnections in the future. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 This is a schematic diagram of the execution flow of the handover and unloading method based on deep reinforcement learning in a D2D environment provided in an embodiment of the present invention;

[0064] Figure 2 This is a schematic diagram of the edge computing offloading environment provided in an embodiment of the present invention;

[0065] Figure 3This is a flowchart illustrating the D2D handover and unloading method provided in this embodiment of the invention.

[0066] Figure 4 This is the cost calculation diagram provided in the embodiments of the present invention;

[0067] Figure 5 This is a neural network structure diagram provided in an embodiment of the present invention. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0069] First Embodiment

[0070] This embodiment provides a handover offloading method based on deep reinforcement learning in a D2D environment, which can be applied to wireless communication fields such as 5G communication and cellular networks. This method can be implemented by an electronic device, which can be a terminal or a server. The execution flow of this method is as follows: Figure 1 As shown, it includes the following steps:

[0071] S1. Establish an edge computing offloading environment that can simultaneously realize D2D offloading between users and task handover between base stations, initialize the random location relationship between users and base stations, establish the channel state relationship between users and users and between base stations, and determine the D2D offloading and base station handover offloading mechanism.

[0072] S2, based on the established edge computing offloading environment, determine the composition of the offloading action and establish the interaction relationship between the offloading action and the offloading environment; the establishment of the interaction relationship between the offloading action and the offloading environment includes: determining the offloading relationship from user to user, from user to base station, and the handover relationship from base station to base station based on the offloading action.

[0073] S3. Calculate the unloading cost based on the interaction between the unloading action and the unloading environment. Use the unloading cost as the target to be optimized and construct a reinforcement learning quadruple. Set reward and penalty items based on the unloading cost.

[0074] S4, based on the constructed reinforcement learning quadruple and reward and penalty terms, uses a preset deep reinforcement learning algorithm to achieve handover offloading in a D2D environment and minimize the weighted sum of latency and energy consumption for all users.

[0075] Specifically, the uninstallation environment is as follows: Figure 2As shown, the scope is a circular space with radius R, containing M users and K base stations, all evenly distributed within this space. The set of users is denoted by M, and the set of base stations by K. To more clearly represent the relationships between devices, this method uses i,j (i∈M, j∈M) to represent a pair of random user devices, and p,q (p∈K, q∈K) to represent a pair of random base station devices. The scope of this environment, the number of users, and the number of base stations can be adjusted according to specific circumstances to conform to actual conditions.

[0076] Channel conditions include distance, bandwidth, channel transmission power, and noise power spectral density. The distance between each device is d. To simplify the model, the transmission bandwidth and transmission power for each interconnection relationship in this environment are assumed to be constant. Therefore, the transmission bandwidth between users is B. d2d The transmission bandwidth between the user and the base station is B. mobile The transmission bandwidth between base stations is B. base The transmission power between users is P. d2d The transmission power between the user and the base station is P. mobile The transmission power between base stations is P. base The noise power spectral density in the environment is n0. Distance is measured in meters (m), bandwidth in megahertz (MHz), transmission power in decibels relative to one milliwatt (dBm), and noise power spectral density in dBm / Hz. P com This refers to the computing power equivalent to each gigahertz (GHz) of computing power, expressed in watts (W). The computing power of each device is set as follows: the user's local computing power is f... local The base station's computing power is f base The unit of computing power is gigahertz (GJ).

[0077] In practice, communication is possible between all devices. Users can perform local offloading, D2D offloading, and base station offloading. This embodiment does not specifically limit the type of offloading or the destination of each user's offloading action; however, when calculating the specific offloading overhead, it is necessary to use encoding to exclude infeasible behaviors. This restriction method will be described below. In local offloading, users can use their own computing resources to perform computational tasks they generate. In D2D offloading, communication is subject to a threshold limit, which is d. th If the threshold is exceeded, D2D offloading cannot be performed. If the target user for D2D offloading has already performed local offloading, D2D offloading will fail. In either of these cases, the data needs to be transmitted back to the source user for local offloading as a penalty for incorrect target user selection. Alternatively, users can offload the task to the base station.

[0078] Base stations can hand over unloading tasks to reduce the unloading pressure on individual base stations. In this embodiment, a base station can hand over a maximum of one task at a time to prevent excessive load on other base stations. Each base station can hand over the computing task with the highest computing power requirement among the computing tasks it undertakes to other base stations. After the unloading process is completed, each base station will evenly distribute computing resources to the tasks it finally receives.

[0079] The specific uninstallation process is as follows: Figure 3 As shown. After the environment is built, the uninstallation action needs to be determined. The process of generating the uninstallation action, and determining the uninstallation and handover relationship based on the uninstallation action, is as follows:

[0080] Based on the environment settings, the total number of tasks is M, which is the same as the number of users. The task for each user i can be represented as R. i ={D i ,L i}(i∈M), where D i The task data size is measured in bits, while L... i The number of CPU cycles required for the task, in units; this method does not set latency limits for each task, but optimizes the selection of the target user through a reward-penalty system;

[0081] In this embodiment, the action value range is determined by the total number of devices in the environment. The action value range in this method is determined by the total number of devices in the environment. The action dimension equals the sum of the number of devices, i.e., (M+K), and the value range of each dimension is [0, M+K-1]. The unloading action in this method represents the interconnection relationship between various devices, rather than a yes / no unloading relationship.

[0082] A connection value in the range [0, M-1] indicates that the user performed a D2D connection or local uninstallation. This action includes performing a local uninstallation, performing a D2D uninstallation, or selecting an incorrect D2D uninstallation option, requiring a return to local uninstallation.

[0083] Interconnection relationship values ​​within the range [M, M+K-1] indicate that users are interconnecting with base stations. This action includes situations where the base station performs handover / offload, the base station is unable to perform handover / offload, or the base station does not need to perform handover / offload and only needs to process received tasks.

[0084] As mentioned above, in this method, each dimension's complex sub-actions are represented by a single positive integer, and the relationships within these sub-actions may be restricted, described, and calculated using encoding rules. Therefore, it is evident that using {0,1} type unloading actions in the context of this method would require (M+K). (M+K)Describing each action individually is extremely difficult to optimize. This method, by decomposing the action space, allows us to use only (M+K) actions. 2 All actions can be represented by a single action, greatly reducing the computational power required for optimization.

[0085] Therefore, in this embodiment, the specific rules for setting and dividing the action range are as follows:

[0086] The total uninstallation action of the entire uninstallation environment is represented by A, with dimensions (M+K), indicating a total of (M+K) devices. That is, A = [A0,…,A0]. M-1 ,…,A M+K-1 The uninstallation action of user i (i∈M) is A. i ∈[A0,…,A M-1 The handover action of base station p (p∈K) is A. p ∈[A M ,…,A M+K-1 ];

[0087] Each A i The following are uninstallation sub-actions:

[0088] 1)A i,local This indicates whether user i has performed a local uninstallation;

[0089] 2) (i,j∈M), which indicates whether user i performs D2D uninstallation with user j as the target;

[0090] 3) This indicates whether user i performed an incorrect D2D uninstallation action, that is, whether user i uninstalled to a user j that had already performed a local uninstallation.

[0091] 4) (i∈M,p∈K), which indicates whether user i has performed base station offloading with base station p as the target;

[0092] All the above sub-actions take values ​​of 0 or 1, and the sum of these sub-actions is 1; specifically, in this embodiment, A i,local =1 indicates that user i uninstalled locally, A i,local =0 indicates that user i did not uninstall locally; This indicates that user i performed a D2D uninstallation targeting user j. This indicates that user i has not uninstalled D2D; This indicates that user i performed an incorrect D2D uninstallation action. This indicates that user i did not perform the erroneous D2D uninstallation action; This indicates that user i performed a base station offloading operation with base station p as the target. This indicates that user i has not performed base station offloading;

[0093] Each A p There is a unique handover action. (p,q∈K), which indicates whether base station p will hand over its task to base station q, A p The value of is 0 or 1; in this embodiment, This indicates that base station p is handing over its task to base station q. This indicates that base station p will not hand over its tasks to base station q.

[0094] It should be noted that the values ​​of these unloading sub-actions or base station transmission sub-actions are all 0 or 1, but the action values ​​of each user and base station in this method are obviously not {0,1} combinations, but integers in [0,M+K-1]. Therefore, a mapping transformation is required to convert the action values ​​of user equipment and base stations into {0,1} combinations corresponding to these sub-actions in the actual calculation. The term "actual action" will be used below to refer to the transformed action.

[0095] User uninstallation action A i The value range of (i∈M) is [0,M+K-1]. This range can be further subdivided, namely [0,M-1] for user internal selection and [M,M+K-1] for base station offloading selection.

[0096] Let's take an arbitrary set of user devices i,j (i,j∈M) as an example for illustration. When A i When the value of is in [0, M-1], user i may have two uninstallation options.

[0097] The first scenario: The user's action value is equal to the user's ID, i.e., A. i =i, in this case, user i performs a local uninstallation, the actual action is A. i,local =1.

[0098] The second scenario: The user's action value is not equal to the user ID, i.e., A. i =j (i,j∈M,i≠j). In this case, user i performs D2D unloading, and the actual action is... In this D2D offloading scenario, if the distance between the two users exceeds the transmission threshold, or if user i performs D2D offloading but ends up offloading to a user j who has already performed local offloading, i.e., A... i =j,A j If the value is j, then user i cannot perform D2D uninstallation and needs to perform a postback to return to the local machine for uninstallation. These two states represent incorrect D2D uninstallation actions, and the actual action at this time is... .

[0099] When A iWhen the value of is in [M, M+K-1], user i will offload the task to base station p (p∈K). The actual action at this time is... (i∈M, p∈K).

[0100] Base station handover action A p The value range remains [0, M+K-1]. Maintaining the same value range for base stations as for user equipment facilitates the construction of branch-depth Q-networks. Simultaneously, since the number of base stations is K, and K is generally less than M, it is necessary to construct a branch-depth Q-network from A... p Mapping to K.

[0101] Let's take an arbitrary set of base stations p and q (p, q ∈ K) as an example for illustration. For base station action A... p Perform the following processing: Calculation And the remainders, we can obtain remainders in the range [0, K-1]. These remainders are then cyclically numbered, corresponding to base stations numbered [0, K-1]. If A p If +1 is not divisible by K, the remaining number can be used to continue numbering sequentially or masked as an illegal action, without affecting the action selection. This yields the desired mapping relationship. This indicates the base station handover action after the mapping transformation.

[0102] Not all base station handover actions can be performed. Invalid handover actions need to be excluded. The exclusion method is: if base station p has not received any tasks and its handover indication action is not directed to itself (i.e., ... In this case, it means that the base station has not received a task and has no tasks to hand over. Therefore, it cannot participate in the handover action designated to other base stations, and the actual handover action of this base station needs to be marked as invalid. If the action is invalid, a handover invalid tag is added. This indicates that base station p does not actually hand over the connection to base station q. In this case, the actual action of the base station is... .

[0103] After eliminating invalid actions, the corresponding actual actions can be obtained based on the mapped actions. When and At that time, actual actions Conversely .

[0104] Since it's common for multiple users to offload to a single base station, it's crucial to determine which task to perform each handover. This embodiment selects the task with the highest computational resource requirement among those received by each base station for handover.

[0105] Simultaneously, for ease of calculating the total cost later, it is necessary to pre-calculate the final number of tasks z allocated to each base station after the offloading operation is completed. Calculating the number of tasks requires determining in advance which users participated in base station offloading. This can be done using the set N = {N0, N...}. p ...,N K-1}(p∈K) represents the number of tasks initially assigned to each base station after the action is selected.

[0106] For base station p, N p It can be divided into two components. The first is the users who are directly offloaded to base station p, i.e. The second is other base stations that are handed over to base station p, i.e. Therefore, we can conclude that... Where i∈M, p∈K, q∈K. Considering that some base station handover actions are invalid, the final number of tasks z assigned to each base station can be calculated as follows: p :

[0107]

[0108] In the formula, 1 is an indicator function, indicating that the function value is 1 when the condition inside the parentheses is met. This indicates the number of base stations that committed illegal actions among all base stations that were handed over to base station p. This represents the scenario where base station p hands over the tasks it received. If it hands over the tasks, the number of tasks received by base station p needs to be subtracted. After removing these two scenarios, we can obtain the final number of tasks allocated to each base station after the handover and allocation process.

[0109] Furthermore, after determining the action, the specific unloading overhead needs to be calculated to construct the quadruple.

[0110] according to Figure 4 The unloading overhead consists of the following components, and the calculation process for the unloading overhead is as follows:

[0111] The offloading delay for user i (i∈M) consists of local offloading delay, D2D offloading delay, D2D error offloading delay, and base station offloading delay. The sub-action of user i (i∈M) is A. i,local , (i,j∈M) , (i∈M, p∈K), corresponding to these four time delay configurations respectively. The base station handover action is as follows: (p,q∈K) indicates whether base station p will hand over the task of base station p to base station q.

[0112] When A i,local When = 1, user i performs local uninstallation, with a latency of f in the formulalocal It refers to local computing power.

[0113] when When (i,j∈M), user i performs D2D offloading, and the latency consists of the D2D transmission latency and the D2D processing latency. The transmission rate between users i and j (i,j∈M) can be calculated using Shannon's formula. B d2d It is the communication bandwidth between two users, P d2d It refers to the communication power between the two users, and the path loss. It can be done To calculate, where This refers to the distance between two users. The calculation method for the transmission rate between base stations and between users and base stations, as described below, is the same as the calculation method for the transmission rate between users. Simply substitute the corresponding device distance, transmission power, and bandwidth into the calculation.

[0114] when When (i,j∈M), user i incorrectly performed a D2D selection targeting user j (e.g., offloading user i to a base station outside the D2D transmission range), therefore, the time for the task to be transmitted back from user j to user i is calculated additionally. The offloading delay for user i at this time is: .

[0115] when When (i∈M, p∈K)=1, user i performs base station offloading. In this method, p (p∈K) represents the target base station of user i. q (q∈K) represents the handover object of base station p. There are two possibilities for base station offloading; this method considers the base station action... This indicates that one possibility is that base station p does not hand over to base station q. Then, the base station offloading delay for user i consists only of the offloading transmission delay and the base station processing delay. Another possibility is that base station p is handed over to base station q, in which case... Then, the base station offloading delay for user i consists of the offloading transmission delay, the handover transmission delay, and the base station processing delay.

[0116] Therefore, based on the above analysis, the base station offload delay for user i can be obtained as follows: Among them, f base It is the computing power of the base station, z p (p∈K) and z q (q∈K) represents the number of tasks that base station p and base station q are ultimately allocated after the offloading operation is completed. This represents the transmission rate between user i and base station p. This represents the transmission rate between base station p and base station q.

[0117] The total uninstallation latency for each user i can be obtained from the above formula as follows:

[0118]

[0119] The sum of the four uninstallation destination actions for each user i is 1 (i.e.,

[0120] The offloading energy consumption of user i (i∈M) consists of local offloading energy consumption, D2D offloading energy consumption, D2D error offloading energy consumption, and base station offloading energy consumption. Energy consumption is calculated by multiplying the delay by the corresponding power. For example, transmission energy consumption is the product of transmission delay and transmission power, and processing energy consumption is the product of processing delay and processing power.

[0121] The energy consumption for local uninstallation by user i is:

[0122] The energy consumption for user i's D2D unloading is:

[0123] The energy consumption for user i's D2D error offloading is:

[0124] The base station offloading energy consumption for user i is:

[0125]

[0126] Combining the above formulas, the total energy consumption for uninstallation per user is:

[0127]

[0128] Therefore, the total uninstallation cost per user can be calculated as C. i =η1T i +η2E i .

[0129] Among them, T i E represents the total uninstallation delay for user i. i Let η1 and η2 be the total energy consumption for user i to unload. η1 and η2 are the delay and energy consumption coefficients, respectively, both ranging from (0,1), and η1+η2=1.

[0130] Furthermore, that is<s,a,reward,s′> Where s is the current state, a is the total unloading action A, reward is the reward, and s′ is the state after the action interacts with the current environment. In this method, s′ is the weighted sum of the total user overhead obtained after the action interacts with the environment, i.e. When the next action is taken and the interaction with the environment occurs, s′ becomes the state s in this situation. The reward is set as the negative of the total cost, i.e. With this setup, the lower the cost, the greater the reward.

[0131] The rewards also include penalties, which are implicitly set in the overhead of each user. If a user performs an illegal action during the uninstallation process, the overhead is set to a maximum value, and the corresponding reward is a negative of that maximum value, thus significantly reducing the reward. Such illegal actions generally refer to handover actions occurring when the base station has not received any tasks. These actions are not only interconnected but also difficult to exclude through hard coding, so the reward settings are used to avoid these actions. The optimization goal is to maximize the reward, that is, to find the optimal reward while avoiding illegal actions. question.

[0132] Furthermore, in this embodiment, a preset deep reinforcement learning algorithm is used to achieve handover offloading in a D2D environment and minimize the weighted sum of latency and energy consumption for all users, specifically including:

[0133] Construct a branched, dual-depth Q-network for dueling and initialize the neural network parameters;

[0134] Train and update the Q-network, then save the model.

[0135] The neural network structure of this method is as follows: Figure 5 As shown, its structural concept and training method are as follows:

[0136] As the number of users capable of D2D and wireless transmission capabilities continues to increase in reality, the growth in user numbers has led to a rapid expansion of action dimensions. This method addresses this issue using a branch-deep Q-network approach. The branch-deep Q-network maps each dimension's action to a set of output nodes, which in turn correspond to each sub-action within that dimension. Both the target network and the evaluation network employ the same branch-deep Q-network structure. Furthermore, to improve training performance, this embodiment also incorporates a duel structure and combines it with a dual-deep Q-network approach to mitigate the estimation problem inherent in over-Q networks.

[0137] Figure 5 This is a neural network structure diagram according to an embodiment of the present invention. Figure 5As shown, the neural network input in this embodiment is the current system state. It first passes through two fully connected hidden layers with P neurons and ReLU activation function. According to the duel structure, the action value function, i.e., the Q function, needs to be decomposed into the sum of the state value function and the advantage function. The state value function is a scalar, so it needs to be output through a separate neural network layer. The advantage function needs to be output through an advantage function layer, producing an (M+K)×(M+K) dimensional advantage function, representing (M+K) sub-actions in (M+K) dimensions. This advantage function is then aggregated with the state value function through an aggregation layer to form the Q function. Thus, the Q function is also (M+K)×(M+K) dimensional. Then, by calculating the argmax of the Q value for each dimension and aggregating them again, the required action value can be obtained.

[0138] After the neural network structure is set up, the method for training the network and updating the parameters is as follows:

[0139] For each dimension, this method employs an ε-greedy strategy for action selection, with ε decaying as the number of training steps increases to enhance the exploration effect. After selecting actions through the neural network, they are interacted with the environment. If the action is successfully executed, the environment is reset; otherwise, it is not. Regardless of whether the action is completed, the resulting quadruples are stored in a memory pool for subsequent training. The quadruples required for training and updating the network are extracted from the memory pool. When calculating the action value function, constructing the temporal difference objective, and calculating the temporal difference error, the quadruples in the formulas are all selected from the memory pool.

[0140] When calculating the temporal difference error, the temporal difference error is calculated separately for each dimension, and the Duel Q network algorithm is combined to improve the convergence speed and optimization effect. When calculating the total loss function, the mean square value of the loss of all dimensions is used as the total loss function, and the target network is updated at certain intervals to save the training model.

[0141] As described above, in the case of a duel-Q network, the behavior value function of each dimension w can be decomposed into the sum of the state value function and the advantage function. The behavior value function of each dimension is as follows:

[0142] Q w (s,a w )=V(s)+X w (s,a w ).

[0143] Where w∈[0,M+K-1] represents the various dimensions of the total action, a w It evaluates the w-th dimension sub-action in the network, Q w X is the behavior-value function that evaluates the w-th dimensional sub-action in the network. wLet be the advantage function of the w-th sub-action. The advantage functions of each dimension need to be constrained; the advantage functions of all actions using a greedy strategy will be set to zero to optimize training performance. That is, the advantage function of each dimension needs to be subtracted from the average advantage function of the actions in each branch dimension.

[0144] The behavior value function for each dimension after constraint is:

[0145]

[0146] Construct a temporal difference objective based on a dual-depth Q-network approach:

[0147] The objective of time-series difference is:

[0148] in, Let be the behavior value function of the w-th dimension sub-action in the target network. γ∈[0,1] is a discount factor, representing the importance of future rewards; the closer γ is to 1, the more important the future rewards are. The loss function is then calculated based on this:

[0149] The loss function is:

[0150] The loss function is the expected mean square of the temporal difference errors of each dimension of the action. Then, gradient descent is applied to the loss function to update the evaluation network parameters. The parameters of the target network are updated at regular intervals.

[0151] In summary, the handover offloading method based on deep reinforcement learning in the D2D environment provided in this embodiment, based on the real-world scenario of multiple users, multiple base stations, and multiple D2D users, establishes a D2D edge heterogeneous network structure capable of handover between users and between base stations; establishes random location relationships between users and base stations; establishes channel state relationships between users and between base stations; determines D2D offloading and base station handover offloading mechanisms; determines the reward function by comprehensively considering offloading energy consumption and latency; determines the penalty by combining internal and external constraints; and establishes a branched deep Q-network, combined with a duel structure to approximate the state-action function, trains the network, and ultimately achieves the effect of minimizing the weighted sum of latency and energy consumption for all users. Furthermore, addressing the current problem of multiple users, multiple base stations, and difficulty in discretizing offloading actions, this embodiment improves the structure of the deep Q-network, decomposing high-dimensional discrete actions into lower-dimensional action combinations, greatly simplifying the problem and facilitating implementation.

[0152] Second Embodiment

[0153] This embodiment provides a handover and unloading device based on deep reinforcement learning in a D2D environment, including:

[0154] The offloading environment construction module is used to establish an edge computing offloading environment that can simultaneously realize D2D offloading between users and task handover between base stations, initialize the random location relationship between users and base stations, establish the channel state relationship between users and users and between base stations, and determine the D2D offloading and base station handover offloading mechanism.

[0155] The unloading action and device relationship determination module is used to determine the composition of unloading actions and establish the interaction relationship between unloading actions and the unloading environment based on the edge computing unloading environment established by the unloading environment construction module. The establishment of the interaction relationship between unloading actions and the unloading environment includes: determining the unloading relationship between users, between users and base stations, and between base stations based on the unloading actions.

[0156] The reinforcement learning quadruple construction module is used to determine the relationship between the unloading action determined by the module and the unloading environment based on the unloading action and device relationship, calculate the unloading cost, use the unloading cost as the target to be optimized, construct reinforcement learning quadruples, and set reward and penalty items based on the unloading cost.

[0157] The reinforcement learning module is used to implement handover and offloading in a D2D environment based on the reinforcement learning quadruple constructed by the reinforcement learning quadruple construction module and the reward and penalty terms, using a preset deep reinforcement learning algorithm, and minimizing the weighted sum of latency and energy consumption for all users.

[0158] The handover and unloading device based on deep reinforcement learning in the D2D environment of this embodiment corresponds to the handover and unloading method based on deep reinforcement learning in the D2D environment described above. The functions implemented by each functional module in the handover and unloading device based on deep reinforcement learning in the D2D environment correspond one-to-one with the process steps in the handover and unloading method based on deep reinforcement learning in the D2D environment described above. Therefore, it will not be described again here.

[0159] Third Embodiment

[0160] This embodiment provides an electronic device, which includes a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment.

[0161] The electronic device can vary considerably depending on its configuration or performance, and may include one or more processors (central processing units, CPUs) and one or more memories, wherein the memories store at least one instruction that is loaded by the processor and executed in accordance with the above method.

[0162] Fourth embodiment

[0163] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc. The instruction stored therein can be loaded and executed by a processor in a terminal.

[0164] Furthermore, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0165] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0166] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0167] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0168] Finally, it should be noted that the above description represents a preferred embodiment of the present invention. It should be pointed out that although preferred embodiments have been described, those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles described herein. These improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A handover offloading method based on deep reinforcement learning in a D2D environment, characterized in that, The handover offloading method based on deep reinforcement learning in the D2D environment comprises the following steps: An edge computing offloading environment is established, which can simultaneously realize D2D offloading among users and task handover among base stations, and the random position relationship between users and base stations is initialized, the channel state relationship between users and base stations is established, and the D2D offloading and base station handover offloading mechanism is determined; Based on the established edge computing offloading environment, the composition of offloading actions is determined, and the relationship between offloading actions and offloading environment interaction is established; the relationship between offloading actions and offloading environment interaction comprises: determining the offloading relationship between users and users, the offloading relationship between users and base stations, and the handover relationship between base stations according to the offloading actions; According to the relationship between offloading actions and offloading environment interaction, the offloading cost is calculated, the offloading cost is taken as the target to be optimized, and the reinforcement learning four-tuple is constructed; and the reward and penalty items are set according to the offloading cost; Based on the constructed reinforcement learning four-tuple and the reward and penalty items, a preset deep reinforcement learning algorithm is used to realize the handover offloading in the D2D environment, and the weighted sum of all user delays and energy consumptions is minimized; The calculation process of the offloading cost comprises: The total offload latency T for each user i is calculated by the following formula i is: wherein A i,local denotes whether user i performs local offloading; denotes whether user i performs D2D offloading with user j as the offloading target; denotes whether user i performs an erroneous D2D offloading action; denotes whether user i performs base station offloading with base station p as the offloading target; i is the number of CPU cycles required for the task; local is the user's local computing capability; base is the base station's computing capability; i is the task data size; denotes the transmission rate between user i and user j, denotes the transmission rate between user i and base station p, denotes the transmission rate between base station q and base station p, p denotes the number of tasks that base station p finally gets after the offloading action ends, q denotes the number of tasks that base station q finally gets after the offloading action ends, m is the local computing capability, takes the value of 0 or 1, and when the value is 1, it indicates that base station p will hand over its tasks to base station q, and when the value is 0, it indicates that base station p will not hand over its tasks to base station q; The total energy consumption E of each user is calculated by the following formula i is: where P com is the computing power corresponding to the computing capability per gigahertz; The total offload cost C for user i is calculated by the following formula i is: C i = η1T i + η2E i Wherein, η1, η2 are delay and energy consumption coefficients, the value range is (0, 1), and η1+η2=1; The reinforcement learning four-tuple is <s, a, reward, s'>; wherein, s is the state at the current time, a is the total offloading action, reward is the reward, s' is the state after the action and the current time environment interaction, which is the weighted sum of the total offloading cost of the users after the action and the environment interaction, and when the next action and environment interaction is performed, s' becomes the state s in this case; The reward is set as the opposite number of the total offloading cost; the reward also has a penalty item, which is implicitly set in the offloading cost of each user; if an illegal action occurs in the offloading action, the offloading cost of this time is set as a maximum value, and the corresponding reward is the negative value of the maximum value, the illegal action refers to the handover action of the current base station without receiving any task; the optimization target is to maximize the reward. 2.The handover offloading method based on deep reinforcement learning in D2D environment of claim 1, wherein, The offloading environment is a circular space with a radius R, there are M users and K base stations in the offloading environment, and each user and base station is uniformly distributed in the circular space; The channel state includes distance, bandwidth, channel transmission power and noise power spectral density; wherein the distance between each device is d; the transmission bandwidth and transmission power of each interconnection relationship in the offloading environment are assumed to be constant; the transmission bandwidth between users is B d2d , the transmission bandwidth between users and base stations is B mobile , the transmission bandwidth between base stations is B base ; the transmission power between users is P d2d , the transmission power between users and base stations is P mobile , the transmission power between base stations is P base ; the noise power spectral density in the environment is n0; The computing power of each device is set as follows: the user local computing power is f local , and the base station computing power is f base . 3.The handover offloading method based on deep reinforcement learning in D2D environment of claim 2, wherein, The D2D offloading and base station handover offloading mechanism comprises: The user's offloading action includes local offloading, D2D offloading, and base station offloading; in the local offloading, the user uses own computing resource to calculate the computing task generated by the user; in the D2D offloading, the communication distance has a threshold limit, and the threshold is d th , and the D2D offloading cannot be performed when the threshold is exceeded; if the target user of the D2D offloading has performed the local offloading, the D2D offloading fails, and when the D2D offloading cannot be performed or the D2D offloading fails, the source user needs to be transmitted back to perform the local offloading, as a punishment for the wrong selection of the target user. The handover action of the base station comprises offloading tasks to other base stations to reduce the offloading pressure of a single base station; wherein, the base station offloads at most one task each time to prevent other base stations from being overloaded; after the offloading behavior ends, each base station evenly allocates the computing resources to the tasks finally received. 4.The handover offloading method based on deep reinforcement learning in D2D environment of claim 3, wherein, The generation process of the offloading action, and the offloading relationship between users and users, the offloading relationship between users and base stations, and the handover relationship between base stations determined according to the offloading action comprise: According to the offloading environment setting, the total number of tasks is M, and the task of each user i is represented as R i = {D i , L i}, wherein D i is the task data size, and L i is the number of CPU cycle numbers required by the task The total offloading action of the whole offloading environment is denoted as A, whose dimension is equal to the total number of devices, i.e., M+K, i.e., A=[A0,...,A M-1 ,...,A M+K-1 ], and the offloading action of user i is A i ∈[A0,...,A M-1 ], the handover action of base station p is A p ∈[A M ,...,A M+K-1 ]. each A i There are the following offloading sub-actions: 1) A i,local which indicates whether user i performs local offloading or not; 2) which indicates whether user i performs D2D offloading with user j as the offloading target; 3) which indicates whether the user i has made a wrong D2D offloading action, i.e. whether the user i offloaded to a user j who has already made a local offloading. 4) which indicates whether or not the user i has performed base station offloading with the base station p as the target; All the above sub-action values are 0 or 1, and the sum of these sub-actions is 1; A i,local = 1 indicates that user i performs local offloading, i,local = 0 indicates that user i does not perform local offloading; indicates that user i performs D2D offloading with user j as the offloading target, indicates that user i does not perform D2D offloading; indicates that user i performs an error D2D offloading action, indicates that user i does not perform an error D2D offloading action; indicates that user i performs base station offloading with base station p as the offloading target, indicates that user i does not perform base station offloading; Each A p has a unique handover action which indicates whether base station p will hand over its tasks to base station q, A p has a value of 0 or 1. indicates that base station p will hand over its tasks to base station q, indicates that base station p will not hand over its tasks to base station q. 5.The handover offloading method based on deep reinforcement learning in D2D environment of claim 1, wherein, Using a preset deep reinforcement learning algorithm, the handover offloading in the D2D environment is realized, and the weighted sum of all user delays and energy consumptions is minimized, which comprises: A double deep duel Q network with a branch structure is built, and the neural network parameters are initialized; The Q network is trained and updated, and the model is saved.

6. The handover offloading method based on deep reinforcement learning in the D2D environment according to claim 5, wherein, The double-depth duel Q network of the branch structure corresponds each dimension of action to a group of output nodes, each output node of the group corresponds to each sub-action of the dimension corresponding to the dimension, and is finally aggregated into an M+K-dimensional action through the action selection layer again; when performing action selection, an ε-greedy strategy is used to select the action, and ε is decayed with the increase of the training period to enhance the exploration effect; The target network and the evaluation network both adopt the same branch-depth Q network structure.

7. The method of claim 6, wherein, The training method of the double-depth duel Q network of the branch structure is as follows: After selecting the action through the neural network, the action is interacted with the environment, if the unloading action is completed, the environment is reset, if not, the environment is not reset; whether the action is completed or not, the four-tuple obtained through the interaction is stored in the memory pool for subsequent training; The four-tuple required for training and updating the network is extracted from the memory pool, and when calculating the time difference error, the time difference error of each dimension is calculated separately, that is, the loss function is the mean square expectation value of the time difference error of each dimension action, and then the gradient descent method is used to update the evaluation network parameters, and the parameters of the target network are updated every certain number of steps until the training is completed.

8. A handover offloading device based on deep reinforcement learning in a D2D environment, characterized in that, The D2D environment-based deep reinforcement learning handover unloading device comprises: An unloading environment construction module is configured to establish an edge computing unloading environment capable of simultaneously realizing D2D unloading among users and task handover between base stations, initialize random position relationships of the users and the base stations, establish channel state relationships between the users and between the base stations, and determine D2D unloading and base station handover unloading mechanisms; An unloading action and device relationship determination module is configured to determine a composition of unloading actions based on the edge computing unloading environment established by the unloading environment construction module, and establish an interaction relationship between the unloading actions and the unloading environment; the establishment of the interaction relationship between the unloading actions and the unloading environment includes determining user-to-user, user-to-base station unloading relationships and base station-to-base station handover relationships according to the unloading actions; A reinforcement learning four-tuple construction module is configured to calculate unloading overhead based on the interaction relationship between the unloading actions and the unloading environment determined by the unloading action and device relationship determination module, take the unloading overhead as a target to be optimized, and construct a reinforcement learning four-tuple; and set rewards and penalty items based on the unloading overhead; A reinforcement learning module is configured to use a preset deep reinforcement learning algorithm to realize handover unloading in a D2D environment and minimize the weighted sum of all user delays and energy consumptions based on the reinforcement learning four-tuple constructed by the reinforcement learning four-tuple construction module and the rewards and penalty items; The calculation process of the unloading overhead includes: The total offload latency T for each user i is calculated by the following formula i is: wherein A i,local represents whether user i performs local offloading; represents whether user i performs D2D offloading whose offloading target is user j; represents whether user i performs error D2D offloading action; represents whether user i performs base station offloading whose offloading target is base station p;L i is the number of CPU cycles required for the task;f local is the user local computing capability;f base is the base station computing capability;D i is the task data size; represents the transmission rate between user i and user j, represents the transmission rate between user i and base station p, represents the transmission rate between base station q and base station p,z p represents the final number of tasks that base station p finally gets after the offloading action,z q represents the final number of tasks that base station q finally gets after the offloading action,f m is the local computing capability, takes the value of 0 or 1, and when the value is 1, it represents that base station p will hand over its tasks to base station q, and when the value is 0, it represents that base station p will not hand over its tasks to base station q; The total energy consumption E of each user is calculated by the following formula i is: where P com is the computing power corresponding to the computing capability per gigahertz; The total offload cost C for user i is calculated by the following formula i is: C i = η1T i + η2E i wherein η1 and η2 are delay and energy consumption coefficients, and the value ranges of η1 and η2 are both (0, 1), and η1+η2=1. The reinforcement learning four-tuple is <s, a, reward, s'>; wherein, s is the state at the current time, a is the total offloading action, reward is the reward, and s' is the state after the action interacts with the environment at the current time, which is the weighted sum of the total offloading cost of the user after the action interacts with the environment, and s' becomes the state s in this case when the next action interacts with the environment; The reward is set as the opposite number of the total offloading cost; the reward also has a penalty term, which is implicitly set in the cost of each user; if the user appears an illegal action in the offloading action, the cost of this time is set as a maximum value, and the corresponding reward is the negative value of the maximum value, the illegal action refers to the handover action of the current base station in the case of not receiving any task; and the optimization goal is to maximize the reward.

Citation Information

Patent Citations

  • METHOD AND APPARATUS FOR DIFFERENTIALLY OPTIMIZING QUALITY OF SERVICE QoS

    US20220400062A1