A greedy DDPG collaborative offloading decision and resource allocation method

CN117692967BActive Publication Date: 2026-09-11GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311557038.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2026-09-11
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

[0003]本发明的目的在于提供一种基贪婪DDPG协同卸载决策与资源分配方法,旨在解决UDN-MEC场景中,用户终端有限的算力和能量资源不能满足用户计算密集型和时延敏感型任务的需求的技术问题,消除微基站覆盖区域用户数量的时变性对卸载决策的影响

Benefits of technology

[0034] This invention provides a greedy DDPG collaborative offloading decision-making and resource allocation method. Addressing the difficulty in adapting to frequent user handovers within and outside the micro base station coverage area during computational offloading in UDN-MEC scenarios, this method utilizes a greedy algorithm to distribute and obtain individual user benefits. A global benefit threshold is used to describe the collaborative behavior among individual users based on global awareness. This overcomes the impact of time-varying user numbers caused by frequent handovers between micro base stations on offloading decisions, as well as the lack of cooperation between different agents in distributed single-agent methods. Furthermore, the DDPG algorithm intelligently adjusts the global benefit threshold, improving the algorithm's adaptability to dynamic changes in user numbers and reducing the latency- and energy-weighted total cost in dynamically heterogeneous UDN-MEC systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117692967B_ABST
    Figure CN117692967B_ABST
Patent Text Reader

Abstract

This invention relates to the field of edge computing technology, specifically to a greedy DDPG collaborative offloading decision-making and resource allocation method. Addressing the difficulty in adapting to frequent user handovers within and outside the micro base station coverage area during the computational offloading solution process in ultra-dense edge computing (UDN-MEC) scenarios, this method utilizes a greedy algorithm to distribute and obtain individual user benefits. A global benefit threshold is used to describe the collaborative behavior among individual users based on global collaborative perception. This overcomes the impact of time-varying user numbers caused by frequent user handovers between micro base stations on offloading decisions, as well as the lack of cooperation between different agents in distributed single-agent methods. Furthermore, the DDPG algorithm intelligently adjusts the global benefit threshold, improving the algorithm's adaptability to dynamic changes in user numbers and reducing the latency- and energy-weighted total cost in dynamically heterogeneous UDN-MEC systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of edge computing technology, specifically to a greedy DDPG collaborative offloading decision and resource allocation method. Background Technology

[0002] In recent years, while significant research has been conducted on Mobile Edge Computing (MEC) in areas such as computation offloading, resource allocation, and latency and energy consumption optimization, research on its application in ultra-dense heterogeneous cellular network (UDN-MEC) scenarios remains limited. In UDN-MEC scenarios, an improved beetle whisker algorithm can maximize system benefits with low complexity requirements by solving computation offloading strategies. An improved particle swarm optimization algorithm can alleviate the increased system latency and energy consumption caused by deploying MEC servers in 5G communication scenarios. However, in each iteration, intelligent algorithms such as the beetle whisker algorithm and particle swarm optimization need to interact with the UDN-MEC environment model, thus consuming substantial computing power in solving offloading decisions. Heuristic greedy offloading schemes can solve the computation offloading problem in UDN-MEC scenarios in a distributed manner; however, lacking a global perspective and cooperative behavior, they can only find locally optimal solutions. Multi-agent reinforcement learning methods can enhance users' global perspective and collaborative behavior. However, the frequent switching of users within and outside the micro base station coverage area leads to time-varying user numbers within the micro base station coverage area, making it difficult for multi-agent reinforcement learning methods to adapt to such dynamic environments. Distributed single-agent methods can adapt to dynamic changes in the number of users, but single agents often make decisions independently, lacking cooperation and communication between different single agents. Therefore, in the context of frequent micro base station switching by multiple users in UDN-MEC scenarios, solving the offloading decision and resource allocation while adapting to dynamic changes in the number of users is a key challenge. Summary of the Invention

[0003] The purpose of this invention is to provide a greedy DDPG collaborative offloading decision and resource allocation method, which aims to solve the technical problem in UDN-MEC scenarios where the limited computing power and energy resources of user terminals cannot meet the needs of users' computationally intensive and latency-sensitive tasks, and to eliminate the impact of the time-varying number of users in the micro base station coverage area on offloading decisions.

[0004] To achieve the above objectives, this invention provides a greedy DDPG collaborative unloading decision-making and resource allocation method, comprising the following steps:

[0005] Step 1: Deploy an ultra-dense network edge computing scenario including macro base stations, micro base stations, and users;

[0006] Step 2: Construct a communication model based on NOMA technology;

[0007] Step 3: Build a local computing model;

[0008] Step 4: Construct the MEC computing model;

[0009] Step 5: Build a cost model and optimization objectives;

[0010] Step 6: Estimate the local computing cost and offloading computing cost for each user task;

[0011] Step 7: Obtain the global reward using a greedy algorithm;

[0012] Step 8: Adjust the cooperation between agents using the DDPG algorithm to solve the unloading decision;

[0013] Step 9: Obtain resource allocation decisions based on unloading decisions.

[0014] Optionally, the micro base station is equipped with an MEC server. The MEC server allocates its computing resources according to the proportion of task data of all unloading users within the coverage area of ​​the micro base station, and sends the calculation results back to the unloading users through the micro base station after completing the unloading task.

[0015] Optionally, in step 2, users within the signal coverage area of ​​each micro base station are divided into a channel-multiplexed NOMA cluster. Based on NOMA technology and serial interference decoding characteristics, for user terminals in the same cluster, the one with a larger channel gain will be interfered with by the one with a smaller channel gain.

[0016] Optionally, in the local computing model, the time for computing tasks to be executed locally. Local computing energy consumption of user terminal x per unit time Latency of local computation for all user tasks in a micro base station NOMA cluster Energy consumption of all user tasks executed locally in a micro base station NOMA cluster The relational expression is shown below:

[0017]

[0018] Where, γ x This indicates the binary offload policy of user terminal x in the current time slot, D. x S represents the amount of task data for user terminal x in the current time slot. x f represents the number of CPU cycles required for user terminal x to perform local computation on each unit of task data. x K represents the local computing power of user terminal x, κ represents the effective switched capacitor in the chip, and K bsThis represents the total number of users offloaded within the NOMA cluster of the bth-th microbase station where user x is located.

[0019] Optionally, during the execution of step 4, it is first assumed that the offloading transfer between user tasks performing offloading computation is parallel, and the computing resources χ allocated by the MEC server for offloading user x are utilized. x MEC server computation latency for user task x User terminal u's binary offloading strategy γ in the current time slot u The amount of task data D of user terminal u in the current time slot u The computing power f of the MEC server deployed in the bth micro base station bs Offloading transmission delay of user task x The transmit power p of user terminal x x Total computational latency of MEC servers And the total transmission energy consumption of all offloaded users within the coverage area of ​​the micro base station. The expression is as follows:

[0020]

[0021] in, For the offloading transmission delay of user task x, This refers to the system's unloading and transmission delay.

[0022] Optionally, the optimization objective in step 5 is to minimize the total cost of the system latency and energy consumption weighted sum in the ultra-dense network edge computing scenario by optimizing user offloading decisions within each micro base station and the allocation of computing resources to the MEC server.

[0023] Optionally, the execution process of step 6 includes the following steps:

[0024] Step 6.1: Assume the coverage area of ​​the micro base station is K. bs Each user selects to perform unloading calculations. The following formula can be used to estimate the user task transmission speed.

[0025]

[0026] Among them, the estimated signal-to-interference-plus-noise ratio (SIR) for the x-th user within the micro base station coverage area is: B represents bandwidth;

[0027] Step 6.2: Calculate the user resource pre-allocation ratio Estimate the local computing cost of user tasks and uninstallation costs The expression is as follows:

[0028]

[0029] Where β is the weighting coefficient in cost calculation.

[0030] Optionally, in step 7, the user's individual greedy reward can be obtained using a greedy algorithm. The global revenue set of micro base station user clusters The expression is as follows:

[0031]

[0032] Optionally, during the execution of step 8, the global revenue threshold bias of the micro base station bs is intelligently adjusted using the DDPG algorithm. bs This enhances the ability to adapt to dynamic changes in the number of users and strengthens the collaborative behavior of users when making distributed uninstallation decisions.

[0033] Optionally, the DDPG algorithm's exploration strategy balances exploration behavior in stages and utilizes this behavior to improve the algorithm's convergence performance. By adjusting the time parameters ρ1 and ρ2 of the agent to strengthen exploration behavior and suppress exploitation behavior, the agent can strengthen exploration behavior in the early stage of training, strengthen exploitation behavior in the middle stage of training, and suppress exploration behavior in the later stage of training.

[0034] This invention provides a greedy DDPG collaborative offloading decision-making and resource allocation method. Addressing the difficulty in adapting to frequent user handovers within and outside the micro base station coverage area during computational offloading in UDN-MEC scenarios, this method utilizes a greedy algorithm to distribute and obtain individual user benefits. A global benefit threshold is used to describe the collaborative behavior among individual users based on global awareness. This overcomes the impact of time-varying user numbers caused by frequent handovers between micro base stations on offloading decisions, as well as the lack of cooperation between different agents in distributed single-agent methods. Furthermore, the DDPG algorithm intelligently adjusts the global benefit threshold, improving the algorithm's adaptability to dynamic changes in user numbers and reducing the latency- and energy-weighted total cost in dynamically heterogeneous UDN-MEC systems. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart illustrating the steps of a greedy DDPG collaborative unloading decision and resource allocation method according to the present invention.

[0037] Figure 2 This is a schematic diagram of the UDN-MEC system model in a specific embodiment of the present invention.

[0038] Figure 3 This is the observation state of the present invention. A schematic diagram of the data distribution.

[0039] Figure 4 This is a schematic diagram of the DDPG algorithm training framework construction of the present invention.

[0040] Figure 5 This is a schematic diagram illustrating the interaction between the DDPG algorithm of this invention and the environment.

[0041] Figure 6 This is a schematic diagram illustrating the convergence of a specific embodiment of the present invention during the training process.

[0042] Figure 7 This is a schematic diagram illustrating the impact of each method on the total cost of the proposed model system in a specific embodiment of the present invention.

[0043] Figure 8 This is a schematic diagram illustrating the impact of each method on the average system cost in a specific embodiment of the present invention.

[0044] Figure 9 This is a schematic diagram illustrating the impact of the MEC server's computing power resources on the average system cost in various embodiments of the present invention. Detailed Implementation

[0045] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0046] Please see Figure 1 This invention provides a greedy DDPG collaborative unloading decision and resource allocation method, comprising the following steps:

[0047] S1: Construct an ultra-dense network edge computing scenario that includes macro base stations, micro base stations, and users;

[0048] S2: Construct a communication model based on NOMA technology;

[0049] S3: Construct a local computing model;

[0050] S4: Construct the MEC computing model;

[0051] S5: Construct a cost model and optimization objectives;

[0052] S6: Estimate the local computing cost and offloading computing cost for each user task;

[0053] S7: Obtain global benefit through a greedy algorithm;

[0054] S8: Adjust the cooperation between agents through the DDPG algorithm, and then solve the unloading decision;

[0055] S9: Obtain resource allocation decisions based on unloading decisions.

[0056] The following provides further explanation in conjunction with the specific implementation steps:

[0057] The ultra-dense network edge computing scenario in S1 mainly consists of a macro base station (MBS), several micro base stations (SBS) deployed with MEC servers, and several users. The model structure is as follows: Figure 2 As shown in the diagram, the macro base station centrally controls the MEC offloading process and is connected to the micro base station via optical fiber. The macro base station does not provide computing resources to user terminals. The MEC server provides computing resources to user terminals within the micro base station's coverage area and is also connected to the micro base station via optical fiber. Assuming the system time is divided into several time slots, each user terminal generates an independent and indivisible data-intensive or computationally intensive task in each time slot. Considering the limited energy resources and computing power of user terminal devices, user tasks are offloaded to the MEC server deployed on their associated micro base station. The MEC server allocates its computing resources based on the proportion of task data volume of all offloading users within the micro base station's coverage area. After completing the offloading task, the MEC server sends the calculation results back to the offloading user via the micro base station. At the beginning of each time slot, each micro base station collects information such as the location and task data size of users within its signal coverage area and uses intelligent algorithms deployed on the micro base station to solve for each user's offloading decision and computational cost. Finally, the micro base station's computational cost is forwarded to the macro base station, which collects the computational costs of all micro base stations to construct the total cost of the UDN-MEC system.

[0058] S2, construct the communication model;

[0059] A communication model based on NOMA technology is constructed. Users within the signal coverage area of ​​each micro base station are divided into a channel-multiplexed NOMA cluster. Based on NOMA technology and the characteristics of serial interference decoding, for user terminals within the same cluster, the one with higher channel gain will experience interference from the one with lower channel gain. Assume that the signal coverage area of ​​the micro base station where user x is located has K... bs Each unloaded user, the user's channel gain is based on... Arranged in descending order, the average reference channel power gain α, additional attenuation η, small-scale Rayleigh fading ζ, and path loss exponent η are used. β The distance d between the user and the micro base stationx The transmit power p of user terminal u u Channel gain h of user terminal u u Additive white Gaussian noise σ 2 Intra-cluster interference I NOMA,x Bandwidth B and transmit power p of user terminal x x The channel gain h of user terminal x within this cluster can be expressed using the following formula. x Intra-cluster interference I NOMA,x and data rate R x ;

[0060]

[0061] S3, builds a local computing model;

[0062] S31, Assuming that user tasks are independent of each other, define γ x This indicates the binary offload policy of user terminal x in the current time slot, D. x S represents the amount of task data (in bits) of user terminal x in the current time slot. x K represents the number of CPU cycles required for local computation of each unit of task data in user terminal x, and κ represents the effective switched capacitor in the chip. bs This represents the total number of users offloaded within the NOMA cluster of the b-th microbase station where user x resides. x The local computing power of user terminal x is expressed in CPU cycles per second. The following formula represents the execution time of this computing task locally. Local computing energy consumption of user terminal x per unit time Latency of local computation for all user tasks in a micro base station NOMA cluster Energy consumption of all user tasks executed locally in a micro base station NOMA cluster

[0063]

[0064] S4, Construct the MEC computing model;

[0065] S41, assuming the offloading and transfer between user tasks performing offloading calculations are parallel, the offloading and transfer delay of user task x can be expressed by the following formula. and system offload transmission latency

[0066]

[0067] S42, In order to improve the utilization of computing resources, the computing power f of the MEC server deployed by the bth micro base station is utilized. bsBinary offloading strategy γ of user terminal u in the current time slot u The amount of task data D of user terminal u in the current time slot u Offloading transmission delay of user task x The transmit power p of user terminal x x The following formula can be used to represent the computing resources χ allocated by the MEC server to the unloading user x. x MEC server computation latency for user task x Total computational latency of MEC servers Total transmission power consumption of all offloaded users within the coverage area of ​​the micro base station

[0068]

[0069] S5, Build the cost model and optimization objectives;

[0070] S51, in order to characterize the system performance, the total delay T of all user terminals in the coverage area of ​​the micro base station bs is used to perform tasks. bs Total energy consumption E bs The weighted sum is defined as the cost function C. bs Let β represent T bs With E bs The weights between them, such as T expressed by the following formula bs E bs C bs ;

[0071]

[0072] S52, the overall objective of this invention is to minimize the total cost of the proposed UDN-MEC system by optimizing user offloading decisions and MEC server computing resource allocation within each micro base station, as expressed by the following formula;

[0073]

[0074]

[0075]

[0076] C3:γ x ∈{0,1}

[0077]

[0078] C5:χ x ∈[0,1)

[0079]

[0080] C7:f min ≤f x ≤f max

[0081] C8:D min ≤D x ≤D max

[0082] In the objective function, K bs C1 represents the number of users within the coverage area of ​​the micro base station bs, determined by the dynamic movement of users; C2 represents the number of users wirelessly covered by the micro base station, which is not greater than the total number of users N in the proposed system, where N... B C1 represents the total number of micro base stations; C2 represents the set of offloading decisions for micro base stations bs. bs C3 represents the binary offload decision γ of user terminal x in micro base station bs. x C4 represents the resource allocation decision set χ of micro base stations bs. bs C5 represents the resource allocation decision χ for user terminal x. x The range of values ​​for f; C6 indicates that the total resources allocated to all user terminals do not exceed the total resources owned by the MEC server; C7 and C8 are used to constrain the computing power of user terminals and the data volume of each time slot generation task, respectively, where f min and f max D represents the minimum and maximum values ​​of the user terminal's computing power, respectively. min and D max These represent the minimum and maximum values ​​of the tasks generated by the user terminal in each time slot, respectively.

[0083] In practice, considering that the user clusters in the coverage areas of each micro base station in the proposed UDN-MEC system are independent of each other, the optimization objective P0 can be decomposed into minimizing the cost C of each micro base station coverage area using the following formula. bs ;

[0084]

[0085] stC1,C3,C5,C6,C7,C8

[0086] S6 estimates the latency and energy weighted cost of local computation and offload computation for each user task.

[0087] S61, assuming the micro base station covers an area of ​​K. bs All users choose to perform offload calculations. The estimated signal-to-interference-plus-noise ratio (SIR) for the x-th user within the micro base station coverage area is: The following formula can be used to estimate the user task transmission speed.

[0088]

[0089] S62, Calculate the user resource pre-allocation ratio The following formula can be used to estimate the local computing cost of a user task. and uninstallation costs

[0090]

[0091] Where β is the weighting coefficient in cost calculation.

[0092] S7 obtains the individual greedy reward for each user through a greedy algorithm. Global benefits of micro base station user clusters

[0093]

[0094] S8 intelligently adjusts the cooperative behavior among agents through the DDPG algorithm, thereby solving the unloading decision;

[0095] S81, within the coverage area of ​​the micro base station bs, comprehensively considers individual greedy gains. and the global profit threshold bias intelligently adjusted by the DDPG algorithm bs Through bias bs To represent collaboration among users, the final reward g for each user can be obtained using the following formula. x and unloading decision γ x ;

[0096]

[0097] S82, Construct the state space. In time slot t, the agent interacts with the environment to obtain the observed state. It contains the global revenue set of user tasks within the current micro base station coverage area. In different time slots, The number of elements changes dynamically with the number of users, therefore This is not suitable as direct input data for the neural network in the DDPG algorithm. To construct suitable observation states s... t Using the neural network input data for the DDPG algorithm, it is necessary to extract data from different time slots. Their common characteristics. Therefore, firstly, regarding... Sort the data in descending order. Then, extract... Geometric features construct the observation state s of the DDPG algorithm t In order to describe The distribution of the data, respectively for Figure 3The horizontal and vertical axis scales are divided by 0.25 and 0.75 respectively to sample revenue data, and the sampled data is used to jointly describe... The geometric characteristics. Simultaneously, to normalize the magnitude information describing the returns, the ratio of average return to maximum return is introduced. and the mapping value of average return in This represents the maximum value in the current list of micro base station revenues; This represents the average value in the current list of micro base station revenues; Take a constant value for normalization. Finally, the observed state s of the DDPG algorithm can be obtained using the following formula. t ;

[0098]

[0099] in, Indicates the 0.25Kth bs Revenue per user terminal Indicates the 0.75Kth point bs Revenue per user terminal Indicates that the return value exceeds The percentage of user terminals, Indicates that the return value exceeds The percentage of user terminals.

[0100] S83, Construct the action space. In this model, the action space is defined by the reward threshold. The DDPG actor network is composed of s t Select action a in the action space. t implement;

[0101] S84, Construct an exploration strategy. A segmented exploration strategy is adopted to balance the agent's exploration and exploitation behaviors. This is achieved by the agent selecting the probability parameter o1 for exploration behavior, the time parameter ρ1 for reinforcing exploration behavior, the time parameter ρ2 for reinforcing exploitation behavior, and the random probability p. r Given ∈(0,1) and the constructed exploration factor b, the exploration strategy of the DDPG algorithm is obtained using the following formula. This strategy allows the agent to strengthen exploration behavior in the early stage of training, strengthen exploitation behavior in the middle stage of training, and suppress exploration behavior in the later stage of training;

[0102]

[0103] S85, Construct the reward function. To highlight the immediate reward r's feedback to the agent's learning policy, the parameter... Divide r into immediate reward or immediate penalty, both represented by r. tThis indicates that, to enhance the convergence performance of the algorithm, an exponential function is used to enhance different immediate rewards r. t The difference between them, when r t When used as an immediate penalty, the penalty intensity is defined to decrease as time increases. The reward for time slot t is then obtained using the following formula;

[0104]

[0105] Where t it It is the number of iterations. It represents the maximum number of iterations.

[0106] S86, Construct the DDPG algorithm training framework. See the framework construction diagram below. Figure 4 Unlike the traditional DRL algorithm, the DDPG algorithm utilizes an Actor-Critic framework. It enhances the algorithm's exploration capabilities by adding behavioral noise to the agent and improves the stability of the training process through an experience replay mechanism. DDPG internally consists of four neural networks: an actor network π and a critic network Q, as well as a target actor network π' and a target critic network Q'. The parameters of the four neural networks are θ. π θ Q θ π' θ Q' The actor and critic networks output the action 'a' and value 'Q', respectively. The target actor and targetcritic networks have the same network structure as the actor and critic networks, outputting the expected action 'a' and value 'Q(s',a') for the next state 's'. During the training of the DDPG algorithm, the actor network first reads the current state 's' from the UDN-MEC environment, and then uses action 'a' as the reward threshold bias. bs The output is sent to the UDN-MEC environment to obtain the immediate reward r and the new state s'. DDPG stores the state s, action a, immediate reward r and new state s' output by the UDN-MEC environment as an experience in the experience pool. Finally, the network parameters of DDPG are trained and updated through experience replay technology.

[0107] S87, using the following formula, the time difference error is calculated based on the batch sample size τ. Solving for the loss function L via backpropagation and optimizing gradient updates. Obtain the parameters θ of the critic network Q ;

[0108]

[0109] Where s iLet γ represent the current state of sample i, and let γ represent the discount factor.

[0110] The input to S88,π is the observation state s of the current time slot. t The output of π is the agent's action a under the observed state. t If the following formula is used, the parameter θ of π is updated according to the determined policy gradient. Q Utilizing θ through soft update strategies π and θ Q Update θ π ' and θ Q ';

[0111]

[0112] Where Γ represents the soft-interval update coefficient.

[0113] After the S89 and DDPG algorithms are trained, save θ. π When applying the method of this invention, the saved θ is used. π Replace the actor neural network parameters in the current algorithm, and then use the actor network of the present invention to obtain the profit threshold bias. bs The interaction with the environment is illustrated as follows: Figure 5 As shown;

[0114] S9, Solve resource allocation decisions based on unloading decisions;

[0115]

[0116] In the greedy DDPG collaborative offloading decision and resource allocation algorithm for ultra-dense network edge computing scenarios of this invention:

[0117] The pre-allocation ratio of user resources is calculated based on the proportion of user data volume. Using a greedy algorithm to predict individual user revenue and global benefits To enhance user collaboration based on global awareness, the global benefit threshold bias is intelligently adjusted using the DDPG algorithm. bs The uninstallation decision g for each user is solved by combining individual and global benefit thresholds. x This reduces the latency and energy-weighted total cost of the UDN-MEC system;

[0118] Furthermore, to verify the superiority of the proposed GOA-RA-DDPG unloading strategy, this invention proposes a specific embodiment, employing the following four strategies for performance comparison with the proposed GOA-RA-DDPG strategy. Specifically, this invention is verified in a Python 3.9 and PyTorch 1.13 environment, comparing the GOA-RA-DDPG algorithm of this invention with the full unloading computation strategy, the random unloading computation strategy, the all-local computation strategy, and the greedy unloading computation strategy.

[0119] 1. Full Unloading Computation Strategy

[0120] All users in the micro base station NOMA cluster have chosen to offload their computational tasks.

[0121] 2. Random Unloading Computation Strategy

[0122] All users in a micro base station NOMA cluster randomly choose to perform their tasks locally or offload the computation.

[0123] 3. All-local computing strategy

[0124] All users in the micro base station NOMA cluster choose to perform their tasks locally.

[0125] 4. Greedy unloading computation strategy

[0126] All users in a micro base station NOMA cluster choose to perform their tasks locally or offload the computation based on cost estimation and the principle of minimizing their own costs.

[0127] Based on actual design requirements, the algorithm parameters are set as follows: the actor network uses a 4-layer fully connected neural network with 18 hidden neurons; the critic network uses a 3-layer fully connected neural network with 21 hidden neurons; users within the microcell use a random walk model for non-linear movement, with the number of users initialized to 30 and the user movement speed initialized to 5 m / s.

[0128] Figure 6 The convergence of the proposed GOA-RA-DDPG algorithm during training was verified. The algorithm reached convergence in approximately 2000 steps. This is attributed to the proposed exploration strategy, which provides suitable training samples to the DDPG algorithm in stages, effectively balancing the algorithm's demands on exploration and exploitation behaviors during training. Furthermore, the proposed reward strategy enhances the discriminative power among advantageous samples, thus promoting algorithm convergence.

[0129] Figure 7This paper describes the impact of the proposed GOA-RA-DDPG strategy, along with the greedy strategy (GOA-RA), the all-offload strategy, the all-local strategy, and the random strategy, on the total system cost of the proposed model. As shown in the figure, the proposed GOA-RA-DDPG strategy outperforms other benchmark strategies in different time slots, indicating its superiority in reducing the cost of micro base station coverage areas. This is because the proposed strategy can predict the offload and local benefits for the current user, and the actor network of the DDPG algorithm can adjust for the benefit threshold bias. bs Intelligent adjustments are made, thereby reducing system costs.

[0130] Figure 8 The impact of the proposed algorithm on the average system cost was compared with other algorithms under different initial user numbers. The results show that the proposed GOA-RA-DDPG algorithm outperforms other algorithms under different initial user numbers, indicating that the proposed algorithm has strong robustness and can adapt to the dynamic changes in the number of users caused by frequent user handovers between micro base stations.

[0131] Figure 9 The impact of MEC server computing power on the average system cost is described. As shown in the figure, the system cost gradually decreases with increasing MEC server computing power. This is because as MEC server computing power increases, it can support more users performing offloaded computations. Furthermore, with increasing MEC server computing power, the advantage of the proposed GOA-RA-DDPGA strategy over other benchmark strategies further diminishes. This is because the number of users performing local computations in the user cluster gradually decreases, and the impact of offload strategies on system cost gradually converges.

[0132] The above description discloses only one preferred embodiment of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A method for greedy DDPG collaborative offloading decision and resource allocation, characterized in that, Includes the following steps: Step 1: Deploy an ultra-dense network edge computing scenario including macro base stations, micro base stations, and users; Step 2: Construct a communication model based on NOMA technology; Step 3: Build a local computing model; Step 4: Construct the MEC computing model; Step 5: Build a cost model and optimization objectives; The optimization objective in step 5 is to minimize the total cost of the system latency and energy consumption weighted sum in the ultra-dense network edge computing scenario by optimizing user offloading decisions within each micro base station and the allocation of computing resources for the MEC server. Step 6: Estimate the local computing cost and offloading computing cost for each user task; The execution process of step 6 includes the following steps: Step 6.1 : Assuming the micro base station coverage area Each user chooses to offload computation, as estimated user task transfer speed using the following formula : wherein the estimated signal-to-interference-and-noise ratio of the user in the micro base station coverage area is , , denotes the bandwidth; Step 6.2: Calculate the user resource pre-allocation ratio Estimate the local computing cost of user tasks and uninstallation costs The expression is as follows: in, These are weighting coefficients used in cost calculation; Step 7: Obtain the global reward using a greedy algorithm; In step 7, the individual greedy reward for each user is obtained using a greedy algorithm. The global revenue set of micro base station user clusters The expression is as follows: ; Step 8: Adjust the cooperation between agents using the DDPG algorithm to solve the unloading decision; During the execution of step 8, the micro base station is intelligently adjusted using the DDPG algorithm. Global revenue threshold This enhances the ability to adapt to dynamic changes in the number of users and strengthens the collaborative behavior of users when making distributed uninstallation decisions. Step 9: Obtain resource allocation decisions based on unloading decisions.

2. The greedy DDPG collaborative unloading decision and resource allocation method as described in claim 1, characterized in that, The micro base station is equipped with an MEC server. The MEC server allocates its computing resources according to the proportion of task data of all unloading users in the coverage area of ​​the micro base station. After completing the unloading task, the calculation results are sent back to the unloading user through the micro base station.

3. The greedy DDPG collaborative unloading decision and resource allocation method as described in claim 2, characterized in that, In step 2, users within the signal coverage area of ​​each micro base station are divided into a channel-multiplexed NOMA cluster. Based on NOMA technology and serial interference decoding characteristics, for user terminals in the same cluster, the one with a larger channel gain will be interfered with by the one with a smaller channel gain.

4. The greedy DDPG collaborative unloading decision and resource allocation method as described in claim 3, characterized in that, In the local computing model, the computing task is executed locally. User terminal Locally calculated energy consumption per unit time The latency of local computation for all user tasks in a micro base station NOMA cluster Energy consumption of all user tasks executed locally in a micro base station NOMA cluster The relational expression is shown below: in, Indicates user terminal The binary offloading strategy in the current time slot, Indicates user terminal The amount of task data in the current time slot, Indicates user terminal Number of CPU cycles required to perform local computation on each unit of task data. Indicates user terminal Local computing power Indicates the effective switched capacitor in the chip. Indicates user The first The total number of users offloaded within the NOMA cluster of each micro base station.

5. The greedy DDPG collaborative unloading decision and resource allocation method as described in claim 4, characterized in that, During the execution of step 4, it is first assumed that the offloading transfer between user tasks performing offloading calculations is parallel, and the MEC server is used for offloading users. Allocated computing resources User tasks MEC server computation latency Total computational latency of MEC servers and the total transmission energy consumption of all offloaded users within the coverage area of ​​the micro base station. The expression is as follows: in, For user tasks The offloading transmission delay For the system's offload transmission delay, Indicates user terminal The binary offloading strategy in the current time slot, Indicates user terminal The amount of task data in the current time slot, Indicates the first The computing power of the MEC server deployed in each micro base station Indicates user task The offloading transmission delay Indicates user terminal The transmission power.

6. The greedy DDPG collaborative unloading decision and resource allocation method as described in claim 5, characterized in that, The DDPG algorithm's exploration strategy balances exploration behavior in stages and leverages this behavior to improve convergence performance; it also adjusts the time parameters by enhancing the agent's exploration behavior and suppressing the exploitation of behavior. and time parameters This allows the agent to strengthen exploratory behavior in the early stages of training, strengthen exploitation behavior in the middle stages of training, and suppress exploratory behavior in the later stages of training.