A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method

Optimizing the offload decision of the NOMA-MEC system through the SAC agent algorithm, solving the network performance degradation and high energy consumption problems caused by the generalized user packet transmission method, and achieving a significant reduction in system energy consumption and complexity.

CN116193546BActive Publication Date: 2025-09-02XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211440680.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-09-02
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

The existing NOMA-MEC system does not adopt a generalized user packet transmission method, resulting in a degradation of network performance. The existing optimization solutions cannot meet the real-time decision-making needs and high complexity, and cannot effectively reduce system energy consumption.

Method used

The SAC agent algorithm is used to optimize the offload decision decision. Through the deep neural network architecture of Actor and two Critics, combining discount factor, soft copy factor and maximum time slot number, the offload ratio and transmission power are optimized to reduce system energy consumption.

Benefits of technology

Significantly reduce system energy consumption, reduce complexity, meet real-time decision-making needs, and improve network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116193546B_ABST
    Figure CN116193546B_ABST
Patent Text Reader

Abstract

The present invention relates to a dynamic generalized user (NOMA) grouping CCHN-MEC network offloading decision optimization method. First, under a given computational offloading ratio, the optimal solution to the local energy consumption minimization problem is derived, namely, the optimal local CPU frequency allocation for the secondary user (SU). #imgabs0# Next, a convex optimization tool is used to solve the offloading energy minimization problem, obtaining the CPU frequency #imgabs2# allocated for each SU task calculation within each NOMA group #imgabs1#, the SU's transmit power and transmission time #imgabs3#, and finally, a deep reinforcement learning algorithm based on SAC is used to learn the user computational offloading ratio distribution for each time slot to obtain the optimal offloading decision. The present invention can significantly save system energy consumption and has low complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of CCHN-MEC network resource allocation, and in particular to a dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method. Background Art

[0002] The Internet of Things (IoT) has become an integral part of our daily lives, giving rise to a variety of computationally intensive and latency-sensitive applications, such as facial recognition and natural language processing. To meet the latency-sensitive computing needs of low-complexity devices, the academic community has proposed Mobile Edge Computing (MEC). Furthermore, due to limited wireless spectrum, researchers are developing new technologies to support MEC, such as Non-Orthogonal Multiple Access (NOMA) and Cognitive Radio (CR). Consequently, the combination of MEC and CR, as well as the combination of MEC and NOMA, has become a hot research topic.

[0003] Recently, scholars have proposed a new CR infrastructure, the Cognitive Capacity Harvesting Network (CCHN), to allow handheld lightweight devices without management / awareness capabilities to enjoy the benefits of the CR network (CRN). In the CCHN, a Secondary Server Provider (SSP) is introduced, which deploys a group of CR routers to monitor / detect the CR spectrum and purchases a small piece of licensed spectrum to build a reliable public control channel. Through the public control channel, the SSP collects the management / awareness results of the CR routers, guides the CR routers to form a CRN, and allocates CR spectrum. Secondary Users (SUs) without management / awareness capabilities can access nearby CR routers through the allocated CR spectrum. In effect, the CCHN introduces a new network operator, which is responsible for building the infrastructure, acquiring spectrum from the primary network operator that owns the spectrum, and providing services to SUs without management / awareness capabilities.

[0004] Assuming that CR routers are already equipped with computing resources, CCHN can provide MEC services. In CCHN, to reduce interference to the primary network, considering the limited transmission power of SUs, and to meet the delay constraints of SU tasks, it is necessary to introduce "generalized user grouping" to improve network performance. That is, a SU is allowed to join multiple NOMA groups and offload different parts of its data to different CR routers through different transmission channels.

[0005] However, existing research on NOMA-MEC does not adopt a generalized user grouping transmission method, which greatly reduces network performance. Furthermore, due to the inherently high complexity of traditional optimization schemes, they cannot meet the real-time decision-making requirements of MEC systems. Therefore, in CCHN-MEC systems based on generalized user grouping NOMA, it is very important to design an optimization scheme that supports dynamic computation offloading decisions. Currently, the dynamic offloading decision optimization schemes that can be used in CCHN-MEC systems based on generalized user grouping NOMA mainly include the TD3 algorithm, pure local computation (LO), and random allocation (RA) algorithm. Because the TD3 algorithm aims to find a deterministic strategy, it is difficult to apply to the randomly changing NOMA-CCHN-MEC scenario, and thus energy consumption may be greatly increased. The LO algorithm does not consider the full utilization of the computing resources of the MEC server, and when latency requirements are very high, energy consumption increases significantly. The RA algorithm does not consider the coupling between system variables and has poor performance. Summary of the Invention

[0006] In response to the problems existing in the prior art, the purpose of the present invention is to provide a dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method, which can obtain more reasonable offloading decisions and significantly reduce system energy consumption.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method is proposed. The CCHN consists of a group of SUs, a group of CR routers, and an SSP. The SSP centrally manages the SUs and CR routers through an established common control channel. The CR routers are equipped with computing resources and act as MEC servers. The uplink cellular spectrum of adjacent cells is divided into a series of cellular resource blocks (CRBs). The CRBs are allocated, and each CR router has been assigned a CRB for data offloading from its connected SUs.

[0009] The decision optimization method adopts SAC algorithm to solve the problem. SAC agent includes an Actor and two Critic. Actor is a fully connected DNN with several layers, denoted as in Represents the weight parameters of DNN; by observing the input state The mean of the Actor output policy distribution and standard deviation Since the strategy distribution is fitted into a Gaussian distribution, Sampling can get actionable actions Each critic contains two fully connected DNNs with the same network architecture, namely the main DNN and the target DNN; each DNN is used to evaluate and The Q value is where θ is the weight of the DNN; using and To represent the Q value of the main DNN and target DNN of Critic 1, the weights are θ1 and use and To represent the Q value of the main and target DNN of Critic2, the weights are θ2 and

[0010] The decision optimization method is specifically as follows:

[0011] Step 1: Set the discount factor λ, the soft copy factor ι, the maximum number of time slots T, and the maximum number of rounds Γ;

[0012] Step 2: Randomly initialize the Actor's neural network parameters Critic's main neural network parameter θ i (i=1,2), initialize the target neural network parameters of Critic Clear the replay experience pool. The current time slot number is t=1, and the current round number is e=1;

[0013] Step 3: Randomly generate an offload ratio action to obtain the state in a CCHN-MEC environment

[0014] Step 4: According to the status The SAC agent outputs a decision action to calculate the unloading ratio

[0015] Step 5: Decision based on uninstall ratio For each SU m l , solve the local energy consumption minimization problem and obtain the local optimal CPU frequency allocation

[0016] Step 6: Decision based on uninstall ratio For each NOMA group get CPU frequency of each SU

[0017] Step 7: Decision based on uninstall ratio For each NOMA group Solve the unloading energy consumption optimization problem and obtain the group The transmission power and transmission time of each SU

[0018] Step 8. Calculate the total system energy consumption E based on steps 5, 6, and 7. total (t);

[0019] Step 9. According to step 8, calculate the current system reward r t , and get the state of the next state

[0020] Step 10: The experience of the current time slot Store in replay experience pool middle;

[0021] Step 11: If you are replaying the experience pool The number of experiences in is greater than the minimum batch size From the replay experience pool Randomly extract data Perform network training to update network parameters θ i (i=1,2), and the temperature coefficient ∈;

[0022] Step 12, t=t+1; if the current time slot number t>T, then t=1, e=e+1, if e>Γ, proceed to step 13; otherwise, return to step 3;

[0023] Step 13: Output the optimal parameters of the Actor's neural network Through this parameter, the Actor can output the optimal decision action in each state.

[0024] In step 5, the local energy consumption minimization problem is:

[0025]

[0026] in, Indicates the maximum CPU frequency of the SU; constraint C1 indicates that the local computing time cannot exceed the length of a time slot, and constraint C2 indicates that the local computing CPU frequency of the SU cannot exceed the maximum CPU frequency.

[0027] In step 6, the following formula (7) is used to obtain CPU frequency of each SU

[0028]

[0029] In step 7, the unloading energy consumption optimization problem is:

[0030]

[0031] in, represents the channel gain from transmitter u1 to receiver u2 in time slot t, Indicates that in time slot t, the transmitter m l To the receiver h k The normalized channel gain, P k INT (t) represents the CR router h in time slot t k Maximum interference and noise power levels at Indicates CR router h k The maximum CPU frequency, SU m l The total number of NOMA groups that join in time slot t, SU m l The maximum transmit power, Indicates that k The maximum tolerated interference at the associated qth BS, N k Indicates that k The number of all associated BSs;

[0032] Constraint C1 indicates that the transmission time of the NOMA group cannot exceed the length of each time slot. Constraint C2 indicates that the computing CPU frequency allocated to the SUs in the NOMA group cannot exceed the maximum CPU frequency of the CR router. Constraint C3 indicates the rate requirement for each SU to offload data. Constraint C4 indicates the limit on the transmission power of each SU. Constraint C5 indicates that the interference power of the SUs in the NOMA group cannot exceed the maximum tolerable interference value of each BS.

[0033] The total energy consumption of the system in step 8 is E total (t) is:

[0034]

[0035] in, SU m l The energy consumption weight, Indicates that the CR router is calculating SU m l The energy weight consumed by the task, SU m l Local energy consumption, κ0 represents the computational energy consumption factor of SU, β l,k (t) represents the time slot tSU m l Offload to CR router c k The task offloading ratio, w l (t) represents the time slot tSUm lThe number of CPU cycles required to calculate the 1-nat task data, R l (t) represents the time slot tSUm l The total amount of task data, f l loc (t) represents the time slot tSUm l Local CPU frequency; SU m l Offload to CR router c k Energy consumption per hour, p l,k (t) represents the time slot tSUm l Offload to CR router c k The transmission power at time d k (t) represents the unloading to CR router h at time slot t k The NOMA transmission time of SU, SU m l Offload to CR router c k When it reaches CR router c k The computational energy consumption of CR router is represented by κ1, and f l,k (t) In time slot tSUm l Offload to CR router h k CR router h k Assigned CPU frequency.

[0036] In step 9, the current system reward r is calculated using formula (10) t :

[0037]

[0038] Among them, and They respectively indicate whether the constraints (1), (2), and (4) in time slot t are violated and whether problems P1 and P2 have solutions in time slot t. If constraint (1) is satisfied, then otherwise, If constraint (2) is satisfied, then otherwise If constraint (4) is satisfied, then otherwise, If problem P1 has a solution, then otherwise, If problem P1 has a solution, then otherwise, If there is no solution to problem P1 / P2 in time slot t, then E total (t) is set to +∞.

[0039] In step 11, the network parameters θ are updated according to formulas (12)-(15). i (i=1,2), And the temperature coefficient ∈:

[0040] The minimum Bellman residual value is calculated as follows:

[0041]

[0042] in, here, It is from Resampled actions;

[0043] Actor's DNN is trained by minimizing the KL divergence, i.e. minimizing

[0044]

[0045] in is a parameter The expression of ε t is the input Gaussian noise vector;

[0046] The temperature parameter ∈ is dynamically adjusted by minimizing the following formula

[0047]

[0048] in is the target entropy constant;

[0049] The weight parameters of the target DNN in each critic are updated by the soft copy method, that is,

[0050]

[0051] After adopting the above scheme, the present invention first derives the optimal solution to the local energy consumption minimization problem under a given computation offloading ratio, that is, the optimal CPU frequency allocation of the secondary user (SU) locally. Secondly, the convex optimization tool is used to solve the problem of minimizing the unloading energy consumption to obtain the Calculate the CPU frequency allocated to each SU task SU's transmission power and transmission time Finally, a deep reinforcement learning algorithm based on SAC is used to learn the user calculation offloading ratio distribution in each time slot to obtain the optimal offloading decision. The present invention can significantly save system energy consumption and has low complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1The energy consumption comparison between the present invention (SAC) and the TD3 algorithm, LO algorithm and random allocation RA algorithm under the change of SU number is shown;

[0053] Figure 2 The energy consumption comparison between the present invention (SAC) and the TD3 algorithm, LO algorithm and RA algorithm under the change of the total amount of unloaded data per SU;

[0054] Figure 3 The figure compares the energy consumption of the present invention (SAC) with that of the TD3 algorithm, the LO algorithm, and the RA algorithm under the change of the length of each time slot. DETAILED DESCRIPTION

[0055] The present invention discloses a dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method, in which the CCHN under consideration includes a group of SUs, a group of CR routers and an SSP. The SSP centrally manages the SUs and CR routers through an established common control channel. The CR router is equipped with computing resources and acts as a MEC server. Consider that the CCHN provides services to the SUs by sharing the uplink cellular spectrum in adjacent cells. The uplink cellular spectrum of the adjacent cells is divided into a series of cellular resource blocks (CRBs). Due to the use of spatial reuse technology, each CRB can be used by cellular users (CUs) in several adjacent cells according to the spatial reuse factor. In order to simplify the analysis, it is assumed that the allocation of CRBs has been completed and each CR router has been allocated a CRB for data offloading of the SUs connected to it.

[0056] System time is slotted as τ = {1, 2, ..., T}, where τ is the length of each slot. In each slot, each SU has a computational task to complete. Assuming the SU is lightweight, it is necessary to offload some of its tasks to the CR router. Considering the use of NOMA technology, SUs that select the same CR router for data offload form a NOMA group. Considering generalized user grouping, each SU can join multiple NOMA groups and offload different parts of its tasks to multiple CR routers simultaneously, reducing interference with the primary network and lowering offloading latency.

[0057] To reduce the complexity of Successive Interference Cancellation (SIC) decoding, a system parameter is set to limit the maximum number of SUs connected to a CR router. In addition, as with many existing MEC works, it is assumed that the size of the task calculation results is small enough that the energy consumption involved in downloading the task calculation results from the CR router to the SU is negligible.

[0058] make and Denote the SU set and the available CRB set respectively. The allocation of CRBs has been completed, and each CRB has been allocated to a CR router. Therefore, the number of CR routers is equal to the number of CRBs. Without loss of generality, assume that CRBc k Assigned to CR router h k .

[0059] Let β l,k (t) represents SU m in time slot t l Offload to CR router h k The task data volume ratio. Since the total amount of offloaded data of each SU in each time slot cannot exceed the entire task, we have

[0060]

[0061] Let L denote the CR router h k The maximum number of SUs connected. Then, we have

[0062]

[0063] Where I(y) is a function of y, defined as

[0064]

[0065] In order to reduce the optimization complexity, we set a parameter to limit the number of CRBs that each SU is allowed to use. That is,

[0066]

[0067] Where Y is the maximum number of CRBs allowed to be used by each SU. Note that Y = 1 represents the case where there is no generalized user grouping.

[0068] The total energy consumption E in each time slot t in the system total can be calculated as

[0069]

[0070] in, SU m l The energy consumption weight, Indicates that the CR router is calculating SU m l The energy weight consumed by the task, SU m l Local energy consumption, κ0 represents the computational energy consumption factor of SU, β l,k (t) represents the time slot tSUm l Offload to CR router c kThe task offloading ratio, w l (t) represents the time slot tSUm l The number of CPU cycles required to calculate the 1-nat task data, R l (t) represents the time slot tSUm l The total amount of task data, f l loc (t) represents the time slot tSUm l Local CPU frequency, SU m l Offload to CR router c k Energy consumption per hour, p l,k (t) represents the time slot tSUm l Offload to CR router c k The transmission power at time d k (t) represents the unloading to CR router h at time slot t k The NOMA transmission time of SU, SU m l Offload to CR router c k When it reaches CR router c k The computational energy consumption of CR router is represented by κ1, and f l,k (t) In time slot tSUm l Offload to CR router h k CR router h k Assigned CPU frequency.

[0071] (1) When the offloading ratio of a given SU ​​is When , for time slot t, the local computation energy consumption minimization of all SUs is independent of each other. Let Indicates m l The total amount of task data calculated locally in time slot t. Then, SU m in time slot t l The problem of minimizing the local computing energy consumption can be expressed as

[0072]

[0073] in, Indicates the maximum CPU frequency of SU. Constraint C1 indicates that the local computing time cannot exceed the length of a time slot, and constraint C2 indicates that the local computing CPU frequency of SU cannot exceed the maximum CPU frequency. The decrease of is monotonically decreasing, so the optimal solution f of problem P1 is l loc_opt (t) is equal to exist In this case, problem P1 has no solution.

[0074] (2) When the offloading ratio of a given SU ​​is When , the CR router h can be obtained by the following formula k For each NOMA group Calculate the optimal CPU frequency assigned to all SU tasks within

[0075]

[0076] (3) When the offloading ratio of a given SU When , the existing convex optimization toolbox can be used to solve the following unloading energy consumption optimization problem P2, and obtain the energy consumption of each NOMA group The optimal power of all SUs in and time allocation k (t);

[0077]

[0078] in, represents the channel gain from transmitter u1 to receiver u2 in time slot t, Indicates that in time slot t, the transmitter m l To the receiver h k The normalized channel gain, P k INT (t) represents the CR router h in time slot t k Maximum interference and noise power levels at Indicates CR router h k The maximum CPU frequency, SU m l The total number of NOMA groups that join in time slot t, SU m l The maximum transmit power, Indicates that the same k The maximum tolerated interference at the associated qth BS, N k Indicates that k The number of all associated BSs. Constraint C1 indicates that the transmission time of the NOMA group cannot exceed the length of each time slot. Constraint C2 indicates that the computing CPU frequency allocated to the SUs in the NOMA group cannot exceed the maximum CPU frequency of the CR router. Constraint C3 indicates the rate requirement for each SU to offload data. Constraint C4 indicates the limit on the transmit power of each SU. Constraint C5 indicates that the interference power of the SUs in the NOMA group cannot exceed the maximum tolerable interference value of each BS.

[0079] Obviously, when the calculated uninstall ratio is given Under this condition, the system energy consumption of each time slot can be calculated using formula (5) according to (6)(7)(8). However, in a dynamic environment, the offloading ratio The optimization of is a key issue. To this end, the SAC DRL algorithm is used to solve this problem. First, the system state, action, and reward are defined as follows:

[0080] 1) State: In order to save the state space size, the state in time slot t, is defined as the reward r of the previous time slot t-1 ;

[0081] 2) Action: The action of the agent in time slot t is defined by the extended calculation offloading ratio of SU, where ν is the positive parameter of the expansion. and β l,k The relationship between (t) is defined as:

[0082]

[0083] 3) Reward: The instantaneous reward definition r of the DRL agent in time slot t t for

[0084]

[0085] Among them, and They represent whether the constraints (1), (2), and (4) in time slot t are violated and whether problems P1 and P2 have solutions in time slot t. If constraint (1) is satisfied, then otherwise, If constraint (2) is satisfied, then otherwise If constraint (4) is satisfied, then otherwise, If problem P1 has a solution, then otherwise, If problem P1 has a solution, then otherwise, If there is no solution to problem P1 / P2 in time slot t, then E total (t) is set to +∞. From Equation (10), we can see that if the computation offloading decision in the current time slot violates constraints (1) / (2) / (4), or makes the problem P1 / P2 unsolvable, the immediate reward will be smaller, which prompts the DRL agent to choose a more reasonable strategy in the next time slot. Its goal is to obtain an optimal strategy π * It can maximize the long-term expected reward while maximizing the action entropy of each state, that is,

[0086]

[0087] in, Represents the state under strategy π The action entropy of .

[0088] ρ π represents the trajectory distribution of action-state pairs caused by unloading policy π in the considered CCHN-MEC environment. λ∈(0,1) is a discount factor that reflects the importance of future rewards. ∈ is a temperature parameter that balances the importance of entropy on system rewards. By introducing entropy regularization, the actions selected by the SAC agent become more random, which gives the algorithm a strong ability to explore actions, thereby obtaining the optimal policy π with a higher probability. * .

[0089] Considering that the state and action space of the system are continuous, in order to obtain a near-optimal unloading ratio decision, the SAC algorithm is used to solve it. In order to handle high-dimensional continuous state and action space, the Actor and Critic of the SAC agent are approximated using a deep neural network (DNN). The Actor is a fully connected DNN with several layers, denoted as in Represents the weight parameters of DNN. By observing the input state Actor can output the mean of the policy distribution and standard deviation Since the strategy distribution is fitted into a Gaussian distribution, Sampling can get actionable actions Each critic contains two fully connected DNNs with the same network architecture, namely the main DNN and the target DNN. Each DNN is used to evaluate and The Q value is Where θ is the weight of DNN. To distinguish, use and To represent the Q value of the main and target DNN of Critic 1, the weights are θ1 and use and To represent the Q value of the main and target DNN of Critic 2, the weights are θ2 and

[0090] In order to train the main DNN of each Critic, it is necessary to minimize the Bellman residual, which can be calculated by the following formula

[0091]

[0092] in here, It is from Resampled action.

[0093] Actor's DNN is trained by minimizing the KL divergence, i.e. minimizing

[0094]

[0095] in is a parameter The expression of ε t is the input Gaussian noise vector.

[0096] The temperature parameter ∈ is dynamically adjusted by minimizing the following formula

[0097]

[0098] in, is the target entropy constant.

[0099] The weight parameters of the target DNN in each critic are updated by the soft copy method, that is,

[0100]

[0101] The method for optimizing the calculation of the unloading ratio of the present invention specifically comprises the following steps:

[0102] Step 1: Discount factor λ, soft copy factor ι, maximum number of time slots T, maximum round Γ;

[0103] Step 2: Randomly initialize the Actor's neural network parameters Critic's main neural network parameter θ i (i=1,2), initialize the target neural network parameters of Critic Will replay the experience pool Clear, that is The current time slot number is t=1, and the current round number is e=1;

[0104] Step 3: Randomly generate an action to obtain a state in a CCHN-MEC environment

[0105] Step 4: According to the status The SAC agent outputs a decision action to calculate the unloading ratio

[0106] Step 5: Decision based on uninstall ratio For each SU m l , solve problem P1 to obtain its local optimal CPU frequency allocation

[0107] Step 6: Decision based on uninstall ratio For each NOMA group Obtained by formula (7) CPU frequency of each SU

[0108] Step 7: Decision based on uninstall ratio For each NOMA group By solving P2, we can obtain the group The transmission power and transmission time of each SU

[0109] Step 8: Based on steps 5, 6 and 7, use formula (5) to calculate the total energy consumption E of the system. total (t);

[0110] Step 9: Based on step 8, use formula (10) to calculate the reward r of the current system t , and get the state of the next state

[0111] Step 10: The experience of the current time slot Store in replay experience pool middle;

[0112] Step 11: If you are replaying the experience pool The number of experiences in is greater than the minimum batch size From the replay experience pool Randomly extract data Perform network training and update network parameters θ using equations (12), (13), (14), and (15) i (i=1,2), temperature coefficient∈, and

[0113] Step 12, t=t+1; if the current time slot number t>T, then t=1, e=e+1, if e>Γ, proceed to step 13; otherwise, return to step 3;

[0114] Step 13: Output the optimal parameters of the Actor's neural network Through this parameter, the Actor can output the optimal decision action in each state.

[0115] In order to evaluate the performance of the present invention, the following simulation is performed, and the simulation parameters are set as follows: it contains a central cell, surrounded by six adjacent cells, and the frequency reuse factor is 7. The radius of each cell is set to 500m. The CCHN considered is located in the central cell and uses the cellular spectrum of the six adjacent cells for data unloading. In each adjacent cell, the BS is located in the center and the CUs are evenly distributed. In the central cell, the SU and CR routers are evenly distributed. Each CRB in the adjacent cell is randomly assigned an active CU. The transmission power and rate requirements of each CU are set to 23dBm and 200knats / s, respectively, which are used to calculate the maximum interference and noise power levels of each CR router and the maximum allowed interference power in each CRB. The hardware-related calculation energy constants κ0 and κ1 of the SU and CR routers are set to 10, respectively. -26 and 10 -28 The number of CPU cycles required to calculate the 1-nat offload data is randomly selected in the range [100,1500]. The energy consumption weight of each SU and The maximum number of state units (SUs) connected to a CR router is set to 10. The maximum number of CRBs allowed per SU is set to 5. In the SAC algorithm, the actor and critic DNNs for the master and target actors consist of an input layer, an output layer, and two hidden fully connected layers with 128 and 64 neurons, respectively. Other parameters are shown in Table 1.

[0116] parameter value parameter value Number of CRBs 5 Maximum CPU frequency of SU 1GHz Channel bandwidth per CRB 1MHz Maximum CPU frequency of CR routers 20GHz Maximum transmit power of SU 20dBm Soft copy coefficient 0.001 Noise power spectral density -174dBm / Hz Discount Factor 0.99 Path loss coefficient 4 Actor learning rate 0.0001 Channel correlation factor 0.6 Replay experience pool size <![CDATA[10 6 ]]> Number of time slots per round 100 Maximum number of rounds 1000

[0117] Table 1

[0118] Figures 1 to 3 The energy consumption of the proposed SAC algorithm is compared with that of the TD3 algorithm, the LO algorithm and the RA algorithm under the changes of the number of SUs, the total amount of unloaded data per SU and the length of each time slot. In the LO algorithm, the entire computational task of each SU is calculated locally. In the RA algorithm, the computational offloading ratio of each SU in each CRB is randomly assigned. At each point, the average energy consumption of each algorithm is calculated over 30 experimental rounds, each containing 100 time slots. The total amount of unloaded data per SU and the length of each time slot are set to 160knats and 500ms, respectively. The number of SUs and the length of each time slot are set to 16 and 500ms, respectively. The number of SUs and the total amount of unloaded data are 16 and 160knats, respectively. From Figure 1-3, two observations can be obtained. First, the energy consumption of all algorithms increases with the increase in the number of SUs and the total amount of data offloaded by each SU, or the decrease in the length of each time slot. The reason is that when the total computing demand of the SU becomes larger, or the delay constraint becomes stricter, a larger computing frequency needs to be allocated to each SU, resulting in increased energy consumption. Second, compared with the TD3, LO, and RA algorithms, the SAC algorithm can save an average of 56.3%, 88.9%, and 73.2% of energy consumption, respectively. The reason is that the SAC algorithm uses entropy regularization to increase the randomness of action selection, which makes the SAC algorithm have a greater probability of finding the optimal action and obtaining a more reasonable offloading decision, thereby significantly reducing system energy consumption.

[0119] The above description is merely an embodiment of the present invention and does not limit the technical scope of the present invention. Therefore, any minor modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method. The CCHN consists of a group of SUs, a group of CR routers, and an SSP. The SSP centrally manages the SUs and CR routers via an established common control channel. The CR routers are equipped with computing resources and act as MEC servers. The uplink cellular spectrum of adjacent cells is divided into a series of cellular resource blocks (CRBs). CRB allocation is completed, and each CR router has been allocated a CRB for data offloading from its connected SUs. The method is characterized by: The decision optimization method is specifically as follows: Step 1. Set the discount factor , soft copy factor , maximum number of time slots , maximum round ; Step 2: Randomly initialize the Actor's neural network parameters , the main neural network parameters of Critic , initialize the target neural network parameters of Critic , clear the replay experience pool, that is , the current time slot number is , the current round number is ; Step 3: Randomly generate an offload ratio action to obtain the state in a CCHN-MEC environment ; Step 4: According to the status , the SAC agent outputs a decision action to calculate the unloading ratio ; Step 5: Decide on an action based on the uninstall ratio , for each SU , solve the local energy consumption minimization problem and obtain the local optimal CPU frequency allocation ; Step 6: Decision based on uninstall ratio , for each NOMA group ,get CPU frequency of each SU ; Step 7: Decision based on uninstall ratio , for each NOMA group , solve the unloading energy consumption optimization problem, and obtain the group The transmission power and transmission time of each SU ; Step 8. Calculate the total energy consumption of the system based on steps 5, 6, and 7. ; Step 9. Calculate the current system reward based on step 8 , and get the state of the next state ; Step 10: The experience of the current time slot Store in replay experience pool middle; Step 11: If you are replaying the experience pool The number of experiences in is greater than the minimum batch size , then from the replay experience pool Randomly extract data Perform network training to update network parameters , , , and the temperature coefficient ; Step 12 ; If the current time slot number ,but , ,like Then go to step 13; otherwise, return to step 3; Step 13: Output the optimal parameters of the Actor's neural network , through this parameter , Actor can output the optimal decision action in each state; In step 5, the local energy consumption minimization problem is: (6) Among them, the constraints Indicates that the local computing time cannot exceed the length of a time slot, the constraint Indicates that the local computing CPU frequency of the SU cannot exceed the maximum CPU frequency; Indicates SU The energy consumption weight, represents the computational energy consumption factor of SU, Indicates time slot SU Local CPU frequency, express In the time slot The total amount of task data for local calculation, is the length of each time slot, Indicates the maximum CPU frequency of SU; In step 6, the following formula (7) is used to obtain CPU frequency of each SU : (7) in, Indicates time slot Central SU Offloading to CR routers The task data volume ratio, Indicates time slot SU Calculate the number of CPU cycles required for 1-nat task data, Indicates time slot SU The total amount of task data, Indicates time slot Offloading to CR routers The NOMA transmission time of SU, is the length of each time slot.

2. A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method according to claim 1, characterized in that: In step 7, the unloading energy consumption optimization problem is: (8) in, , Indicates that the CR router is calculating SU The energy weight consumed by the task, represents the computational energy consumption factor of the CR router, Indicates time slot Slave Transmitter To the receiver The channel gain, Indicates time slot Slave Transmitter To the receiver The normalized channel gain of Indicates time slot Medium CR router Maximum interference and noise power levels at CR router The maximum CPU frequency, Indicates SU In the time slot The total number of NOMA groups added in, Indicates SU The maximum transmit power, Indicates that it is related to CRB The associated The maximum tolerable interference at each BS, Indicates that it is related to CRB The number of all associated BSs; constraint Indicates that the transmission time of the NOMA group cannot exceed the length of each time slot, constraining Indicates that the computing CPU frequency allocated to the SU in the NOMA group cannot exceed the maximum CPU frequency of the CR router. Indicates the rate requirement for each SU to unload data, constraint Indicates the transmission power limit of each SU, constraint Indicates that the interference power of SUs in the NOMA group cannot exceed the maximum tolerable interference value of each BS.

3. A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method according to claim 1, characterized in that: The total energy consumption of the system in step 8 for: (5) in, Indicates SU The energy consumption weight, Indicates that the CR router is calculating SU The energy weight consumed by the task, Indicates SU Local energy consumption, represents the computational energy consumption factor of SU, Indicates time slot SU Offloading to CR routers The task offloading ratio, Indicates time slot SU Calculate the number of CPU cycles required for 1-nat task data, Indicates time slot SU The total amount of task data, Indicates time slot SU Local CPU frequency; Indicates SU Offloading to CR routers Energy consumption when Indicates time slot SU Offloading to CR routers The transmission power at Indicates time slot Offloading to CR routers The NOMA transmission time of SU, Indicates SU Offloading to CR routers CR router The computing energy consumption, represents the computational energy consumption factor of the CR router, In the time slot SU Offloading to CR routers CR router Assigned CPU frequency.

4. A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method according to claim 1, characterized in that: In step 9, the current system reward is calculated using formula (10) : (10) Among them, , , , and Indicates whether the time slot is violated The constraints (1) (2) (4) and whether there is a problem and In the time slot There is a solution; if constraint (1) is satisfied, then ;otherwise, ; If constraint (2) is satisfied, then ;otherwise ; If constraint (4) is satisfied, then ;otherwise, ; If the problem If there is a solution, then ;otherwise, ; If the problem If there is a solution, then ;otherwise, ; If in time slot Middle Problem / If there is no solution, Set to .

5. A dynamic generalized user NOMA grouping CCHN-MEC network offloading decision optimization method according to claim 1, characterized in that: In step 11, the network parameters are updated according to formulas (12)-(15): , , , and the temperature coefficient : The minimum Bellman residual value is calculated as follows: (12) in, ,here, It is from Resampled actions; Actor's DNN is trained by minimizing the KL divergence, i.e. minimizing (13) in is a parameter The expression, is the input Gaussian noise vector; Temperature parameters Dynamically adjust by minimizing the following formula (14) in is the target entropy constant; The weight parameters of the target DNN in each critic are updated by the soft copy method, that is, (15)。

Citation Information

Patent Citations

  • DRL-based traffic unloading algorithm for NOMA-MEC network and implementation device

    CN112911613A

  • Reinforcement learning resource allocation and task unloading method based on NOMA-MEC

    CN113543342A