A long-term energy consumption optimization method for a non-orthogonal multiple access (NOMA) based mobile edge computing (MEC) system

By designing reasonable state space, action space, and reward function in a non-fully overlapping NOMA-MEC system, and using the SAC algorithm to optimize the allocation of communication and computing resources, the problem of heterogeneous user delay constraints under time-varying channels is solved, and long-term energy consumption optimization of the MEC system is achieved, thereby reducing system energy consumption.

CN115877933BActive Publication Date: 2025-11-21XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211503083.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-11-21
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

Existing research on non-fully overlapping NOMA-MEC systems has mostly focused on resource optimization in the non-fully overlapping NOMA transmission process, failing to effectively consider the computational resource limitations of MEC servers and the heterogeneous delay constraints of users under time-varying channels, resulting in insufficient long-term energy consumption optimization of the system.

Method used

By designing state-space functions, action functions, and reward functions, and using the SAC algorithm to jointly optimize the allocation of communication and computing resources, the long-term energy consumption of the MEC system is optimized. Considering time-varying channels and heterogeneous user delay constraints, the system optimizes user transmit power, transmission duration, and CPU frequency allocation of the MEC server.

Benefits of technology

Under time-varying channel conditions, compared with traditional fully overlapped NOMA and TDMA transmission methods, the average system energy consumption is reduced by 59.3% and 75.5%, respectively, achieving more efficient resource allocation and energy management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115877933B_ABST
    Figure CN115877933B_ABST
Patent Text Reader

Abstract

The application relates to a long-term energy consumption optimization method of a non-fully overlapped NOMA-based MEC system, which converts a long-term energy consumption minimization problem of a non-fully overlapped NOMA-MEC system into an optimal resource allocation problem in each time slot under a time-varying channel, considers heterogeneous delay constraint requirements of users in the same NOMA group, and solves the problem by designing a reasonable S t ,A t ,R t , converts it into a DRL problem, and solves the problem by using a SAC algorithm to obtain close-to-optimal user transmission power p k,j , each transmission period duration d j , CPU frequency allocation f k of the MEC server. Compared with a conventional fully overlapped NOMA and a TDMA transmission mode, the scheme can averagely reduce total energy consumption by 59.3% and 75.5% respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy consumption optimization for MEC systems, and specifically to a long-term energy consumption optimization method for MEC systems based on non-completely overlapping NOMA, applicable to non-completely overlapping NOMA-MEC systems with multiple users in a single packet under time-varying channels. Background Technology

[0002] With the rapid development of mobile internet technology, various computationally intensive and latency-sensitive applications have emerged. To meet the rapidly growing low-latency computing demands of lightweight devices, Mobile Edge Computing (MEC) has been proposed. To improve spectral efficiency, Non-Orthogonal Multiple Access (NOMA) is widely used in the task offloading process of MEC systems. Traditional NOMA-MEC system research typically employs fully overlapping techniques, meaning users within a NOMA group have the same transmission time. However, in latency-sensitive MEC scenarios, users within the same NOMA group may have heterogeneous transmission latency requirements. This has prompted researchers to propose non-fully overlapping NOMA, allowing users within the same NOMA group to have different transmission times. Furthermore, energy consumption is a crucial performance indicator in MEC systems that cannot be ignored. Non-fully overlapping NOMA introduces greater flexibility to the system by controlling the degree of overlap in users' use of resource blocks, helping to reduce transmission energy consumption on the user side and further improving the system's energy-saving effect.

[0003] However, most current research on non-fully overlapped NOMA-MEC focuses only on resource optimization during the non-fully overlapped NOMA transmission process. In reality, the computational resources of MEC servers are typically not unlimited, and these studies are mostly single-cycle resource allocation optimizations for static scenarios, failing to characterize the long-term computational offloading performance of the MEC system. Therefore, under time-varying channel conditions, it is crucial to address the delay constraints of multiple heterogeneous users within the same NOMA packet. Based on a non-fully overlapped NOMA transmission scheme, this research should jointly consider the user-end task transmission process and the MEC server-end task processing process, investigating how to allocate communication and computational resources to minimize the long-term energy consumption of the MEC system. Summary of the Invention

[0004] To address the problems existing in the prior art, the present invention aims to provide a long-term energy consumption optimization method for MEC systems based on non-fully overlapping NOMA. This method considers time-varying channels and, under the premise of satisfying the heterogeneous delay constraints of multiple users within the same NOMA group, optimizes the allocation of communication and computing resources by rationally designing state-space functions, action functions, reward functions, and constraint penalty functions, and using the SAC algorithm to achieve long-term energy consumption optimization of the MEC system.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A long-term energy consumption optimization method for MEC systems based on non-fully overlapping NOMA, the optimization method comprising the following steps:

[0007] Step 1: Place a group of users U = {u} in the MEC system. k A BS equipped with a MEC server, 1≤k≤K, is used to form a NOMA packet with all users, dividing time into multiple time slots labeled t={1,2,…,T}, with each time slot having a length of τ;

[0008] Step 2: Within each time slot t, sort the tasks based on the maximum computation deadline for each user in that time slot, i.e.

[0009] Step 3: Calculate user u in time slot t k List the transmission rates during each available transmission period and the constraints that must be met during the task transmission process;

[0010] (1) User u in time slot t k The transmission rate r during transmission time period j k,j (t) is represented as:

[0011]

[0012] Where B is the uplink cellular channel bandwidth, n0 is the channel noise power spectral density, and p k,j (t) represents user u k The transmit power during transmission period j, g k (t) is u k Channel gain in time slot t; U j (t) represents the set of users that transmit within transmission time period j in time slot t, u k' This indicates that transmission occurs within transmission period j and the channel gain is greater than u. k Bad users, For user u k The sum of interference from other users received during transmission period j;

[0013] (2) Constraint C1: To ensure user u k The offloading must be completed by the end of its last available transmission period, and the total amount of data offloaded must not be less than the number of bits L of its computational task. k (t);

[0014] (3) Constraint C2: User u kIts transmission power during any available transmission period does not exceed its maximum transmission power.

[0015] Step 4: Calculate user u in time slot t k The transmission energy consumed during its available transmission period;

[0016] User u k The transmission energy consumption in time slot t is expressed as the sum of energy consumption generated during all available transmission periods. Then, user u... k The total transmission energy consumption in time slot t can be expressed as:

[0017]

[0018] in, Indicates user u k The transmission energy consumption generated during transmission period j, d j (t) represents the length of transmission time j;

[0019] Step 5: Request the MEC server on the BS to process user u k List the computation time and corresponding computational energy consumption required for the task in time slot t, and list the constraints that should be met when processing user tasks on the MEC server.

[0020] (1) Allocate the BS to user u in time slot t. k The CPU frequency is expressed as f k (t), then u is processed on the MEC server. k The computation time and energy consumption required for the computation task are expressed as follows:

[0021]

[0022]

[0023] Among them, C k (t) represents the computation unit bit user u k The number of CPU cycles required for the computational task in time slot t, κ, is a constant that depends on the MEC server hardware architecture;

[0024] (2) Constraint C3: Considering that the computing resources of MEC servers are usually limited, the sum of the CPU frequencies allocated to all users in any time slot cannot exceed their maximum computing frequency F. max ;

[0025] (3) Constraint C4: User u k The total time spent on task computation and unloading cannot exceed the maximum computation deadline of the computation task.

[0026] Step 6: Considering time-varying channels, in a non-fully overlapping NOMA-MEC system, the long-term energy consumption of the system is minimized by jointly optimizing user transmit power, duration of each available transmission period, and allocation of computational resources. This optimization problem is described as P1.

[0027] First, the overall energy consumption of a non-fully overlapped NOMA-MEC system in time slot t is expressed as the weighted sum E of the transmission energy consumption at the user end and the computing energy consumption at the MEC server end. total (t),

[0028]

[0029] Where ω is the weighting factor, representing the weight value of the user-end transmission energy consumption.

[0030] If the long-term energy consumption of a non-completely overlapping NOMA-MEC system is expressed as the average of the total energy consumption generated over all time slots, then the long-term energy consumption optimization problem of the non-completely overlapping NOMA-MEC system over a time period T can be expressed as P1:

[0031]

[0032] Step 7: Restate the communication and computational resource allocation problem of the non-completely overlapping NOMA-MEC system in each time slot into a deep reinforcement learning problem;

[0033] (1) Define the state space: Define the state space of time slot t as the reward function of the previous time slot t-1, then we have: S t =R t-1 ;

[0034] (2) Define the action space: The actions of an agent in time slot t include the transmit power p of each user in each available transmission period. k,j (t) Duration d of each transmission period in non-perfectly overlapping NOMA j (t) and the CPU frequency allocation of the MEC server f k (t), then the action space of time slot t is defined as:

[0035] A(t) = [p 1,1 (t),…,p k,j (t),…,p K,K (t); d1(t),…,d K (t); f1(t),…,f K (t)];

[0036] (3) Design the reward function: The reward function should be related to the system energy consumption and constraints. The goal of reinforcement learning is to maximize the reward (while minimizing the sum of system energy consumption and penalty for violating constraints). The instantaneous reward R obtained by the agent in time slot t is calculated as follows. t Defined as:

[0037] R t =exp(4*(-E total (t)))-(β1+β2+β3+β4)

[0038] Where β1, β2, β3, and β4 represent the penalty values ​​that will be generated if constraints C1, C2, C3, and C4 are violated within time slot t, respectively.

[0039] Step 8: Learn the near-optimal communication and computing resource allocation of the non-completely overlapping NOMA-MEC system in each time slot t using the SAC-based DRL algorithm.

[0040] Step 8 is described in detail below:

[0041] Given the maximum number of training epochs Γ, the maximum number of time slots T in a single epoch, the discount factor γ, the soft replication factor ζ, and the minimum number of training samples |Z|, clear the experience buffer and randomly initialize the neural network parameters of the actor. critic's main neural network parameters θ i (i = 1, 2), initialize the target neural network parameters for the critic.

[0042] Treating the non-fully overlapping NOMA-MEC system as a whole as the environment, in each time slot t, the SAC agent bases its current state S on the observed non-fully overlapping NOMA-MEC system. t Make the corresponding communication and computing resource allocation decision A designed in step 7.(2). t The environment will affect action A. t Provide feedback and calculate the corresponding immediate reward R based on the reward function designed in step 7.(3). t And transition to the next state S t+1 , will (S t A t ,R t ,S t+1 This data is stored as a set of historical experience data in the experience buffer.

[0043] When the amount of historical experience data in the experience buffer is greater than the minimum number of training samples |Z|, a set of experience samples of size |Z| is extracted to train and update the relevant parameters of the SAC algorithm: θ i (i = 1, 2) After Γ rounds of training, the SAC agent will output the optimal DNN weight coefficients of the actor network. The optimal user transmit power, the optimal duration allocation of each transmission period, and the optimal allocation of computing resources are obtained to minimize the long-term energy consumption of the system.

[0044] In step 3, constraints C1 and C2 are represented as follows:

[0045]

[0046]

[0047] In step 5, constraints C3 and C4 are represented as follows:

[0048]

[0049]

[0050] In step 7, β1, β2, β3, and β4 represent the following:

[0051]

[0052]

[0053]

[0054]

[0055] By adopting the above scheme, this invention, under time-varying channel conditions, considers the heterogeneous time delay constraints of users in the same NOMA group, transforming the long-term energy consumption minimization problem of a non-completely overlapping NOMA-MEC system into an optimal resource allocation problem within each time slot. This is achieved by designing a reasonable S... t A t ,R t This problem is transformed into a DRL problem, and the SAC algorithm is used to obtain a near-optimal user transmit power p. k,j Duration d of each transmission period j CPU frequency allocation for MEC servers k Compared with traditional fully overlapped NOMA and TDMA transmission methods, the proposed solution reduces total energy consumption by an average of 59.3% and 75.5%, respectively. Attached Figure Description

[0056] Figure 1 The average reward obtained by the SAC algorithm for the MEC system based on the three different transmission modes considered in this invention: non-completely overlapping NOMA transmission mode, completely overlapping NOMA, and TDMA.

[0057] Figure 2 This paper compares the long-term energy consumption of the present invention (non-fully overlapping NOMA-MEC) with that of fully overlapping NOMA-MEC and TDMA-MEC under the condition of varying total offload data for each user.

[0058] Figure 3 This is a comparison of the long-term energy consumption of the present invention (non-fully overlapping NOMA-MEC) with that of fully overlapping NOMA-MEC and TDMA-MEC under the variation of the length of a single time slot. Detailed Implementation

[0059] This invention discloses an optimization method for communication and computing resource allocation in a non-completely overlapping NOMA-MEC system with multiple users within a single packet under time-varying channels. The application scenario is an uplink, comprising a BS (base station) equipped with an MEC server module and a group of users U = {u k With the goal of optimizing long-term system energy consumption, considering the latency constraints of heterogeneous users, all users form a single NOMA packet and share a single cellular channel for transmission. Time is divided into multiple time slots, labeled t={1,2,…,T}, with each time slot having a length of τ.

[0060] Within each time slot, each user has an indivisible and latency-sensitive computing task. Users offload their computing tasks to the MEC server on the BS (Browser-Based Service) using a non-fully overlapped (NOMA) transport method. The order in which users complete the offloading is scheduled based on the ascending order of their maximum computation deadlines. Based on this ranking, the BS allocates corresponding non-completely overlapping NOMA available transmission periods to each user. By applying channel gain-based SIC technology to decode the superimposed signals from users in each transmission period, the BS allocates different computation frequencies to each user, processes their offloaded computational tasks in parallel, and generates corresponding computational energy consumption.

[0061] This system considers a time-varying channel model, and the user u in time slot t k Channel gain g to BS k (t) is represented as

[0062]

[0063] Among them l k For user u k The path distance to BS, where α is the path loss factor. It is the large-scale fading coefficient, which remains constant across all time slots; |ε k (t)| 2It is a small-scale fading coefficient that remains constant within a single time slot but varies across different time slots, ε. k (t) can be expressed as

[0064]

[0065] in, ε is the channel correlation coefficient between different time slots. k (0) follows a complex Gaussian distribution with mean and unit variance of 0; v k (t) is an independent and identically distributed random variable with a mean of 0 and a variance of 0. The complex Gaussian distribution.

[0066] This system is based on the fact that the computing resources of MEC servers are usually not unlimited. It also considers the task transmission process on the user side and the task computing process on the MEC server side. In each time slot t, the communication and computing resources of the system are optimized and allocated. The system energy consumption mainly consists of two parts:

[0067] (1) User transmission energy consumption caused by user terminal unloading task data:

[0068]

[0069] Where p k,j (t)≥0 represents user u k The transmit power during transmission period j, d j (t)>0 represents the length of transmission time j.

[0070] (2) Computing energy consumption generated by the MEC server in processing user tasks:

[0071]

[0072] Where κ is a constant that depends on the MEC server hardware architecture, f k (t)>0 indicates that BS allocates time slot t to user u. k CPU frequency, C k (t) represents the computation unit bit user u k Computational task L in time slot t k (t) The number of CPU cycles required.

[0073] The weighting factor ω represents the weight value of the user-end transmission energy consumption. The weighted sum of the user-end transmission energy consumption and the MEC server-side computation energy consumption is used as the total system energy consumption generated in time slot t. Specifically, it can be expressed as:

[0074]

[0075] The long-term energy consumption of a non-fully overlapping NOMA-MEC system is represented by the average of the total energy consumption generated over all time slots. Therefore, the long-term energy consumption of the entire system can be calculated as follows:

[0076]

[0077] The communication and computational resource allocation problem of a non-fully overlapping NOMA-MEC system over a time period T can be represented as P1:

[0078]

[0079]

[0080]

[0081]

[0082]

[0083] var.p k,j ≥0,d j >0,f k >0.

[0084] Where constraint C1 represents user u k The total amount of data offloaded during all available transmission periods must not be less than the number of bits L of its computational task. k (t), where r k,j Indicates user u in time slot t k The transmission rate during transmission period j is calculated as follows:

[0085]

[0086] Where B is the uplink cellular channel bandwidth, n0 is the channel noise power spectral density, and p k,j (t)≥0 represents user u k The transmit power during transmission period j, g k (t) is u k Channel gain in time slot t. U j (t) represents the set of users that transmit within transmission time period j in time slot t, u k' This indicates that transmission occurs within transmission period j and the channel gain is greater than u. k Bad users, For user u k The sum of interference from other users received during transmission period j. Constraint C2 represents user u. k Its transmission power during any available transmission period does not exceed its maximum transmission power. Constraint C3 states that in any time slot, the sum of the CPU frequencies allocated by the BS to all users cannot exceed the maximum computing frequency F of the MEC server. max Constraint C4 represents user u k The total time spent on task computation and unloading cannot exceed the maximum computation deadline of the computation task.

[0087] Clearly, problem P1 is a multi-user, multi-variable, tightly coupled non-convex problem, and it considers time-varying channels. Therefore, we formulate the allocation decision of communication and computing resources for the non-completely overlapping NOMA-MEC system within each time slot as a DRL problem. By reasonably designing the state-space function, action function, reward function, and constraint penalty function, and using the SAC algorithm to jointly optimize the user transmit power p, we can achieve the desired results. k,j Duration d of each transmission period j CPU frequency allocation for MEC servers k This achieves the goal of minimizing the system's long-term energy consumption.

[0088] The optimization method of the present invention specifically includes the following steps:

[0089] Step 1: Place a group of users U = {u k A BS equipped with a MEC server divides time into multiple time slots, labeled as t={1,2,…,T};

[0090] Step 2: Within each time slot t, sort the tasks based on the maximum computation deadline for each user in that time slot, i.e.

[0091] Step 3: Calculate user u in time slot t k List the transmission rates during each available transmission period and the constraints that must be met during the task transmission process;

[0092] (1) User u in time slot t k The transmission rate r during transmission time period j k,j (t) can be expressed as:

[0093]

[0094] (2) Constraint C1: To ensure user u k The offloading must be completed by the end of its last available transmission period, and the total amount of data offloaded must not be less than the number of bits L of its computational task. k (t),

[0095]

[0096] (3) Constraint C2: User u kIts transmission power during any available transmission period does not exceed its maximum transmission power.

[0097]

[0098] Step 4: Calculate user u in time slot t k The transmission energy consumed during its available transmission period;

[0099] User u k The transmission energy consumption in time slot t is expressed as the sum of energy consumption generated during all available transmission periods. Then, user u... k The total transmission energy consumption in time slot t can be expressed as:

[0100]

[0101] in, Indicates user u k The transmission energy consumption generated during transmission period j, d j (t)>0 represents the length of transmission time j.

[0102] Step 5: Request the MEC server on the BS to process user u k List the computation time and corresponding computational energy consumption required for the task in time slot t, and list the constraints that should be met when processing user tasks on the MEC server.

[0103] (1)f k (t)>0 indicates that BS allocates time slot t to user u. k The CPU frequency is then used to process u on the MEC server. k The computation time and energy consumption required for the computation task can be expressed as follows:

[0104]

[0105]

[0106] (2) Constraint C3: Considering that the computing resources of MEC servers are usually limited, the sum of the CPU frequencies allocated to all users in any time slot cannot exceed their maximum computing frequency F. max ,

[0107]

[0108] (3) Constraint C4: User u k The total time spent on task computation and unloading cannot exceed the maximum computation deadline of the computation task.

[0109]

[0110] Step 6: Considering the time-varying channel, calculate the overall energy consumption E of the non-completely overlapping NOMA-MEC system in time slot t. total (t) represents the weighted sum of the transmission energy consumption at the user end and the computing energy consumption at the MEC server end, representing the long-term energy consumption E of the system. total The long-term energy consumption optimization problem P1 of the non-completely overlapping NOMA-MEC system over a time period T is obtained by representing the sum of the average energy consumption generated in all time slots.

[0111] Step 7: Restate the communication and computational resource allocation problem of the non-completely overlapping NOMA-MEC system in each time slot into a deep reinforcement learning (DRL) problem;

[0112] (1) Define the state space: Define the state space of time slot t as the reward function of the previous time slot t-1, then we have: S t =R t-1 ;

[0113] (2) Define the action space: The actions of an agent in time slot t include the transmit power p of each user in each available transmission period. k,j (t) Duration d of each transmission period in non-perfectly overlapping NOMA j (t) and the CPU frequency allocation of the MEC server f k (t), then the action space of time slot t is defined as:

[0114] A(t) = [p 1,1 (t),…,p k,j (t),…,p K,K (t); d1(t),…,d K (t); f1(t),…,f K (t)];

[0115] (3) Design the reward function: The reward function should be related to the system energy consumption and constraints. The goal of reinforcement learning is to maximize the reward (while minimizing the sum of system energy consumption and penalty for violating constraints). The instantaneous reward R obtained by the agent in time slot t is calculated as follows. t Defined as:

[0116] R t =exp(4*(-E total (t)))-(β1+β2+β3+β4)

[0117] Where β1, β2, β3, and β4 represent the penalty values ​​that will be generated if constraints C1, C2, C3, and C4 are violated within time slot t, respectively, and can be specifically represented as follows:

[0118]

[0119]

[0120]

[0121]

[0122] Step 8: Learn the near-optimal communication and computing resource allocation of the non-completely overlapping NOMA-MEC system in each time slot t using the SAC-based DRL algorithm;

[0123] Given the maximum number of training epochs Γ, the maximum number of time slots T in a single epoch, the discount factor γ, the soft replication factor ζ, and the minimum number of training samples |Z|, clear the experience buffer and randomly initialize the neural network parameters of the actor. critic's main neural network parameters θ i (i = 1, 2), initialize the target neural network parameters for the critic.

[0124] Treating the non-fully overlapping NOMA-MEC system as a whole as the environment, in each time slot t, the SAC agent bases its current state S on the observed non-fully overlapping NOMA-MEC system. t Make the corresponding communication and computing resource allocation decision A designed in step 7.(2). t The environment will affect action A. t Provide feedback and calculate the corresponding immediate reward R based on the reward function designed in step 7.(3). t And transition to the next state S t+1 , will (S t A t ,R t ,S t+1 This data is stored as a set of historical experience data in the experience buffer.

[0125] When the amount of historical experience data in the experience buffer is greater than the minimum number of training samples |Z|, a set of experience samples of size |Z| is extracted to train and update the relevant parameters of the SAC algorithm: θ i (i = 1, 2) After Γ rounds of training, the SAC agent will output the optimal DNN weight coefficients of the actor network. The optimal user transmit power, the optimal duration allocation of transmission periods, and the optimal allocation of computing resources are obtained to minimize the long-term energy consumption of the system.

[0126] To evaluate the performance of this invention, the following simulation was performed. The simulation parameters were set as follows: users were uniformly distributed in a cell with a radius of 300m, the weight factor ω was set to 0.5, and the maximum computation cutoff delay of the users was randomly generated in [τ-200ms,τ], where τ represents the length of each time slot in ms. Other parameters are shown in Table 1.

[0127]

[0128] Table 1 Simulation Parameters

[0129] In the simulation, the number of users and the amount of task data per user L are... k The length of a single time slot, τ, was set to 5, 140 kNats, and 400 ms, respectively. For comparison, with the actor network learning rate, critic network learning rate, and minimum number of training samples of the SAC algorithm unchanged, training was performed based on fully overlapping NOMA and TDMA transmission methods, respectively. Figure 1 The average reward obtained by the SAC algorithm based on three different transmission modes—non-perfectly overlapping NOMA, fully overlapping NOMA, and TDMA—is shown. Figure 1 It can be seen that compared with fully overlapped NOMA and TDMA, non-fully overlapped NOMA can obtain a higher average reward. This is because the non-fully overlapped NOMA transmission method takes into account the low latency requirements of heterogeneous users while reusing the spectrum, and can more flexibly adjust the degree of spectrum reuse to achieve better energy saving. Figure 2 , Figure 3 The long-term system energy consumption based on these three different transmission methods was compared when the user task data volume and the length of a single time slot changed. At each point, the average energy consumption over 30 experimental rounds was taken, with each round containing 20 time slots. Figure 2 , Figure 3 First, we can see that as the amount of user task data increases or the length of a single time slot decreases, the system energy consumption increases under all three transmission methods. This is because when the amount of user task data is larger or the user task latency constraints are stricter, the transmission energy consumption at the user end will be greater. The MEC server also needs to allocate more computing frequency to process each user task, resulting in greater computing energy consumption, ultimately leading to an increase in the overall system energy consumption. Second, compared to fully overlapped NOMA and TDMA, the non-fully overlapped NOMA transmission method can reduce energy consumption by an average of 59.3% and 75.5%, respectively.

[0130] The above description is merely an embodiment of the present invention and does not constitute any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. A long-term energy consumption optimization method for MEC systems based on non-completely overlapping NOMA, characterized in that: The optimization method includes the following steps: Step 1: Place a group of users in the MEC system A BS equipped with a MEC server allows all users to collectively form a NOMA packet, dividing time into multiple time slots and labeling them as... Each time slot is [length missing] ; Step 2, in each time slot Within that time slot, tasks are sorted based on their maximum computation deadline for each user. ; Step 3: Determine the time slot Chinese users List the transmission rates during each available transmission period and the constraints that must be met during the task transmission process; (1) Time slot Chinese users During transmission period Internal transmission rate Represented as: in, For uplink cellular channel bandwidth, For channel noise power spectral density, For users During transmission period The transmission power in the middle, for In the time slot Channel gain; Indicates time slot During the transmission period The set of users that transmit within the network. Indicates during the transmission period Internal transmission and channel gain ratio Bad users, For users During transmission period The sum of interference received from other users; (2) Constraints To ensure users The offloading must be completed by the end of its last available transmission period, and the total amount of data offloaded must not be less than the number of bits of its computational task. ; (3) Constraints :user Its transmission power during any available transmission period does not exceed its maximum transmission power. ; Step 4: Determine the time slot Chinese users The transmission energy consumed during its available transmission period; User In the time slot The transmission energy consumption in the context is expressed as the total energy consumption generated during all available transmission periods. Then, the user... In the time slot The total transmission energy consumption can be expressed as: in, Indicates user During transmission period The transmission energy consumption generated in the process Transmission period Length; Step 5: Request the MEC server on the BS to process users. In the time slot The computation time and corresponding energy consumption required for the tasks in the MEC server are listed, along with the constraints that should be met when processing user tasks on the MEC server. (1) Place the BS in the time slot Distributed to users CPU frequency is expressed as Then it will be processed on the MEC server. The computation time and energy consumption required for the computation task are expressed as follows: in, Represents the user of the unit of computation (bit). In the time slot The number of CPU cycles required for the computational tasks in the process. It is a constant that depends on the MEC server hardware architecture; (2) Constraints Given the limited computing resources of the MEC server, the sum of the CPU frequencies allocated to all users in any time slot cannot exceed their maximum computing frequency. ; (3) Constraints :user The total time spent on task computation and unloading cannot exceed the maximum computation deadline of the computation task. ; Step 6: Considering time-varying channels, in a non-fully overlapping NOMA-MEC system, minimize the long-term energy consumption of the system by jointly optimizing user transmit power, the duration of each available transmission period, and computational resource allocation. The optimization problem is described as follows: ; First, the non-fully overlapping NOMA-MEC system is time-slotted. The overall energy consumption is expressed as the weighted sum of the transmission energy consumption at the user end and the computing energy consumption at the MEC server end. , in It is a weighting factor, representing the weight value of the user-end transmission energy consumption; If the long-term energy consumption of a non-completely overlapping NOMA-MEC system is expressed as the average of the total energy consumption generated over all time slots, then the long-term energy consumption of the non-completely overlapping NOMA-MEC system over a period of time... The long-term energy consumption optimization problem within the range can be expressed as: : Step 7: Restate the communication and computational resource allocation problem of the non-completely overlapping NOMA-MEC system in each time slot into a deep reinforcement learning problem; (1) Define the state space: divide the time slots The state space is defined as the previous time slot. The reward function is then: ; (2) Define the action space: the agent in the time slot The actions include each user's transmit power during each available transmission period. Duration of each transmission period in non-perfectly overlapping NOMA CPU frequency allocation for MEC servers So, time slot The action space is defined as: ; (3) Design the reward function: The reward function should be related to the system energy consumption and constraints. The goal of reinforcement learning is to maximize the reward, while minimizing the sum of system energy consumption and penalty for violating constraints, so that the agent can perform tasks in time slots. Instant rewards received Defined as: in, They represent time slots respectively If the internal rules are violated ,constraint ,constraint ,constraint The corresponding penalty value will be generated; Step 8: Learn each time slot using the SAC-based DRL algorithm. The near-optimal allocation of communication and computing resources in this non-fully overlapping NOMA-MEC system.

2. The long-term energy consumption optimization method for an MEC system based on non-completely overlapping NOMA according to claim 1, characterized in that: Step 8 is described in detail below: Given the maximum number of training rounds Maximum number of time slots in a single round Discount Factor soft replication factor Minimum number of training samples Clear the experience buffer and randomly initialize the actor's neural network parameters. The main neural network parameters of the critic Initialize the target neural network parameters for the critic. ; Treating the non-fully overlapping NOMA-MEC system as a whole as an environment, in each time slot In this context, the SAC agent bases its current state on the observed non-fully overlapping NOMA-MEC system. Make corresponding communication and computing resource allocation decisions as designed in step 7 (2). The environment will affect the action. Provide feedback and calculate the corresponding immediate reward based on the reward function designed in step 7.(3). And transition to the next state. ,Will This set of historical experience data is stored in the experience buffer. When the amount of historical experience data in the experience buffer is greater than the minimum number of training samples When, a group of sizes is drawn. Empirical samples are used to train and update the relevant parameters of the SAC algorithm: , , , ,when After each training round, the SAC agent will output the optimal DNN weight coefficients of the actor network. This yields the optimal user transmit power, the optimal duration allocation for each transmission period, and the optimal allocation of computing resources that minimize the long-term energy consumption of the system.

3. The long-term energy consumption optimization method for MEC systems based on non-completely overlapping NOMA according to claim 1, characterized in that: In step 3, constraints and constraints They are represented as follows: 。 4. The long-term energy consumption optimization method for an MEC system based on non-completely overlapping NOMA according to claim 1, characterized in that: In step 5, constraints and constraints They are represented as follows: 。 5. The long-term energy consumption optimization method for an MEC system based on non-completely overlapping NOMA according to claim 1, characterized in that: In step 7 They are represented as follows: , , , 。