A User Grouping and Resource Allocation Method and Device in a NOMA-MEC System
By adopting the hybrid deep reinforcement learning method in the NOMA-MEC system, a hybrid deep reinforcement learning network is built, and user grouping, calculation offloading and bandwidth allocation ratios are dynamically adjusted, which solves the problems of complex calculation of resource allocation strategies and lack of adaptive capabilities in the existing technology, and maximizes system energy efficiency.
Patent Information
- Application Number
- CN202210282489.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-22
AI Technical Summary
In the existing NOMA-MEC system, the user grouping and resource allocation strategies are complex and lack dynamic adaptability, making it difficult to schedule resources in a dynamic system in real time to maximize system energy efficiency.
A hybrid deep reinforcement learning method is adopted to build a hybrid deep reinforcement learning network, and the actions are generated through the time slot state input network, including user grouping, calculation of offloading and bandwidth allocation ratios, and dynamically adjust resource allocation to maximize system energy efficiency.
It realizes real-time scheduling of resources in a dynamic NOMA-MEC system, maximizes system energy efficiency, and makes optimal decisions in a dynamic environment, and overcomes the shortcomings of a single deep reinforcement learning method to deal with continuous and discrete action spaces.
Smart Images

Figure CN114885420B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile communications and deep reinforcement learning, and specifically relates to a computation offloading method and device in a NOMA-MEC system based on hybrid deep reinforcement learning. Background Art
[0002] With the significant increase in the number of smart devices, a large number of user devices generate a large amount of data that needs to be processed. However, due to the size limitations of smart devices, their computing resources and energy resources are very scarce, which makes them face huge challenges in service demand. Therefore, in order to improve the efficiency of task processing and meet service needs, Mobile Edge Computing (MEC) technology came into being. In addition, the explosive growth of data traffic has caused an urgent need for massive access and a sharp shortage of spectrum resources. The Non-Orthogonal Multiple Access (NOMA) technology in the fifth generation (5G) communication is an effective solution to these problems. Therefore, the technical research of NOMA-MEC has attracted widespread attention in recent years.
[0003] At present, most of the research on user grouping and resource allocation strategies in NOMA-MEC systems uses traditional optimization methods to solve them, such as obtaining the optimal solution through iterative algorithm convergence, or obtaining suboptimal solutions through heuristic algorithms. However, these methods either have too high computational complexity or can only obtain suboptimal solutions, and more importantly, lack the ability to adapt to dynamic systems. Summary of the invention
[0004] The purpose of the present invention is to propose a user grouping and resource allocation method in a NOMA-MEC system based on hybrid deep reinforcement learning, which can schedule resources in a dynamic NOMA-MEC system in real time to maximize the system energy efficiency.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for user grouping and resource allocation in a NOMA-MEC system based on hybrid deep reinforcement learning, comprising the following steps:
[0007] Step 1: describe the NOMA-MEC system, which operates in a time slot manner, and the time slot set is denoted as Γ = {1, 2, ..., T};
[0008] Step 2: Define the energy efficiency of the system;
[0009] Step 3: Describe the optimization problem;
[0010] Step 4: Define the state space of deep reinforcement learning and the action space of deep reinforcement learning;
[0011] Step 5: construct a hybrid deep reinforcement learning network; the input of the network is the state and the output is the action;
[0012] Step 6: Input each time slot state into the hybrid deep reinforcement learning network to generate actions;
[0013] Step 7: Train the hybrid deep reinforcement learning network;
[0014] Step 8: Repeat steps 6 and 7 until the number of repetitions reaches the specified number of time slots T, and then output the actions generated at this time, that is, the decisions to be optimized: user grouping, computation offloading, and bandwidth allocation ratio.
[0015] Further, the method of describing the NOMA-MEC system includes:
[0016] The NOMA-MEC system consists of K user devices and a single-antenna base station connected to an edge server, and all users have only a single transmitting antenna to establish a communication link with the base station. The system operates in a time slot manner, and the time slot set is recorded as Γ = {1, 2, ..., T};
[0017] The total system bandwidth B is divided into N orthogonal sub-channels, and the bandwidth of sub-channel n accounts for the proportion of the total bandwidth τ n , Define K = {1, 2, ..., K} and N = {1, 2, ..., N} to represent the user set and orthogonal subchannel set respectively, K ≤ 2N;
[0018] The whole process is divided into time slots, Γ = {1, 2, ..., T}; the channel gain remains unchanged within a time slot and changes between different time slots, h nk , n∈N, k∈K represents the channel gain from user k to the base station on channel n, and let h n1 <h n2 <....<h nK ,n∈N;
[0019] Limit a channel to at most two user signals for simultaneous transmission, and a user can only send signals on one channel in a time slot; nk =1 means channel n is allocated to user k to send signals, m nk =0 means that channel n is not allocated to user k to send signals.
[0020] Furthermore, the method of defining the energy efficiency of the system in step 2 includes:
[0021] Step 2.1) The energy efficiency Y of the system is defined as the sum of the ratio of the computing rate to the computing power of all users, as shown in the following formula:
[0022]
[0023] Among them, R i,off represents the computing rate at which user i offloads computing tasks to edge servers for execution, p i is the transmit power of user i, which does not change over time, and the transmit power of all users is the same; R i,local represents the computing rate of the task executed locally by user i, p i,local represents the power of local execution of user i, x ni = 1 means user i offloads tasks to the edge server through channel n, x ni =0 means that user i does not offload tasks to the edge server through the channel;
[0024] Step 2.2) Since the channel gain h of user i on channel n is ni Greater than the channel gain h of user j nj According to the serial interference cancellation technology, the base station decodes the users in descending order of channel gain, so the unloading rate of user i is Uninstall rate of user j Where N 0 is the power spectral density of the noise,
[0025] Step 2.3) The local execution computation rates of user i and user j are where f i and f j For the user's CPU processing power, is the number of cycles required to process a 1-bit task; the computing power of user i and user j executed locally is p i,local =νf i 3 、p j,local =νf j 3 , where ν is the capacitance effectiveness coefficient of the user equipment chip architecture;.
[0026] Furthermore, the optimization problem in step 3 is described as:
[0027]
[0028] stC1:x nk ∈{0,1},m nk ∈{0,1},
[0029] C2:
[0030] C3:
[0031] C4:
[0032] Furthermore, the method of defining the state space and action space of deep reinforcement learning in step 4 includes:
[0033] Step 4.1) The state space s, s = {h 11 ,h 12 ,...h 1K ,h 21 ,h 22 ,...,h 2K ,h N1 ...h NK};
[0034] Step 4.2) The action space a consists of two stages, a = {a_c, a_d}, where a_c = {τ 1 ,τ 2 ,...,τ N} is the continuous action indicating the system bandwidth allocation ratio, a_d={m 11 ,m 12 ,...,m 1K ,...,m N1 ,m N2 ,...,m NK ,x 11 ,x 12 ,...,x 1K ,...,x N1 ,x N2 ,...,x NK} represents the subchannel allocation scheme for discrete actions;
[0035] Furthermore, the method of constructing a hybrid deep reinforcement learning network in step 5 includes:
[0036] The hybrid deep reinforcement network includes a continuous layer deep reinforcement learning network and a discrete layer deep reinforcement learning network; the continuous layer deep reinforcement learning network is DDPG, and the discrete layer deep reinforcement learning network is DQN.
[0037] Furthermore, in step 6, the method of inputting each time slot state into the hybrid deep reinforcement learning network to generate an action includes:
[0038] Step 6.1) Input the system state into the hybrid deep reinforcement learning network, and the DDPG Actor network generates a_c bandwidth allocation ratio, and the DQN network generates a_d user grouping situation;
[0039] Step 6.2) After the user grouping and bandwidth allocation ratio are determined, the maximum system energy efficiency is decomposed into the maximum energy efficiency Y of each channel n ;
[0040] The problem is transformed into
[0041]
[0042] The matrix X is initialized to a zero matrix at each time step; (x n,i ,x n,j ) has four possible values, namely (0, 0), (1, 0), (0, 1), and (1, 1). The value of x determines the offloading decision. 0 means that the computing task of the user device is not offloaded to the edge server for execution, and 1 means that it is offloaded to the edge server for execution. Substitute the four combinations into the above formula and select n The largest combination resets the value of the corresponding position of X.
[0043] Furthermore, step 7 of training the hybrid deep reinforcement learning network method includes:
[0044] When the base station is in state s, it performs action a = (a_c, a_d) and gets an immediate reward from the environment feedback. And obtain the state s' of the next time slot;
[0045] Store (s, a_c, r, s') in the DDPG experience pool, store the sample (s, a_d, r, s') in the DQN experience pool, and the DDPG network and the DQN network share the state and reward value;
[0046] The DDPG network and DQN network sample D samples from the experience pool to train and update their own parameters.
[0047] In a second aspect, the present invention provides a user grouping and resource allocation device in a NOMA-MEC system based on hybrid deep reinforcement learning, comprising the following steps:
[0048] System description module: used to describe the NOMA-MEC system;
[0049] Efficiency definition module: used to define the energy efficiency of the system;
[0050] Problem description module: used to describe the optimization problem;
[0051] Space definition module: used to define the state space and action space of deep reinforcement learning;
[0052] Network building module: used to build a hybrid deep reinforcement learning network; the input of the network is the state and the output is the action;
[0053] Action generation module: used to input each time slot state into the hybrid deep reinforcement learning network to generate actions;
[0054] Network training module: used to train hybrid deep reinforcement learning networks;
[0055] Output module: After the number of repeated training reaches the specified number of time slots T, the action generated at this time is output, that is, the decision to be optimized: user grouping, computing offloading, and bandwidth allocation ratio.
[0056] In a third aspect, the present invention provides a user grouping and resource allocation device in a NOMA-MEC system based on hybrid deep reinforcement learning, including a processor and a storage medium; the storage medium is used to store instructions;
[0057] The processor is used to operate according to the instructions to execute the steps of the method described in the first aspect.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] 1. Based on the NOMA-MEC system, this paper proposes a novel hybrid deep reinforcement learning algorithm that can solve the problem of having both discrete action space and continuous action space, and dynamically and in real time determine the sub-channel allocation, computational offloading decision, and bandwidth allocation scheme according to the system state to maximize the long-term energy efficiency of the system. The main problem solved is that the algorithm determines the bandwidth allocation ratio, user grouping, and task offloading decision according to the time-varying channel conditions;
[0060] 2. In the NOMA-MEC scenario, the present invention uses the proposed method to determine user grouping, computational offloading decisions, and bandwidth allocation ratios to maximize the ratio of the system's computational rate to power consumption.
[0061] 3. The method of the present invention can make optimal decisions in a dynamic environment, and the proposed hybrid deep reinforcement learning method can overcome the disadvantage that a single deep reinforcement learning method cannot handle tasks with both continuous action space and discrete action space. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 A schematic diagram of a system network of the present invention;
[0063] Figure 2 This is the flow chart of the hybrid deep reinforcement learning algorithm. DETAILED DESCRIPTION
[0064] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0065] Embodiment 1:
[0066] Combination Figure 1 , this embodiment is based on a method for user grouping and resource allocation in a NOMA-MEC system based on hybrid deep reinforcement learning. The method comprises the following steps:
[0067] Step 1: Describe the NOMA-MEC system. The system operates in a time slot manner, and the time slot set is denoted as Γ = {1, 2, ..., T};
[0068] Step 2: Define the energy efficiency of the system.
[0069] Step 3: Describe the optimization problem.
[0070] Step 4: Define the state space of deep reinforcement learning and define the action space of deep reinforcement learning.
[0071] Step 5: Build a hybrid deep reinforcement learning network.
[0072] Step 6: Input each time slot state into the hybrid deep reinforcement learning network to generate actions.
[0073] Step 7: Train the hybrid deep reinforcement learning network;
[0074] Step 8, repeat steps 6 and 7 until the number of repetitions reaches the specified number of time slots T, the algorithm terminates and outputs the action at this time. The action is output according to the constructed algorithm model. The action is the decision to be optimized by the present invention - user grouping, computing offloading, and bandwidth allocation ratio.
[0075] Specifically, step 1 describes the method of the NOMA-MEC system including:
[0076] Step 1.1) The NOMA-MEC system consists of K user devices and a single-antenna base station connected to an edge server, and all users have only a single transmitting antenna to establish a communication link with the base station. The total system bandwidth B is divided into N orthogonal subchannels, and the bandwidth of subchannel n accounts for a ratio of τ to the total bandwidth. n , Define K = {1, 2, ..., K} and N = {1, 2, ..., N} to represent the user set and orthogonal subchannel set, respectively, K ≤ 2N. The present invention divides the entire process into time slots, Γ = {1, 2, ..., T}. The channel gain remains unchanged within a time slot and changes between different time slots, h nk , n∈N, k∈K represents the channel gain from user k to the base station on channel n, and let h n1 <h n2 <....<h nK,n∈N. In the power domain NOMA scenario, multiple users can transmit signals in the same subchannel at the same time. In order to avoid excessive interference between users in the subchannel, the present invention limits a channel to a maximum of two user signals for simultaneous transmission, and the user only sends signals on one channel in a time slot, m nk =1 means channel n is allocated to user k to send signals, m nk =0 means that channel n is not allocated to user k to send signals.
[0077] Specifically, the method of defining the energy efficiency of the system in step 2 includes:
[0078] Step 2.1) The energy efficiency Y of the system is defined as the sum of the ratio of the computing rate to the computing power of all users, as shown in the following formula:
[0079]
[0080] For the convenience of formula expression, the present invention omits the description of time slot t. i,off represents the computing rate at which user i offloads computing tasks to edge servers for execution, p i is the transmit power of user i, which does not change over time, and the transmit power of all users is the same. i,local represents the computing rate of the task executed locally by user i, p i,local represents the power of local execution of user i, x ni = 1 means user i offloads tasks to the edge server through channel n, x ni =0 means that user i does not offload tasks to the edge server through the channel.
[0081] Step 2.2) Since the channel gain h of user i on channel n is ni Greater than the channel gain h of user j nj According to the serial interference cancellation technology, the base station decodes the users in descending order of channel gain, so the unloading rate of user i is Uninstall rate of user j Where N 0 is the power spectral density of the noise.
[0082] Step 2.3) The local execution computation rates of user i and user j are where f i and f j For the user's CPU processing power, is the number of cycles required to process a 1-bit task; the computing power of user i and user j executed locally is p i,local =νf i 3 、p j,local =νfj 3 , where ν is the capacitance effectiveness coefficient of the user equipment chip architecture;
[0083] Specifically, the optimization problem in step 3 is described as
[0084]
[0085] stC1:x nk ∈{0,1},m nk ∈{0,1},
[0086] C2:
[0087] C3:
[0088] C4:
[0089] Specifically, the method of defining the state space and action space of deep reinforcement learning in step 4 includes:
[0090] Step 4.1) The state space s, s = {h 11 ,h 12 ,...h 1K ,h 21 ,h 22 ,...,h 2K ,h N1 ...h NK}.
[0091] Step 4.2) The action space a consists of two stages, a = {a_c, a_d}, where a_c = {τ 1 ,τ 2 ,...,τ N} is the continuous action indicating the system bandwidth allocation ratio, a_d={m 11 ,m 12 ,...,m 1K ,...,m N1 ,m N2 ,...,m NK ,x 11 ,x 12 ,...,x 1K ,...,x N1 ,x N2 ,...,x NK} represents the sub-channel allocation scheme for discrete actions.
[0092] Specifically, the method of constructing a hybrid deep reinforcement learning network in step 5 includes:
[0093] Step 5.1) Construct a hybrid deep reinforcement learning network, which consists of two layers. The continuous layer deep reinforcement learning network is DDPG. The discrete layer deep reinforcement learning network is DQN.
[0094] Step 5.2) The DDPG network consists of the Actor current network, the Actor target network, the Critic current network and the Critic target network. The four network parameters are θ DDPG ,θ' DDPG ,ω DDPG and ω' DDPG The role of the Actor network is to output action decisions based on the input state, and the role of the Critic network is to estimate the value of the Actor network taking a certain action in a certain state - the Q value, and guide the action selection of the next state. The DQN network consists of the DQN current network and the DQN target network. The parameters of the two networks are ω DQN ,ω' DQN . Build a neural network, initialize DDPG network parameters, DQN network parameters, and experience pool capacity E DQN 、E TD3 .
[0095] Specifically, the method of inputting each time slot state into the hybrid deep reinforcement learning network to generate an action in step 6 includes:
[0096] The system state is input into the hybrid deep reinforcement learning network, and the DDPG Actor network generates the bandwidth allocation ratio a_c, and the DQN network generates the user grouping a_d. At this time, according to the channel allocation scheme, that is, the user grouping situation m nk , bandwidth allocation ratio τ nk , decomposing the maximization of system computational efficiency into maximizing the computational efficiency Y of each channel n :
[0097] The problem is transformed into
[0098]
[0099] The matrix X is initialized to a zero matrix at each time step. (x n,i ,x n,j ) has four possible values, namely (0, 0), (1, 0), (0, 1), and (1, 1). Substitute the four combinations into the above formula and select the one that makes Y n The largest combination resets the value of the corresponding position of X.
[0100] Specifically, the method of training the hybrid deep reinforcement learning network in step 7 includes:
[0101] When the base station is in state s, it performs action a = {a_c, a_d} and gets an immediate reward from the environment feedback. And get the state s' of the next time slot. Store (s, a_c, r, s') in the DDPG experience pool, and store the sample (s, a_d, r, s') in the DQN experience pool. The DDPG network and the DQN network share the state and reward value.
[0102] During the operation of the hybrid deep reinforcement learning network, it is necessary to continuously train the network and adjust the network parameters to improve the function fitting ability, so that the algorithm can always output reasonable action decisions in a dynamically changing state environment.
[0103] Adjust the parameters of the DQN current network by minimizing the cost function of the neural network. The cost function is as follows:
[0104]
[0105] Where D is the sample size taken from the experience pool.
[0106] After the DQN current network is updated a certain number of times, the weights of the DQN current network are copied to the DQN target network.
[0107] The update method of each parameter of the DDPG network is shown in the formula. The update formula of the current network parameter of Critic is as follows:
[0108]
[0109]
[0110] The update of the Actor's current network weight depends on the Q value of the Critic's current network. The Actor's current network updates its network parameters in the direction of obtaining greater cumulative rewards. The update formula of the Actor's current network parameters is as follows:
[0111]
[0112] Unlike the DQN algorithm, which directly copies the current network parameters of DQN to the target network parameters of DQN, DDPG uses a soft update method to update the target network parameters. The soft update formula is as follows:
[0113]
[0114] Here τ is generally taken as 0.001.
[0115] Step 8: Repeat steps 6 and 7 until the number of repetitions reaches the total number of time slots T, thereby stopping the algorithm.
[0116] In summary, the present invention establishes a NOMA-MEC system and proposes a new sub-channel allocation, computation offloading decision, and bandwidth allocation scheme based on hybrid deep reinforcement learning to maximize the long-term energy efficiency of the system.
[0117] It should be noted that the above-mentioned embodiments are only specific implementations of the present invention, but the protection scope of the present invention is not limited thereto. Any replacement, improvement, etc. based on the present invention should be included in the claims of the present invention.
[0118] Embodiment 2:
[0119] This embodiment provides a user grouping and resource allocation device in a NOMA-MEC system based on hybrid deep reinforcement learning, including the following steps:
[0120] System description module: used to describe the NOMA-MEC system;
[0121] Efficiency definition module: used to define the energy efficiency of the system;
[0122] Problem description module: used to describe the optimization problem;
[0123] Space definition module: used to define the state space and action space of deep reinforcement learning;
[0124] Network building module: used to build a hybrid deep reinforcement learning network; the input of the network is the state and the output is the action;
[0125] Action generation module: used to input each time slot state into the hybrid deep reinforcement learning network to generate actions;
[0126] Network training module: used to train hybrid deep reinforcement learning networks;
[0127] Output module: After the number of repeated training reaches the specified number of time slots T, the action generated at this time is output, that is, the decision to be optimized: user grouping, computing offloading, and bandwidth allocation ratio.
[0128] The device of this embodiment can be used to implement the method described in the first embodiment.
[0129] Embodiment three:
[0130] This embodiment provides a user grouping and resource allocation device in a NOMA-MEC system based on hybrid deep reinforcement learning, including a processor and a storage medium;
[0131] The storage medium is used to store instructions;
[0132] The processor is used to operate according to the instructions to execute the steps of the method described in embodiment 1.
[0133] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for user grouping and resource allocation in a NOMA-MEC system based on hybrid deep reinforcement learning, characterized in that, it includes the following steps: Step 1, Describe the NOMA-MEC system, which operates in a time-slot manner, and the time-slot set is denoted as Γ = {1, 2,..., T}; Step 2, Define the energy efficiency of the system; Step 3, Describe the optimization problem; Step 4, Define the state space of deep reinforcement learning and the action space of deep reinforcement learning; Step 5, Construct a hybrid deep reinforcement learning network; the input of the network is the state and the output is the action; Step 6, Input each time-slot state into the hybrid deep reinforcement learning network to generate an action; Step 7, Train the hybrid deep reinforcement learning network; Step 8, Repeat Step 6 and Step 7 until the number of repetitions reaches the specified number of time slots T, and then output the action generated at this time, that is, the decision to be optimized: user grouping, computing offloading, and bandwidth allocation ratio; The method for describing the NOMA-MEC system includes: The NOMA-MEC system consists of K user devices and a single-antenna base station connected to an edge server, and all users have only a single transmit antenna to establish a communication link with the base station; the system operates in a time-slot manner, and the time-slot set is denoted as Γ = {1, 2,..., T}; The total system bandwidth B is divided into N orthogonal sub-channels, and the bandwidth of sub-channel n accounts for a proportion τ of the total bandwidth. n , Define and to represent the set of users and the set of orthogonal sub-channels respectively, where K ≤ 2N; Divide the entire process into time slots, Γ = {1, 2,..., T}; the channel gain remains constant within the time period of one time slot and varies between different time slots. denotes the channel gain from user k to the base station on channel n, and let h n1 < h n2 <.... < h nK , n ∈ [1, N]; Limit a channel to allow at most two user signals to be transmitted simultaneously, and a user only sends a signal on one channel within a time slot; m nk = 1 indicates that channel n is allocated to user k for signal transmission, m nk = 0 indicates that channel n is not allocated to user k for signal transmission; The method for defining the energy efficiency of the system includes: Step 2.1) The energy efficiency Y of the system is defined as the sum of the calculation rates and calculation power ratios of all users, as shown in the following formula: Among them, R i,off represents the computing rate at which user i offloads the computing task to the edge server for execution, and p i is the transmission power of user i, which does not change with time, and the transmission powers of all users are the same; R i,local represents the computing rate at which user i executes the task locally, and p i,local represents the power consumed by user i for local execution, and x ni = 1 indicates that user i offloads the task to the edge server for execution through channel n, and x ni = 0 indicates that user i does not offload the task to the edge server for execution through the channel; Step 2.2) Since the channel gain h of user i on channel n ni is greater than the channel gain h of user j nj ; according to the successive interference cancellation technique, the base station decodes in descending order of the channel gains of users, then the offloading rate of user i is the offloading rate of user j where N 0 is the power spectral density of the noise; Step 2.3) The local execution computing rates of user i and user j are respectively where f i and f j are the CPU processing capabilities of the users, is the number of cycles required to process 1 bit of task; the local execution computing powers of user i and user j are respectively p i,local = νf i 3 and p j,local = νf j 3 , where ν is the capacitance effective coefficient of the user device chip architecture; The optimization problem is described as: The method for defining the state space and action space of deep reinforcement learning includes: Step 4.1) The state space s, s = {h 11 , h 12 ,... h 1K , h 21 , h 22 ,..., h 2K , h N1 ... h NK}; The action space a described in step 4.2 consists of two stages, a = {a_c, a_d}, where a_c = {τ 1 , τ 2 ,..., τ N} represents the system bandwidth allocation ratio of continuous actions, and a_d = {m 11 , m 12 ,..., m 1K ,..., m N1 , m N2 ,..., m NK , x 11 , x 12 ,..., x 1K ,..., x N1 , x N2 ,..., x NK} represents the discrete action representation sub-channel allocation scheme; The method for constructing a hybrid deep reinforcement learning network includes: The hybrid deep reinforcement network includes a continuous-layer deep reinforcement learning network and a discrete-layer deep reinforcement learning network; the continuous-layer deep reinforcement learning network is DDPG, and the discrete-layer deep reinforcement learning network is DQN; The method for inputting each time-slot state into the hybrid deep reinforcement learning network to generate an action includes: Step 6.1) Input the system state into the hybrid deep reinforcement learning network. The Actor network of DDPG generates the a_c bandwidth allocation ratio, and the DQN network generates the a_d user grouping situation; Step 6.2) In the case of household grouping m nk , bandwidth allocation ratio τ n After determination, the maximization of system energy efficiency is decomposed into the maximization of energy efficiency Y of each channel n ; The problem is transformed into where the matrix X is initialized as a zero matrix at each time step; (x n,i , x n,j ) has 4 possible values, namely (0, 0), (1, 0), (0, 1), and (1, 1). Among them, the value of x determines the offloading decision. 0 means not offloading the computing task of the user equipment to the edge server for execution, and 1 means offloading it to the edge server for execution. Substitute the 4 combinations into the above formula respectively, and select the combination that makes Y n the largest, and reset the value at the corresponding position of X; The method for training the hybrid deep reinforcement learning network includes: When the base station is in state s and executes action a = {a_c, a_d}, it obtains the immediate reward of the environmental feedback and obtains the state s' of the next time slot; Store (s, a_c, r, s') into the DDPG experience pool, and store the sample (s, a_d, r, s') into the DQN experience pool. The DDPG network and the DQN network share the state and reward value; The DDPG network and the DQN network sample D samples from the experience pool to train and update their own parameters.
2. A device for user grouping and resource allocation in a NOMA-MEC system based on hybrid deep reinforcement learning for executing the method as claimed in claim 1, characterized in that, it includes the following steps: System description module: used to describe the NOMA-MEC system; Efficiency definition module: used to define the energy efficiency of the system; Problem description module: used to describe the optimization problem; Space definition module: used to define the state space of deep reinforcement learning and the action space of deep reinforcement learning; Network construction module: used to construct a hybrid deep reinforcement learning network; the input of the network is the state and the output is the action; Action generation module: used to input the state of each time slot into the hybrid deep reinforcement learning network to generate actions; Network training module: used to train the hybrid deep reinforcement learning network; Output module: after the number of repeated training reaches the specified number of time slots T, output the actions generated at this time, that is, the decisions to be optimized: user grouping, computing offloading, and bandwidth allocation ratio.
3. A user grouping and resource allocation device in a NOMA-MEC system based on hybrid deep reinforcement learning, characterized in that, it includes a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in claim 1.
Citation Information
Patent Citations
Task unloading and bandwidth allocation processing method and system based on physical layer
CN112272390A