A method, system and program product for joint optimization of communication and computing resource allocation in a ris-assisted noma-mec network
By tightly coupling communication and computation in a RIS-assisted NOMA-MEC network, and by using the Lagrange duality algorithm and a deep reinforcement learning model to optimize resource allocation, the problem of system performance degradation caused by ignoring the computation process in existing technologies is solved, achieving energy minimization and performance improvement.
Patent Information
- Application Number
- CN202410995248.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-07-24
AI Technical Summary
Existing technologies in RIS-assisted NOMA-MEC networks neglect the tight coupling between computation and communication processes, making it difficult to obtain the optimal solution for the communication process and reducing the overall performance of the system.
By tightly coupling the communication and computation processes, the original model is decomposed into two nested sub-models. The Lagrange dual algorithm and deep reinforcement learning model are used to optimize the task offloading ratio and reflector phase shift, respectively. The transmission time of the NOMA group, user transmit power, server computing resource allocation and local computing frequency are jointly optimized to minimize system energy consumption.
This achieves long-term energy consumption minimization for both users and servers, improving the overall performance of the NOMA-MEC network.
Smart Images

Figure CN119012279B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of MEC network technology, and specifically to a joint optimization method, system, and program product for the allocation of communication and computing resources in a RIS-assisted NOMA-MEC network. Background Technology
[0002] Within the framework of Mobile Edge Computing (MEC), offloading tasks to MEC servers significantly reduces system latency and energy consumption. Non-orthogonal Multiple Access (NOMA) and Reconfigurable Intelligent Surface (RIS) technologies further improve spectrum utilization efficiency and wireless channel controllability, effectively enhancing the performance of MEC systems.
[0003] Optimizing resource allocation in RIS-assisted NOMA-MEC networks has sparked extensive discussion in academia and industry. Most current solutions only consider the communication process, neglecting the computation process. However, communication and computation are tightly coupled; ignoring computation makes it difficult to find optimal solutions for relevant variables in the communication process, thus reducing the overall system performance. Summary of the Invention
[0004] To address the problems existing in the prior art, the present invention aims to provide a joint optimization method, system, and program product for the allocation of communication and computing resources in a RIS-assisted NOMA-MEC network, which tightly couples the communication process with the computing process, thereby improving the overall performance of the NOMA-MEC network system.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network is disclosed. This method decomposes the original model into two nested sub-models: a task offloading ratio and communication process joint optimization sub-model, and a reflector phase shift optimization sub-model. For the task offloading ratio and communication process joint optimization sub-model, theoretical derivation proves that, given the reflector phase shift, the optimal transmit power is a function of the NOMA group transmission time and the task offloading ratio. Then, the Lagrange dual algorithm is used to solve for the optimal task offloading ratio and the optimal NOMA group transmission time. For the reflector phase shift optimization sub-model, it is expressed as a deep reinforcement learning model and solved using a near-end policy optimization algorithm to obtain the optimal reflector phase shift. Finally, the optimal reflector phase shift is substituted into the task offloading ratio and communication process joint optimization sub-model for further solving to obtain the optimal offloading ratio and the optimal transmission time, thereby obtaining the optimal transmit power, the optimal local computing frequency, and the optimal server computing resource allocation.
[0007] The method is applicable to RIS-assisted NOMA-MEC networks, where the MEC server is connected to the BS, and a K-element RIS assists N single-antenna users in offloading computation from the MEC server; the users form a set. A collection of reflective surface elements The time of the entire NOMA-MEC system is divided into equally spaced discrete time slots. Each time slot has a length of τ. In each time slot t, the user has a task that needs to be computed. The user adopts a partial offloading strategy, offloading some tasks to the edge server for computation, while the rest are computed locally.
[0008] The method specifically includes the following steps:
[0009] Step 1: Construct the original model;
[0010] User n's computational task in time slot t is composed of tuples. Characterization, where L n (t) represents the amount of task data for user n in time slot t, C n (t) represents the number of CPU cycles required for user n to compute each bit of the task in time slot t. This represents the maximum allowed task delay for user n in time slot t, and the task needs to... The calculation is completed within a time slot; all users' tasks can always be completed within a single time slot.
[0011] Let λ be the proportion of data unloaded by user n in time slot t to the total task data volume. n (t), where the local computation frequency of user n in time slot t is The local computation latency at this time Represented as:
[0012]
[0013] Locally calculated energy consumption Represented as:
[0014]
[0015] Wherein, κ1 is the energy consumption coefficient;
[0016] Users' computing power is not unlimited; there are computing resource constraints, therefore:
[0017]
[0018] Furthermore, the total latency of local computation must meet the maximum latency constraint of the user task, that is:
[0019]
[0020] use To represent a complex space with x×y dimensions, use This represents the direct channel between user n and the server in time slot t, using... This represents the channel from user n in time slot t to RIS, using This represents the channel from RIS to the server in time slot t; θ is used. k (t)∈[0,2π] represents the phase shift of the k-th element in time slot t. Let g represent the phase shift matrix of RIS in time slot t. Then, the equivalent channel gain g between user n and the server in time slot t is... n (t) is represented as:
[0021]
[0022] All users in a NOMA group use the NOMA protocol for offloading. Let P be the transmit power of user n in time slot t. n (t), the system bandwidth is B, and the system noise is σ. 2 Then the offloading rate of user n in time slot t is expressed as:
[0023]
[0024] Let the transmission time of the NOMA group be d(t), then the transmission energy consumption of user n in time slot t is... Represented as:
[0025]
[0026] When using the NOMA protocol for uninstallation, it should be ensured that all users complete the uninstallation process. Therefore, the following constraints must be met:
[0027]
[0028] At the same time, users are subject to a maximum transmission power limit, namely:
[0029]
[0030] Suppose that the server computing resources allocated to user n in time slot t are... The following constraints must be satisfied:
[0031]
[0032] in Let be the total computing resources that the server can provide; the server computation latency for user n in time slot t is expressed as:
[0033]
[0034] The computational energy consumption of the server in performing calculations for the offloading task of user n in time slot t is:
[0035]
[0036] Where κ2 is the energy consumption coefficient related to server hardware;
[0037] The total latency of offloading computation includes the transmission latency of the NOMA group and the computation latency of the server. It needs to meet the maximum latency constraint of the user task, therefore:
[0038]
[0039] Based on the above, the optimization objective for each time slot is expressed as:
[0040]
[0041] The original model is then represented as
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] The goal of the optimization is to jointly optimize the transmission time of the NOMA group, local computing frequency, server computing resource allocation, user transmit power, task offloading ratio, and reflector phase shift in order to minimize the long-term energy consumption of the system.
[0053] Step 2, take the original Transformation;
[0054] Step 2.1: Replace the locally calculated frequency variable;
[0055] Lemma 1: The optimal solution for frequency locally calculated by user n in time slot t Always satisfied:
[0056]
[0057] And there are:
[0058]
[0059] Substituting the result of Lemma 1 into the original model Obtain the model
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069] Step 2.2: Replace the server's frequency calculation variable;
[0070] Lemma 2: Model Medium constraints The less than or equal to sign can be replaced by an equal sign without changing the optimal solution of the model; that is:
[0071]
[0072] According to Lemma 2, the server allocated to user n in time slot t calculates the frequency. Represented as:
[0073]
[0074] Will Substituting the expression, we get:
[0075]
[0076] At the same time, in order to ensure that the allocated server computing frequency is positive, a new constraint is introduced:
[0077]
[0078] Substitute the above results into the model The model can be further obtained.
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089] Step 3: Optimize other variables given the phase shift of the reflecting surface;
[0090] Step 3.1: Solve the resource allocation optimization model for each time slot when the phase shift of the reflecting surface is given. Consider each time slot Given the phase shift Θ of the reflecting surface t Energy consumption minimization model under the condition of
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] Step 3.2: Replace the transmit power optimization variable;
[0100] Lemma 3: Model Constraints Replacing the less than or equal to sign with an equal sign does not change the optimal solution of the model; that is:
[0101]
[0102] When the phase shift of the reflecting surface in time slot t is given, the user's channel gain is also determined; let s be the user whose channel gain ranks qth in time slot t. q (Θ t If the channel gain is arranged in ascending order, then the users are as follows:
[0103]
[0104] According to Lemma 3, we obtain the following based on different user serial numbers:
[0105]
[0106] Therefore, the transmission power is calculated using the following formula:
[0107]
[0108] From the above equation, we can see that when 2≤q≤N, The signal strength is always related to the transmit power of users with worse channel quality; in this case, the signal from users with lower channel gain is considered interference. When q ≥ 2, Multiply both sides of the expression by get:
[0109]
[0110] After rearranging the items, we have:
[0111]
[0112] Therefore, when q≥2 Calculated by continuous multiplication:
[0113]
[0114] Substituting the above derivation into... In the expression, when q≥2:
[0115]
[0116] Therefore, user s q (Θ t The transmit power of ) is expressed as d(t). The function, denoted as
[0117] In summary, user s q (Θ t The transmit power of ) is expressed as:
[0118]
[0119] Through the above transformation, each time slot The model below is reformulated as a model That is, a joint optimization sub-model of task offloading ratio and communication process:
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127] d(t)≥0,
[0128] Step 3.3: Write out the model Let the Lagrangian function be:
[0129]
[0130]
[0131] Model The Lagrange function is expressed as:
[0132]
[0133] in,
[0134]
[0135] use Representation Model The set of Lagrange multipliers corresponding to each constraint in the equation. Therefore, the dual function of the original model is expressed as:
[0136]
[0137] The goal of the dual model is to maximize the dual function, i.e.:
[0138]
[0139] The constraint condition for the dual model is that all Lagrange multipliers are greater than or equal to 0;
[0140] Step 3.4: Find the optimal solution given the Lagrange multipliers. At this point, d(t) and λ(t) are unconstrained with respect to the Lagrange function. Use the coordinate descent algorithm to solve for the given Lagrange multipliers. Optimal under the condition and optimal Each variable is updated in a certain order, ensuring that other variables are up-to-date when a variable is updated. Let i be the number of iterations of the coordinate descent method. The coordinate descent algorithm is expressed as:
[0141]
[0142] ...
[0144]
[0145] Repeat the above process until convergence or the maximum number of iterations is reached;
[0146] Step 3.5: Update the Lagrange multipliers;
[0147] get and Next, calculate the user's transmit power, denoted as... Then, update all Lagrange multipliers along the gradient direction of the Lagrange function with respect to the Lagrange multipliers, that is, for each And ψ5, whose update directions are as follows:
[0148]
[0149]
[0150]
[0151]
[0152] and
[0153]
[0154] Let the update direction be ρ is the step size for each update, and the new La Langerange multipliers are updated as follows:
[0155]
[0156] Step 3.6: Repeat steps 3.3-3.5 until the maximum number of iterations is reached;
[0157] Step 4: Represent the phase shift optimization model of the reflector as a deep reinforcement learning model;
[0158] Step 4.1: Define the state space and action space; agents obtain better cumulative rewards based on states, so the state is set as the channel information of the NOMA-MEC system; let... Given the set of all channel gains within time slot t, the state is set as follows:
[0159]
[0160] The goal is to optimize the phase shift of the reflecting surface, therefore the motion space is set as follows:
[0161] a t =[θ1(t),θ2(t),...,θ k (t)],
[0162] Step 4.2: Define the reward. Based on the reward, the agent updates the policy and establishes a mapping from state to action. The goal is to minimize system energy consumption, therefore the reward for time slot t is set as follows:
[0163] r t =exp(-E total (t))+χ(t),
[0164] Where χ(t) represents the penalty, when at Make the model When there is no solution, χ(t) = -1, and E total The value of χ(t) is set to +∞, otherwise χ(t) = 0; in this way, the AI will gradually discard these invalid actions, and the lower system energy consumption can bring higher rewards. By continuously training, maximizing long-term rewards can minimize long-term total energy consumption.
[0165] Step 4.3: Solve for the phase shift of the reflector surface using the near-end strategy optimization algorithm;
[0166] Randomly initialize the parameters φ of the Actor network A With Critic network parameter φ c Clear the experience pool; in each round, based on each state s t Intelligent agents utilize strategies Choose action a t Based on action a t The RIS-assisted NOMA-MEC environment provides feedback rewards to the agent. t And proceed to the next state, and transfer the experience <s t ,a t ,r t ,s t > Add to the experience pool; after each round, perform Γ updates on the PPO agent, using the Adam optimizer to update the Actor network parameters φ. A and Critic network parameters φ c Repeat this process multiple times until the maximum number of training rounds Ω is reached to obtain the phase shift of the reflector surface.
[0167] Step 5: After training is complete, near-optimal phase shifts of the reflector can be generated in real time, and then the model from Step 3 can be further used. The solution method yields the optimal offloading ratio and optimal transmission time, which in turn leads to the optimal transmission power, optimal local computing frequency, and optimal allocation of server computing resources.
[0168] A joint optimization system for communication and computing resource allocation in a RIS-assisted NOMA-MEC network includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network as described above.
[0169] A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of a joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network as described above.
[0170] By adopting the above scheme, this invention theoretically proves the relationship between local computing frequency, server computing frequency, and other optimization variables when the original model achieves its optimum, reducing the number of variables that need to be directly optimized. Then, the problem is decomposed into two nested sub-models: a joint optimization sub-model of task offloading ratio and communication process, and a reflector phase shift optimization sub-model. These are solved using the Lagrange dual algorithm and the PPO algorithm, respectively, ultimately minimizing the long-term energy consumption of users and servers. Thus, this invention tightly couples the communication process with the computing process, improving the overall performance of the NOMA-MEC network system. Attached Figure Description
[0171] Figure 1 This is a scenario diagram illustrating the applicability of the present invention.
[0172] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0173] like Figure 1 As shown, this invention is applicable to RIS-assisted NOMA-MEC networks, where the MEC server is connected to the BS. A K-element RIS assists N single-antenna users in offloading computations to the MEC server. It is assumed that direct communication between users and the BS is not possible, possibly due to service dead zones or obstructions. Therefore, tasks can only be offloaded to the BS with the assistance of the RIS, using the NOMA protocol for uplink transmission. The user set... A collection of reflective surface elements The time of the entire NOMA-MEC system is divided into equally spaced discrete time slots. Each time slot has a length of τ. In each time slot t, the user has a task that needs to be computed. The user adopts a partial offloading strategy, offloading part of the task to the edge server for computation, while the rest is computed locally.
[0174] The joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network of the present invention specifically includes the following steps:
[0175] Step 1: Construct the original model.
[0176] User n's computational task in time slot t is composed of tuples. Characterization, where L n (t) represents the amount of task data for user n in time slot t, in bits. n (t) represents the number of CPU cycles required for user n to compute each bit of task in time slot t, in units of cycle / bit. This represents the maximum allowed task latency for user n in time slot t, in milliseconds. The task needs to... The calculation is completed within a time slot. Consider the scenario where all users' tasks can always be completed within one time slot.
[0177] Let λ be the proportion of data unloaded by user n in time slot t to the total task data volume. n (t), which employs dynamic voltage frequency scaling technology, assuming that the local calculation frequency of user n in time slot t is... The local computation latency at this time It can be represented as:
[0178]
[0179] Locally calculated energy consumption It can be represented as:
[0180]
[0181] Where κ1 is the energy consumption coefficient, which is related to the user's hardware. The user's computing power is not unlimited; there are computing resource constraints, therefore:
[0182]
[0183] Furthermore, the total latency of local computation must meet the maximum latency constraint of the user task, that is:
[0184]
[0185] use To represent a complex space with x×y dimensions, use This represents the direct channel between user n and the server in time slot t, using... This represents the channel from user n in time slot t to RIS, using This represents the channel from the RIS to the server in time slot t. By adjusting the phase shift of the RIS, the equivalent channel between the user and the server can be dynamically adjusted. Using θ... k (t)∈[0,2π] represents the phase shift of the k-th element in time slot t. Let g represent the phase shift matrix of RIS in time slot t. Then, the equivalent channel gain g between user n and the server in time slot t is... n (t) can be represented as:
[0186]
[0187] In the scenario of this invention, all users in a NOMA group use the NOMA protocol for offloading. Let the transmit power of user n in time slot t be P. n (t), the system bandwidth is B, and the system noise is σ. 2 Then the offloading rate of user n in time slot t can be expressed as:
[0188]
[0189] Let the transmission time of the NOMA group be d(t), then the transmission energy consumption of user n in time slot t is... It can be represented as:
[0190]
[0191] When using the NOMA protocol for uninstallation, the uninstallation needs of all users must be considered, meaning that the uninstallation of all users should be guaranteed to be completed. Therefore, the following constraints must be satisfied:
[0192]
[0193] At the same time, the user's transmission power is not unlimited; there is a maximum transmission power limit, namely:
[0194]
[0195] For data offloaded to the server, the MEC server allocates computing resources to different users to complete the computation. Compared to cloud servers, the computing resources of the MEC server are very limited. Let's assume that the server computing resources allocated to user n in time slot t are... The following constraints must be satisfied:
[0196]
[0197] in This represents the total computing resources that the server can provide. Similarly, the server computation latency for user n in time slot t is expressed as:
[0198]
[0199] The computational energy consumption of the server in performing calculations for the offloading task of user n in time slot t is:
[0200]
[0201] Where κ2 is the energy consumption coefficient related to server hardware. Since the calculated result is usually much smaller than the input data, the energy consumption and latency of downloading from the MEC to the user are ignored. The total latency of offloading the computation includes the transmission latency of the NOMA group and the server computation latency, which needs to meet the maximum latency constraint of the user task; therefore:
[0202]
[0203] Based on the above discussion, the optimization objective for each time slot can be expressed as:
[0204]
[0205] In summary, the model is represented as follows:
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214]
[0215]
[0216] The optimization objective is to jointly optimize the NOMA group's transmission time, local computing frequency, server computing resource allocation, user transmit power, task offloading ratio, and reflector phase shift to minimize the system's long-term energy consumption. Constraints This indicates that the user's local computing frequency is limited, constraining... This indicates that the local computation latency cannot exceed the task's maximum allowable latency, a constraint. This indicates that when using the NOMA protocol for transmission, it is necessary to ensure that each user completes the offloading process, which is a constraint. This indicates that the user's transmit power is limited, constrained. This indicates that the server's computing resources are limited and constrained. This indicates that the unloading delay cannot exceed the task's maximum allowed delay.
[0217] Step 2: Transform the original model.
[0218] Step 2.1: Replace the locally calculated frequency variable.
[0219] Lemma 1. The optimal solution for frequency locally calculated by user n in time slot t. Always satisfied:
[0220]
[0221] And there are:
[0222]
[0223] Proof. By constraints and The following inequality can be obtained:
[0224]
[0225] If the original model has a solution, then the lower bound in the above formula must be less than or equal to the upper bound, that is... It holds true, and because the objective function changes with... Since it is increasing, the optimal solution for locally calculated frequency must be... The lower bound, that is Q.E.D.
[0226] Substitute the result of Lemma 1 into the model The model can be obtained
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236] Step 1.2: Replace the server's frequency calculation variable.
[0237] Lemma 2. Model Medium constraints The less than or equal to sign can be replaced with an equal sign without changing the optimal solution of the original model. That is:
[0238]
[0239] Proof. Assume that there exists a user n in time slot t, and the corresponding optimal solution d(t). # , λ n (t) # and For the inequality to hold strictly in the constraints, that is:
[0240]
[0241] So you can always find Make constraints The equality holds in the equation, provided that the condition is met. Related constraints While this holds true, it also reduces the total energy consumption of the system. This contradicts the assumption of the optimal solution; therefore, when obtaining the optimal solution, constraints... The intermediate form is strictly true. Q.E.D.
[0242] By Lemma 2, the server allocated to user n in time slot t calculates the frequency. It can be represented as:
[0243]
[0244] Will Substituting the expression, we get:
[0245]
[0246] At the same time, in order to ensure that the allocated server computing frequency is positive, new constraints need to be introduced:
[0247]
[0248] Substitute the above results into the model The model can be further obtained.
[0249]
[0250]
[0251]
[0252]
[0253]
[0254]
[0255]
[0256]
[0257]
[0258]
[0259] Step 3: Optimize other variables when the phase shift of the reflecting surface is given.
[0260] Step 3.1: Solve the resource allocation optimization model for each time slot when the phase shift of the reflecting surface is given.
[0261] Consider each time slot Given the phase shift Θ of the reflecting surface t The energy consumption minimization model under the given conditions is expressed as follows:
[0262]
[0263]
[0264]
[0265]
[0266]
[0267]
[0268]
[0269]
[0270]
[0271] Step 3.2: Replace the transmit power optimization variable
[0272] Lemma 3. Model Constraints The less than or equal to sign can be replaced with an equal sign without changing the optimal solution of the original model. That is:
[0273]
[0274] Proof. Suppose that for a certain user n, there is a corresponding optimal solution d(t). # , λ n (t) # and P n (t) # This makes the constraint The inequality sign holds strictly, that is:
[0275]
[0276] d(t) can always be found. * <d n (t) # This makes the constraint The equality holds. At this point, the constraints related to d(t) include:
[0277]
[0278] therefore Established; and because Therefore:
[0279]
[0280] Therefore, constraints This means that all constraints related to d(t) are valid. At this point, local computation energy consumption remains unchanged, while transmission energy consumption and server computation energy consumption are both lower, resulting in a lower total system energy consumption. This contradicts the assumption of an optimal solution. Therefore, when the optimal solution is obtained, the constraints... The intermediate form is strictly true. Q.E.D.
[0281] When the phase shift of the reflecting surface in time slot t is given, the user's channel gain is also determined. Let s be the user whose channel gain ranks qth in time slot t. q (Θ t If the channel gain is arranged in ascending order, then the users are as follows:
[0282]
[0283] According to Lemma 3, we can obtain the following based on the different user serial numbers:
[0284]
[0285] Therefore, the transmission power can be calculated using the following formula:
[0286]
[0287] From the above equation, we can see that when 2≤q≤N, The transmit power of users with worse channel quality is always related to the channel power output. This is because a channel gain-based SIC decoding order is used, prioritizing the decoding of signals with higher channel gain. In this case, user signals with lower channel gain are considered interference. Further derivation follows. When q ≥ 2, Multiply both sides of the expression by We can obtain:
[0288]
[0289] After rearranging the items, we have:
[0290]
[0291] Therefore, when q≥2 It can be calculated by continuous multiplication:
[0292]
[0293] Substituting the above derivation into... From the expression, we can obtain the condition when q≥2:
[0294]
[0295] Therefore, user s q (Θ t The transmit power of ) can be expressed as d(t), The function, denoted as In summary, user s q (Θ t The transmit power of ) can be expressed as:
[0296]
[0297] Through the above transformation, each time slot The model below is reformulated as a model
[0298]
[0299]
[0300]
[0301]
[0302]
[0303]
[0304]
[0305] d(t)≥0.
[0306] Step 2.3: Write out the model Let the Lagrangian function be:
[0307]
[0308]
[0309] Model The Lagrange function can be expressed as:
[0310]
[0311] in,
[0312]
[0313] use Representation Model The set of Lagrange multipliers corresponding to each constraint in the equation. Therefore, the dual function of the original model can be expressed as:
[0314]
[0315] The goal of the dual model is to maximize the dual function, i.e.:
[0316]
[0317] The dual model constraint is that all Lagrange multipliers are greater than or equal to 0.
[0318] Step 3.4: Find the optimal solution given the Lagrange multipliers. At this point, d(t) and λ(t) are unconstrained with respect to the Lagrange function. The coordinate descent algorithm is used to solve for the given Lagrange multipliers. Optimal under the condition and optimal Each variable is updated in a certain order, ensuring that other variables are up-to-date when a variable is updated. Let i be the number of iterations of the coordinate descent method. The coordinate descent algorithm can be expressed as:
[0319]
[0320]
[0321] ...,
[0322]
[0323] Repeat the above process until convergence or the maximum number of iterations is reached.
[0324] Step 3.5: Update the Lagrange multipliers. (The result is...) and Then, the user's transmit power can be calculated, denoted as... Then, update all Lagrange multipliers along the gradient direction of the Lagrange function with respect to the Lagrange multipliers, that is, for each And ψ5, whose update directions are as follows:
[0325]
[0326]
[0327]
[0328]
[0329] and
[0330]
[0331] Let the update direction be ρ is the step size for each update, and the new La Langerange multipliers are updated as follows:
[0332]
[0333] Step 3.6: Repeat steps 3.3-3.5 until the maximum number of iterations is reached.
[0334] Step 4: Express the phase shift optimization model of the reflector as a deep reinforcement learning model.
[0335] Step 4.1: Define the state space and action space. Intelligent agents obtain better cumulative rewards based on states; therefore, the state is set as the channel information of the NOMA-MEC system. Let... Given the set of all channel gains within time slot t, the state is set as follows:
[0336]
[0337] The goal is to optimize the phase shift of the reflecting surface, therefore the motion space is set as follows:
[0338] a t =[θ1(t),θ2(t),...,θ k (t)].
[0339] Step 4.2: Define the reward, which reflects which behaviors are beneficial to the system. Based on the reward, the agent updates the policy, establishing a mapping from state to action. The goal is to minimize system energy consumption; therefore, the reward for time slot t is set as follows:
[0340] r t =exp(-E total (t))+χ(t),
[0341] Where χ(t) represents the penalty, when a t Make the model When there is no solution, χ(t) = -1, and E total The value of χ(t) is set to +∞, otherwise χ(t) = 0. In this way, the AI will gradually discard these ineffective actions, and the lower system energy consumption will bring higher rewards. Through continuous training, maximizing long-term rewards can minimize long-term total energy consumption.
[0342] Step 4.3: Use the Proximal Policy Optimization (PPO) algorithm to solve for the phase shift of the reflector.
[0343] Randomly initialize the parameters φ of the Actor network A With Critic network parameter φ c Clear the experience pool. In each round, based on each state s... t Intelligent agents utilize strategies Choose action a t Based on action a t The RIS-assisted NOMA-MEC environment provides feedback rewards to the agent. t And proceed to the next state, and transfer the experience <s t ,a t ,r t ,s t > Add to the experience pool. After each round, perform Γ updates on the PPO agent, using the Adam optimizer to update the Actor network parameters φ. A and Critic network parameters φ c Repeat this process multiple times until the maximum number of training rounds Ω is reached.
[0344] Step 5: After training is completed, the near-optimal reflector phase shift can be generated in real time. Then, the solution method of the model in Step 3 is used to obtain the optimal unloading ratio and the optimal transmission time, thereby obtaining the optimal transmission power, the optimal local computing frequency, and the optimal allocation of server computing resources.
[0345] The key to this invention lies in its decomposition of the model into two nested sub-models: a joint optimization sub-model of task offloading ratio and communication process, and a reflector phase shift optimization sub-model. For the former, theoretical derivation proves that, given the reflector phase shift, the optimal transmit power can be expressed as a function of the NOMA group transmission time and the task offloading ratio. Then, the optimal task offloading ratio and optimal NOMA group transmission time are solved using the Lagrange dual algorithm. For the latter, it is expressed as a deep reinforcement learning model and solved using the PPO algorithm. Thus, this invention tightly couples the communication process with the computation process, improving the overall performance of the NOMA-MEC network system.
[0346] The present invention also provides a joint optimization system for communication and computing resource allocation in a RIS-assisted NOMA-MEC network, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network as described above.
[0347] The present invention also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist alone and not assembled into the electronic device.
[0348] The aforementioned computer-readable medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the methods described in the above embodiments.
[0349] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0350] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network, characterized in that: The method is applicable to RIS-assisted NOMA-MEC networks, where the MEC server is connected to the BS, one RIS assistance of components The computational offloading from individual antenna users to the MEC server; user groups A collection of reflective surface elements The time of the entire NOMA-MEC system is divided into equally spaced discrete time slots. Each time slot is [length missing] ; in each time slot Each user has a task that needs to be computed. The user adopts a partial offloading strategy, offloading part of the task to the edge server for computation, while the rest is computed locally. The method specifically includes the following steps: Step 1: Construct the original model; user In the time slot The computational task is performed by tuples Characterization, in which Indicates user In the time slot The amount of task data, Indicates user In the time slot Calculate the number of CPU cycles required per bit of task. Indicates user In the time slot The maximum allowed task delay, the task needs to be The calculation is completed within a time slot; all users' tasks can always be completed within a single time slot. Set user In the time slot The proportion of unloaded data to the total task data is ,user In the time slot The local computing frequency is Then the local computation latency at this time Represented as: Locally calculated energy consumption Represented as: in, Energy consumption coefficient; Users' computing power is not unlimited; there are computing resource constraints, therefore: Furthermore, the total latency of local computation must meet the maximum latency constraint of the user task, that is: use express Complex space of dimensions, using Indicates user In the time slot Direct connection channel with server, using Indicates user In the time slot The channel to RIS uses Indicates in time slot The channel from the RIS to the server; using Indicates the first Each component in the time slot phase shift, Indicates RIS in time slot The phase shift matrix, then the user In the time slot Equivalent channel gain between the server and the server Represented as: All users in a NOMA group use the NOMA protocol for offloading, configured in a time slot. user The transmission power is The system bandwidth is The system noise is Then the user In the time slot The unloading rate is expressed as: Let the transmission time of the NOMA group be... Then the user In the time slot Transmission energy consumption Represented as: When using the NOMA protocol for uninstallation, it should be ensured that all users complete the uninstallation process. Therefore, the following constraints must be met: At the same time, users are subject to a maximum transmission power limit, namely: Set in time slot For users The allocated server computing resources are Then the following constraints must be satisfied: in The total computing resources that the server can provide; users In the time slot The server computation latency in the data is expressed as: The server is in the time slot For users The computational energy consumption generated by performing the unloading task is: in The energy consumption coefficient related to server hardware; The total latency of offloading computation includes the transmission latency of the NOMA group and the computation latency of the server. It needs to meet the maximum latency constraint of the user task, therefore: Based on the above, the optimization objective for each time slot is expressed as: The original model is then represented as : The goal of the optimization is to jointly optimize the transmission time of the NOMA group, local computing frequency, server computing resource allocation, user transmit power, task offloading ratio, and reflector phase shift in order to minimize the long-term energy consumption of the system. Step 2, take the original Transformation; Step 2.1: Replace the locally calculated frequency variable; Lemma 1: User In the time slot The optimal solution for local computation frequency Always satisfied: And there are: Substituting the result of Lemma 1 into the original model To obtain the model : Step 2.2: Replace the server's frequency calculation variable; Lemma 2: Model Medium constraints The less than or equal to sign can be replaced by an equal sign without changing the optimal solution of the model; that is: According to Lemma 2, in the time slot For users Allocated server computing frequency Represented as: Will Substituting the expression, we get: At the same time, in order to ensure that the allocated server computing frequency is positive, a new constraint is introduced: Substitute the above results into the model The model can be further obtained. : Step 3: Optimize other variables given the phase shift of the reflecting surface; Step 3.1: Solve the resource allocation optimization model for each time slot when the phase shift of the reflecting surface is given. ; Consider each time slot Given the phase shift of the reflecting surface Energy consumption minimization model under the condition of : Step 3.2: Replace the transmit power optimization variable; Lemma 3: Model Constraints Replacing the less than or equal to sign with an equal sign does not change the optimal solution of the model; that is: And when the gap Given a phase shift of the reflector, the user's channel gain is also determined; the time slot... The mid-channel gain is ranked as the top The user is recorded as 1. Then, the users are arranged in ascending order of channel gain as follows: According to Lemma 3, we obtain the following based on different user serial numbers: Therefore, the transmission power is calculated using the following formula: From the above equation, we can obtain that when hour, The signal strength is always related to the transmit power of users with worse channel quality; in this case, the signal from users with lower channel gain is considered interference. hour, Multiply both sides of the expression by get: , After rearranging the items, we have: , Therefore when hour, Calculated by continuous multiplication: Substituting the above derivation into... In the expression, we get hour: Therefore, users The transmit power is expressed as The function, denoted as ; In summary, users The transmit power is expressed as: Through the above transformation, each time slot The model below is reformulated as a model , That is, a joint optimization sub-model of task unloading ratio and communication process: Step 3.3: Write out the model Let the Lagrangian function be: Model The Lagrange function is expressed as: in, use Representation Model The set of Lagrange multipliers corresponding to each constraint in the equation. Therefore, the dual function of the original model is expressed as: The goal of the dual model is to maximize the dual function, i.e.: The constraint condition for the dual model is that all Lagrange multipliers are greater than or equal to 0; Step 3.4: Find the optimal solution given the Lagrange multipliers; at this point... and The Lagrange function is unconstrained; the coordinate descent algorithm is used to solve for the given Lagrange multipliers. Optimal under the condition and optimal Update each variable in a certain order while ensuring that other variables are up-to-date when updating a variable. Let... Let be the number of iterations for the coordinate descent method. The coordinate descent algorithm is expressed as: Repeat the above process until convergence or the maximum number of iterations is reached; Step 3.5: Update the Lagrange multipliers; get and Next, calculate the user's transmit power, denoted as... Then, update all Lagrange multipliers along the gradient direction of the Lagrange function with respect to the Lagrange multipliers, that is, for each , , , as well as Their update directions are as follows: and Let the update direction be , For each update step, update the new La Langerange multipliers as follows: Step 3.6: Repeat steps 3.3-3.5 until the maximum number of iterations is reached; Step 4: Represent the phase shift optimization model of the reflector as a deep reinforcement learning model; Step 4.1: Define the state space and action space; agents obtain better cumulative rewards based on states, so the state is set as the channel information of the NOMA-MEC system; let... For time slots The set of all channel gains within the range is then set to the following state: The goal is to optimize the phase shift of the reflecting surface, therefore the motion space is set as follows: Step 4.2: Define the reward. Based on the reward, the agent updates the policy and establishes a mapping from state to action. The goal is to minimize system energy consumption, therefore time slots are... The reward is set as follows: in, Indicates punishment, when Make the model When there is no solution and will The value is set to ,otherwise In this way, the AI will gradually discard these ineffective actions, and the lower system energy consumption will bring higher rewards. By continuously training, maximizing long-term rewards can minimize long-term total energy consumption. Step 4.3: Solve for the phase shift of the reflector surface using the near-end strategy optimization algorithm; Randomly initialize the parameters of the Actor network. With Critic network parameters Clear the experience pool; in each round, based on each state Intelligent agents utilize strategies Choose an action Based on action The RIS-assisted NOMA-MEC environment provides rewards to the agent. And proceed to the next state, and gain experience. Add to the experience pool; after each round, perform a process on the PPO agent. This update uses the Adam optimizer to update the Actor network parameters. and Critic network parameters Repeat this process multiple times until the maximum number of training rounds is reached. This yields the phase shift of the reflecting surface; Step 5: After training is complete, near-optimal phase shifts of the reflector can be generated in real time, and then the model from Step 3 can be further used. The solution method yields the optimal offloading ratio and optimal transmission time, which in turn leads to the optimal transmission power, optimal local computing frequency, and optimal allocation of server computing resources.
2. A joint optimization system for communication and computing resource allocation in a RIS-assisted NOMA-MEC network, characterized in that, It includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network as described in claim 1.
3. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the joint optimization method for communication and computing resource allocation in a RIS-assisted NOMA-MEC network as described in claim 1.
Citation Information
Patent Citations
SWIPT-assisted NOMA-MEC system resource allocation modeling method
CN114521023A
Mobile edge calculation safety rate maximization method based on TD3 reinforcement learning and intelligent reflecting surface
CN117979314A