Airborne star-ris empowered mobile edge computing system energy consumption optimization method

By constructing an airborne STAR-RIS-enabled MEC system and utilizing deep reinforcement learning to optimize energy consumption, the problem of high system energy consumption was solved, resulting in a significant reduction in energy consumption and an improvement in communication quality.

CN120186682BActive Publication Date: 2026-03-31HEBEI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing research lacks energy consumption optimization methods for STAR-RIS-enabled mobile edge computing systems, making it difficult to effectively reduce system energy consumption, especially in urban environments with dense high-rise buildings.

Method used

By constructing an airborne STAR-RIS-enabled MEC system and combining it with deep reinforcement learning (DRL) technology, the trajectory, task offloading ratio, amplitude coefficients, phase shift coefficients, and power allocation of the airborne STAR-RIS are optimized. The deep deterministic policy gradient method (DDPG) is used to minimize energy consumption.

Benefits of technology

It significantly reduces the energy consumption of the airborne STAR-RIS-enabled MEC system and improves the system's energy efficiency and communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186682B_ABST
    Figure CN120186682B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of aerial STAR-RIS empowerment mobile edge computing system energy consumption optimization method, comprising the following steps: S1, constructs an aerial STAR-RIS empowerment MEC system;S2, by solving the flight energy consumption of each time slot of aerial STAR-RIS and the local computing energy consumption of IoT equipment k each time slot and the task offloading energy consumption of IoT equipment k in a flight cycle, determine the energy consumption of aerial STAR-RIS empowerment MEC system;S3, determine the constraint condition of minimizing the energy consumption of aerial STAR-RIS empowerment MEC system;S4, according to constraint condition, the energy consumption of STAR-RIS empowerment MEC system is minimized processing.The present application is integrated by UAV and STAR-RIS, auxiliary establishment communication link, promotes the communication between IoT equipment and MEC server;By jointly optimizing the trajectory of aerial STAR-RIS, task offloading proportion, transmission and reflection mode amplitude coefficient, transmission and reflection mode phase shift coefficient and power allocation, the low energy consumption target of aerial STAR-RIS empowerment MEC system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a wireless communication method, specifically an energy consumption optimization method for an airborne STAR-RIS-enabled mobile edge computing system. Background Technology

[0002] With the rapid development of Internet of Things (IoT) technology, its applications have widely penetrated into various fields such as smart homes, wearable devices, medical devices, and home appliances. However, the promotion of this technology faces many challenges, such as network latency and insufficient computing resources for low-power devices. To address these challenges, Mobile Edge Computing (MEC) technology has emerged, allowing IoT devices to offload some computing tasks to edge servers, thereby achieving efficient localized processing at the network edge. The introduction of MEC technology provides adaptive and reliable service support for fields such as smart cities, industrial automation, and healthcare, effectively meeting the stringent requirements of these fields for low latency and high-performance computing. Furthermore, Reconfigurable Smart Metasurfaces (RIS) technology can significantly improve signal coverage and quality by intelligently controlling the propagation path of electromagnetic waves, thus effectively coping with complex and ever-changing communication environments. The STAR-RIS, a novel architecture of RIS, simultaneously transmits and reflects incident signals, providing services to IoT devices located on both sides of the STAR-RIS simultaneously, achieving 360° omnidirectional coverage and higher degrees of freedom (DoF) for signal propagation, introducing flexibility into MEC. Nevertheless, establishing stable line-of-sight (LoS) connections remains a significant challenge in densely populated urban environments with tall buildings. To address this issue, unmanned aerial vehicle (UAV) technology has been introduced into the IoT field. With its rapid deployment capabilities and high flexibility, UAVs offer a new solution for IoT connectivity in urban environments, significantly improving network coverage and service quality.

[0003] Airborne STAR-RIS is a technology that utilizes UAVs equipped with STAR-RIS. Airborne STAR-RIS-enabled MECs offer a range of advantages, including significantly improved signal coverage, enhanced communication between IoT devices and MECs, and solutions to insufficient computing resources. Based on these advantages, researchers have begun studying airborne STAR-RIS-enabled MEC systems. However, existing research lacks consideration of the energy consumption constraints of such systems. To date, researchers have not designed low-energy optimization methods for airborne STAR-RIS-enabled MEC systems. Summary of the Invention

[0004] The purpose of this invention is to provide an energy consumption optimization method for airborne STAR-RIS-enabled mobile edge computing systems, so as to reduce the energy consumption of airborne STAR-RIS-enabled mobile edge computing systems.

[0005] The objective of this invention is achieved as follows:

[0006] A method for optimizing energy consumption in an airborne STAR-RIS-enabled mobile edge computing system includes the following steps:

[0007] S1. Construct an airborne STAR-RIS-enabled MEC system. The system includes: a single STAR-RIS mounted on a UAV, a BS integrating an MEC server and several antennas, and a group of K IoT devices k. The airborne STAR-RIS flies at an altitude of H, with coordinates: q = {x} s ,y s Let K be K IoT devices, of which R are located in the reflection region, forming reflective IoT devices r; and T are located in the transmission region, forming transmission IoT devices t; where K = R + T.

[0008] S2. By solving the flight energy consumption of STAR-RIS in each time slot and the local computing energy consumption of IoT device k in each time slot, and the energy consumption of IoT device k in one flight cycle. The energy consumption of in-flight task offloading was determined to assess the energy consumption of the STAR-RIS-enabled MEC system.

[0009] S3. Set constraints to minimize the energy consumption of the STAR-RIS-enabled MEC system in the air.

[0010] S4. Minimize the energy consumption of the STAR-RIS-enabled MEC system according to the constraints to obtain the energy consumption optimization design scheme of the STAR-RIS-enabled mobile edge computing system in the air.

[0011] Furthermore, the airborne STAR-RIS consists of a group It consists of several units, each containing a reflection / transmission coefficient and a phase shifter, used to guide the incident signal in the desired direction.

[0012] Furthermore, the specific operation method of step S2 is as follows:

[0013] S2-1 Calculates the flight energy consumption E of STAR-RIS in each time slot. flight (n):

[0014] The flight cycle of STAR-RIS in the air Discretized into a set of N equally spaced time intervals This yields the flight energy consumption E of STAR-RIS in each time slot. flight (n) is:

[0015]

[0016] in, It refers to the quality of the STAR-RIS in the air, including the payload of unmanned aerial vehicles.

[0017] S2-2 Solving for the local computing energy consumption of IoT device k in each time slot

[0018]

[0019] Among them, f k It is the local computing power of IoT device k; λ k (n) is the proportion of tasks offloaded to the MEC server; I k (n) is the input data size for each computation task; G k It is the amount of computing resources required to process 1 bit of input data; It depends on the coefficient of the processor chip architecture.

[0020] S2-3 Solving for IoT device k in a flight cycle Internal task unloading energy consumption

[0021] S2-3-1 uses airborne STAR-RIS to assist in establishing a communication link, and sets the communication link between IoT device k and BS to be entirely assisted by airborne STAR-RIS.

[0022] S2-3-2 Calculate channel gain:

[0023] Channel gain h from reflective IoT device r to BS r (n) is:

[0024]

[0025] Among them, h M,B (n) is the channel response vector between STAR-RIS and BS in the air; h r,M (n) is the channel response vector from the reflective IoT device r to the airborne STAR-RIS; This is the reflection and transmission coefficient matrix.

[0026] Channel gain h from transmission IoT device t to BS t (n) is:

[0027]

[0028] Among them, h M,B (n) is the channel response vector between STAR-RIS and BS in the air; h t,M(n) is the channel response vector from the transmission IoT device t to the airborne STAR-RIS.

[0029] The achievable data rate r of S2-3-3 IoT device k k (n) is:

[0030]

[0031] Where W is the bandwidth of each subcarrier; γ δ (n) represents the signal-to-interference-plus-noise ratio (SIR) of reflective IoT device r and transmittive IoT device t.

[0032] S2-3-4IoT device k unloading task λ k uplink transmission time of (n) for:

[0033]

[0034] S2-3-5IoT device k in a flight cycle Internal task unloading energy consumption for:

[0035]

[0036] The energy consumption E(λ,q,φ,β,p) of the S2-4 airborne STAR-RIS-enabled MEC system is:

[0037]

[0038] Where N is the flight cycle of STAR-RIS in the air. The number of discrete time intervals; K is the number of IoT devices k.

[0039] Furthermore, step S3 sets eight constraints, expressed by equations (10b) to (10i), with the aim of minimizing the energy consumption of the airborne STAR-RIS-enabled MEC system as expressed by equation (10a):

[0040]

[0041] q(1)=q(N), (10i)

[0042] Among them, constraint (10b) limits the delay in task completion to no more than the maximum tolerable delay. Constraint (10c) limits the computing resources of the MEC server to no more than the maximum computing resources F. max Constraint (10d) limits the feasible phase shift value for airborne STAR-RIS transmission. With the feasible phase shift value of reflection All are within (0, 2π); (10e) defines the transmission amplitude coefficient of STAR-RIS in the air. With reflection amplitude coefficient The sum of these values ​​is 1; constraint (10f) limits the task unloading ratio λ. k (n) is within (0,1); constraint (10g) limits the maximum power budget allowed for each IoT device when transmitting data to not exceed its maximum power. The constraint (10h) limits the speed of STAR-RIS in the air to no more than the maximum speed V. max The constraint (10i) is to ensure that the STAR-RIS in the air can return to its starting point at the end of the flight.

[0043] Furthermore, the specific operation method of step S4 is as follows:

[0044] S4-1 defines an MDP, whose components include:

[0045] ①State Space State at each moment The setting is a tuple, represented as:

[0046]

[0047] ② Action Space The action a of n at each moment n ∈A is defined as the increment of the current value, including the task unloading ratio λ. k The position q of STAR-RIS in the air, the phase shift coefficient vector φ, the amplitude vector β, and the transmit power p;

[0048] ③ Reward r n Defined as r n =r(s n ,a n );

[0049] ④ Discount factor γ: a parameter used to adjust the importance of future rewards in the agent's decision-making process;

[0050] S4-2 runs the DDPG algorithm, which includes a policy network (Actor) and a value network (Critic).

[0051] The S4-2-1 intelligent agent collects information from the environment to generate an initial state.

[0052] S4-2-2 Experience Replay and Value Network Evaluation: The master policy network performs experience playback and value network evaluation based on the current state s n Current strategy σ and random noise sequence Choose an action In obtaining action an Then, using the experience replay buffer, the agent transfers the sample to a tuple (s n ,a n ,r n The values ​​, s′) are stored in the replay buffer; simultaneously, based on the current state and action, the main value network updates the value network parameters ξ and calculates the current Q value; the main policy network simulates the Bellman equation by constructing a convolutional neural network to solve the recursive Q function, and updates the value network by randomly selecting 64-128 sample transition tuples from the experience replay memory; the target Q value y j It is calculated from the target value network.

[0053] S4-2-3 Update the value network and the main policy network: The agent randomly selects 64-128 sample transfer tuples from the experience playback memory (s n ,a n ,r n The process is iterated over by y, s′); in each iteration, y is calculated. j The target policy network and target value network are combined and sent to the main value network to minimize the loss function L(ξ):

[0054]

[0055] Where m represents the number of sample transition tuples.

[0056] S4-2-4 By using the value network parameters ξ and the sample transition tuple from the replay buffer, the main policy network updates the current policy using the policy gradient:

[0057]

[0058] S4-3 uses a soft update method to update parameters until the desired policy network is trained.

[0059] ξ′←τξ+(1-τ)ξ′ (14)

[0060] θ′←τθ+(1-τ)θ′ (15)

[0061] Where τ represents the update parameter.

[0062] Mobile edge computing (MEC) effectively addresses the limitations of Internet of Things (IoT) devices in terms of computing power and battery life by allowing devices to offload computing tasks to edge servers. Integrated unmanned aerial vehicles (UAVs) provide enhanced data exchange, rapid deployment, and mobility, effectively overcoming the difficulties of establishing line-of-sight (LoS) connections. The use of reconfigurable smart metasurfaces (RIS), particularly simultaneous transmission and reflection reconfigurable smart surfaces (STAR-RIS) technology, further extends the signal coverage of wireless communication systems and introduces flexibility into MEC. This invention, through the integration of UAVs and STAR-RIS, assists in establishing communication links, facilitating communication between IoT devices and MEC servers; by jointly optimizing the trajectory of the aerial STAR-RIS, the task offloading ratio, the amplitude coefficients of transmission and reflection modes, the phase shift coefficients of transmission and reflection modes, and power allocation, it minimizes the energy consumption of aerial STAR-RIS-enabled MEC systems.

[0063] This invention proposes an airborne STAR-RIS-enabled MEC system, primarily by developing a novel framework to minimize the energy consumption of such systems. This goal is achieved through joint optimization of the airborne STAR-RIS trajectory, task offloading ratio, amplitude coefficients of transmission and reflection modes, phase shift coefficients of transmission and reflection modes, and power allocation. Considering the dynamic nature and unpredictable outcomes of the environment, this invention utilizes deep reinforcement learning as an effective solution. Within DRL technology, this invention employs a deep deterministic policy gradient method, which can accurately and efficiently handle continuous actions, thus significantly reducing the energy consumption of the airborne STAR-RIS-enabled MEC system. Attached Figure Description

[0064] Figure 1 This is a schematic diagram of the system configuration of the STAR-RIS-enabled MEC system in the air. Detailed Implementation

[0065] The invention will now be described in further detail with reference to the accompanying drawings.

[0066] The energy consumption optimization method for the STAR-RIS-enabled mobile edge computing system of this invention includes the following steps:

[0067] S1. Construct an airborne STAR-RIS-enabled MEC system:

[0068] like Figure 1 As shown, the airborne STAR-RIS-enabled MEC system includes: a single STAR-RIS mounted on a UAV, a BS integrating an MEC server and several antennas, and a group of K IoT devices k. The airborne STAR-RIS flies at a fixed altitude H, with coordinates: q = {x}s ,y s Let K be K IoT devices. R of them are located in the reflection region, forming a reflective IoT device r; and T of them are located in the transmission region, forming a transmission IoT device t; where K = R + T.

[0069] The airborne STAR-RIS consists of a group It consists of several units, each containing a reflection / transmission coefficient and a phase shifter, used to guide the incident signal in the desired direction.

[0070] S2. By solving the flight energy consumption of STAR-RIS in each time slot and the local computing energy consumption of IoT device k in each time slot, and the energy consumption of IoT device k in one flight cycle. The energy consumption of in-flight STAR-RIS-enabled MEC systems is determined by offloading tasks within the system. The specific operational method is as follows:

[0071] S2-1 Calculates the flight energy consumption E of STAR-RIS in each time slot. flight (n).

[0072] The flight cycle of STAR-RIS in the air Discretized into a set of N equally spaced time intervals This yields the flight energy consumption of STAR-RIS in each time slot:

[0073]

[0074] in, It refers to the quality of the STAR-RIS in the air, including the payload of unmanned aerial vehicles.

[0075] S2-2 Solving for the local computing energy consumption of IoT device k in each time slot

[0076] Assuming that each IoT device k is assigned a set of computational tasks in each time slot, the local computation execution latency of each IoT device k in each time slot can be expressed as:

[0077]

[0078] Among them, f k It is the local computing power of IoT device k, λ k (n)∈[0,1] is the proportion of tasks offloaded to the MEC server, I k (n) is the input data size for each computational task, G k It is the amount of computing resources required to process 1 bit of input data.

[0079] This yields the local computing power consumption of IoT device k per time slot. for:

[0080]

[0081] in, It depends on the coefficient of the processor chip architecture.

[0082] S2-3 Solving for IoT device k in a flight cycle Internal task unloading energy consumption

[0083] S2-3-1 Using airborne STAR-RIS to assist in establishing a communication link: When IoT device k is offloading its tasks, establishing a Loss of Service (LoS) communication link between the BS and the IoT device may be extremely challenging, or even impossible, due to obstacles. Therefore, it is assumed that there is no direct link between IoT device k and the BS, and the communication link is entirely assisted by airborne STAR-RIS.

[0084] S2-3-2 Calculate channel gain:

[0085] Channel gain h from reflective IoT device r to BS r (n) is:

[0086]

[0087] in, It is the channel response vector between STAR-RIS and BS in the air; It is the channel response vector from the reflective IoT device r to the airborne STAR-RIS; The reflection and transmission coefficient matrices can be specifically represented as follows:

[0088]

[0089] in, These are the amplitude coefficients of the reflected and transmitted signals of the m-th unit, respectively. These are the phase shift values ​​for reflection and transmission of the m-th element, respectively. Using an energy partitioning model, the energy of the incident signal on each element is divided into the energy of the reflected signal and the energy of the transmitted signal. According to the law of conservation of energy, the following is derived:

[0090]

[0091] Channel gain h from transmission IoT device t to BS t (n) is:

[0092]

[0093] Among them, h M,B (n) is the channel response vector between STAR-RIS and BS in the air; h t,M (n) is the channel response vector from the transmission IoT device t to the airborne STAR-RIS.

[0094] S2-3-3 Calculate the achievable data rate r of IoT device k. k (n): h is obtained through S2-3-2 r (n) and h t After (n), respectively for h M,B (n), h r,M (n) and h t,M (n) Using the Ricean channel model, for example: h M,B (n) can be represented as:

[0095]

[0096] in, It is the Rice factor. It is a deterministic line-of-sight component. It is a non-line-of-sight component.

[0097] In channel h r,M With channel h t,M Using the Ricean channel model in the same manner as above, the signal received by the BS can be obtained as follows:

[0098]

[0099] Where, x k (n)=s k p k (n) is the signal sent by IoT device k, s k It is unit power information, p k (n) is the transmit power of IoT device k, and n0 is a power with a mean of 0 and a variance of σ. 2 Additive white Gaussian noise.

[0100] Therefore, the signal-to-interference-plus-noise ratio (SINR) of a reflective IoT device r can be expressed as:

[0101]

[0102] Similarly, the SINR of a transmission-type IoT device t can be expressed as:

[0103]

[0104] Therefore, the achievable data rate of IoT device k can be determined as follows:

[0105]

[0106] In this system, the radio frequency of the BS is divided into orthogonal subcarriers, where W is the bandwidth of each subcarrier; γ δ (n) represents the signal-to-interference-plus-noise ratio (SIR) of reflective IoT device r and transmittive IoT device t.

[0107] S2-3-4 Calculate the task unloading time: r is obtained from S2-3-3. k After (n), we can derive the task λ for unloading IoT device k. k uplink transmission time of (n) for:

[0108]

[0109] S2-3-5IoT device k in a flight cycle Internal task unloading energy consumption for:

[0110]

[0111] S2-4 obtains E through S2-1, S2-2, and S2-3 respectively. flight (n) and Then, by adding up the energy consumption calculated above, the final energy consumption E(λ,q,φ,β,p) of the airborne STAR-RIS-enabled MEC system is:

[0112]

[0113] Where N is the flight cycle of STAR-RIS in the air. The number of discrete time intervals; K is the number of IoT devices k.

[0114] S3. Set constraints to minimize the energy consumption of the STAR-RIS-enabled MEC system in the air.

[0115] Step S3 sets eight constraints, expressed by formulas (10b) to (10i), for the purpose of minimizing the energy consumption of the airborne STAR-RIS-enabled MEC system as expressed by formula (10a), so as to minimize the system energy consumption using DRL.

[0116]

[0117]

[0118] Among them, constraint (10b) limits the delay in task completion to no more than the maximum tolerable delay. Constraint (10c) limits the computing resources of the MEC server to no more than the maximum computing resources F. max Constraint (10d) limits the feasible phase shift value for airborne STAR-RIS transmission. With the feasible phase shift value of reflection All are within (0, 2π); (10e) defines the transmission amplitude coefficient of STAR-RIS in the air. With reflection amplitude coefficient The sum of these values ​​is 1; constraint (10f) limits the task unloading ratio λ. k (n) is within (0,1); constraint (10g) limits the maximum power budget allowed for each IoT device when transmitting data to not exceed its maximum power. The constraint (10h) limits the speed of STAR-RIS in the air to no more than the maximum speed V. max The constraint (10i) is to ensure that the STAR-RIS in the air can return to its starting point at the end of the flight.

[0119] Because the objective function and constraints are coupled and contain nonlinear constraints, this problem is nonconvex. Therefore, this invention adopts a DRL-based scheme, namely a DDPG-based joint optimization design scheme to minimize the energy consumption of the airborne STAR-RIS-enabled MEC system.

[0120] S4. Minimize the energy consumption of the STAR-RIS-enabled MEC system according to the constraints to obtain the energy consumption optimization design scheme of the STAR-RIS-enabled mobile edge computing system in the air. The specific operation method is as follows:

[0121] S4-1 To implement DRL, we first need to define a Markov Decision Process (MDP), which is the basic framework for modeling and solving sequential decision problems in stochastic environments. The components of an MDP include a state space. Action space Reward r n And discount factor γ.

[0122] ①State Space State at each moment The setting is a tuple, represented as:

[0123]

[0124] ② Action Space Action a at each time step n n ∈A is defined as the increment of the current value, including the task unloading ratio λ. kThe STAR-RIS in the air has a position q, a phase shift coefficient vector φ, an amplitude vector β, and a transmit power p. These parameters are defined as increments of their current values. Specifically, position q is represented as: q(n+1) = q(n)⊙Δq(n); phase shift coefficient vector φ is represented as: β(n+1) = β(n)⊙Δβ(n); amplitude vector β is represented as: φ(n+1) = φ(n)⊙Δφ(n); and transmit power p is represented as: p(n+1) = p(n)⊙Δp(n). Here, ⊙ is the Hadamard product, and Δ represents the increment.

[0125] ③ Reward r n To minimize total energy consumption while ensuring tolerable latency, computing resources, and speed conditions are met, the reward r will be... n Defined as r n =r(s n ,a n ).

[0126] ④ Discount factor γ: A parameter used to adjust the importance of future rewards in the agent's decision-making process.

[0127] S4-2 runs the DDPG algorithm, which includes a policy network (Actor) and a value network (Critic).

[0128] The S4-2-1 intelligent agent collects information from the environment to generate an initial state.

[0129] S4-2-2 Experience Replay and Value Network Critic Evaluation: The main policy network Actor evaluates the current state s n Current strategy σ and random noise sequence Choose an action In obtaining action a n Then, using the experience replay buffer, the agent can transfer sample tuples (s n ,a n ,r n The values ​​, s′) are stored in the replay buffer. Simultaneously, based on the current state and action, the main value network Critic can update the value network parameters ξ and calculate the current Q-value. The main value network Critic simulates the Bellman equation by constructing a convolutional neural network (CNN) to solve the recursive Q-function, and updates the value network Critic by randomly selecting 64-128 sample transition tuples from the empirical replay memory. Therefore, the target Q-value y j It can be calculated from the target value network Critic.

[0130] S4-2-3 Update the value network Critic and the principal policy network Actor: To stabilize the learning process and improve the convergence speed of the algorithm, the agent randomly selects 64-128 sample transfer tuples from the experience playback memory (s n ,a n ,r n ,s′), to eliminate coupling in the time series training data. In each iteration, by calculating y j The target policy network (Actor) and the target value network (Critic) are combined and sent to the main value network (Critic) to minimize the loss function L(ξ), which can be expressed as:

[0131]

[0132] Where m represents the number of sample transition tuples.

[0133] S4-2-4 By using the value network parameters ξ and the sample transition tuples from the replay buffer, the main policy network Actor updates the current policy using the policy gradient, which can be expressed as:

[0134]

[0135] S4-3 DDPG uses a soft update method to partially update parameters until the desired policy network is trained:

[0136] ξ′←τξ+(1-τ)ξ′ (14)

[0137] θ′←τθ+(1-τ)θ′ (15)

[0138] Where τ represents the update parameter.

[0139] This process is iterative, continuing until a well-trained policy network (Actor) is reached, up to the maximum number of iterations. For each iteration's network parameter update, the value network (Critic) interacts with the environment and updates its own Q-network reward parameter ξ; it also needs to estimate the Q-values ​​of the current state and action to guide the policy network (Actor) updates.

[0140] By running this DDPG optimization strategy, the mission offloading, airborne STAR-RIS trajectory, amplitude and phase shift coefficients, and transmission power can be optimized, resulting in an optimization strategy σ. out This strategy can minimize the energy consumption of the airborne STAR-RIS-enabled MEC system while ensuring system performance.

[0141] The detailed steps of energy consumption optimization for the STAR-RIS-enabled MEC system based on DDPG are shown in Algorithm 1 of the table below. In the table, the DDPG algorithm runs iteratively until the set number of iterations is reached (row 2). In each round, the simulation environment is initialized and an initial state is randomly generated (rows 3-4). In each time slot, the main policy network Actor will adjust the current state based on the current state s. n The agent generates corresponding actions based on the current policy σ, and performs actions accordingly to generate new states and rewards. Then, the agent can transfer sample transition tuples (s... n ,a n ,r n The values ​​ξ and Q' are stored in the experience replay buffer (lines 5-8). Based on the current state and action, the main value network Critic can update the value network parameters ξ and calculate the current Q value (lines 10-11). Then, the target value network Critic sends the current Q value to the main value network Critic to minimize the loss function L(ξ) (lines 12-13). Based on the value network parameters ξ and the sample transition tuple from the replay buffer, the main policy network Actor updates the current policy using the policy gradient (lines 14-15). At the end of the time slot, the parameters of the target network are updated (lines 17-18).

[0142]

Claims

1. An aerial STAR-RIS empowered mobile edge computing system energy consumption optimization method, characterized in that, Comprising the following steps: S1, constructing an aerial STAR-RIS enabled MEC system; the system includes a single STAR-RIS installed on a UAV, a BS integrated with a MEC server and a plurality of antennas, and a group of K IoT devices k; the flight height of the aerial STAR-RIS is H, and the coordinates are: , the constraint condition is: , to ensure that the aerial STAR-RIS can return to its starting point at the end of the flight; The aerial STAR-RIS is composed of a group of units, each unit containing a reflection / transmission coefficient and a phase shifter for directing the incident signal to the desired direction; R of the K IoT devices k are located in the reflection area, constituting the reflection-type IoT device r; T are located in the transmission area, constituting the transmission-type IoT device t; wherein, K = R + T; the communication link between the IoT device k and the BS is completely assisted by the aerial STAR-RIS; S2, determining the energy consumption of the aerial STAR-RIS empowered MEC system by solving the flight energy consumption of the aerial STAR-RIS per time slot and the local computing energy consumption of the IoT device k per time slot and the task offloading energy consumption of the IoT device k in one flight cycle S2, determining the energy consumption of the aerial STAR-RIS empowered MEC system by solving the flight energy consumption of the aerial STAR-RIS per time slot and the local computing energy consumption of the IoT device k per time slot and the task offloading energy consumption of the IoT device k in one flight cycle S3, setting constraint conditions for minimizing the energy consumption of the STAR-RIS-enabled MEC system in the air; S4, minimizing the energy consumption of the STAR-RIS-enabled MEC system according to the constraint conditions to obtain an optimal design scheme for the energy consumption of the STAR-RIS-enabled mobile edge computing system in the air; The specific operation mode of step S4 is: S4-1 defines MDP, and its components include: : state at each time n is a tuple, denoted by:​ . (11) is the channel response vector between the aerial STAR-RIS and the BS, is the channel response vector from the reflective IoT device r to the aerial STAR-RIS, is the channel response vector from the transmissive IoT device t to the aerial STAR-RIS, is the coordinates of the aerial STAR-RIS, I k (n) is the input data size of each computing task, G k is the amount of computing resources required to process 1-bit input data, T k max (n) is the maximum tolerable delay for task completion, , is the reflective or transmissive coefficient matrix when δ = r, is the reflective coefficient matrix when δ = t, is the transmissive coefficient matrix; ② Action space : Action at each time n is defined as the increment of the current value, containing the task offloading ratio λ k , the position q of the aerial STAR-RIS, the phase shift coefficient vector , the amplitude vector β and the transmit power p; wherein the position is expressed as: ; the phase shift coefficient vector is expressed as: ; the amplitude vector is expressed as: ; the transmit power is expressed as: ; wherein, is the Hadamard product, represents the incremental value; iii. reward r n : defined as r n = r(s n , a n ); (4) Discount factor γ: as a parameter for adjusting the importance of future rewards in the decision-making process of the agent; S4-2 runs the DDPG algorithm including the policy network Actor and the value network Critic: S4-2-1 the agent will collect information from the environment to generate an initial state; S4-2-2 Empirical playback and value network evaluation: the main policy network selects an action a according to the current state s n , the current policy σ and a random noise sequence , and obtains the action a ; after obtaining the action a n , the agent stores the sample transition tuple (s n , a n , r n , s') into the playback buffer using the empirical playback buffer; at the same time, based on the current state and the action, the main value network updates the value network parameters ξ and calculates the current Q value; the main policy network simulates the Bellman equation by constructing a convolutional neural network to solve the recursive Q function, and updates the value network by randomly selecting 64-128 sample transition tuples from the empirical playback memory; the target Q value y j is calculated by the target value network; S4-2-3 Update value network and main policy network: the agent randomly selects 64-128 sample transition tuples (s n , a n , r n , s') from the experience replay memory; in each iteration, the target policy network and the target value network are combined by calculating y j , and sent to the main value network to minimize the loss function L(ξ): (12) Where m represents the number of sample transition tuples; S4-2-4 the main policy network updates the current policy using the policy gradient by using the value network parameter ξ and the sample transition tuples from the replay buffer: (13) S4-3 updates the parameters in a soft update manner until the required policy network is trained: (14) (15) wherein denotes an update parameter; The DDPG algorithm is iteratively run until the set number of iterations is reached.

2. The method of claim 1, wherein the method is performed by an aerial STAR-RIS empowered mobile edge computing system. The specific operation mode of step S2 is: S2-1 solves the flight energy consumption of each time slot of the aerial STAR-RIS : Flight cycle of aerial STAR-RIS is discretized into a set of N equally spaced time intervals The flight energy consumption of aerial STAR-RIS in each time slot is derived as : (1) wherein, , is the mass of the aerial STAR-RIS including the drone payload; S2-2 Solving the local computation energy consumption of IoT device k per time slot : (3) wherein f k is the local computing capability of the IoT device k; λ k (n) is the proportion of tasks offloaded to the MEC server; is a coefficient depending on the processor chip architecture; S2-3 solving for the task offload energy consumption of IoT device k in one flight cycle :​ S2-3-1 uses the STAR-RIS in the air to assist in establishing a communication link, and sets that the communication link between IoT device k and BS is completely assisted by the STAR-RIS in the air; S2-3-2 calculates the channel gain: Channel gain h from reflective IoT device r to BS r (n) is: (4) Channel gain h from transmission-type IoT device t to BS t (n) is: (5) S2-3-3 Data rate r achievable by IoT device k k (n) is: (6) where W is the bandwidth of each subcarrier; γ δ (n) is the signal-to-interference-plus-noise ratio of the IoT device in the reflection region, and when δ = r, γ δ (n) is the signal-to-interference-plus-noise ratio of the IoT device in the reflection region, and δ (n) is the signal-to-interference-plus-noise ratio of the IoT device in the transmission region; S2-3-4 IoT device k offloads task λ k (n) the uplink transmission time t k off (n) is: (7) S2-3-5 IoT device k in one flight cycle Task offloading energy consumption E k off (n) is: (8) S2-4 Energy consumption of an aerial STAR-RIS empowered MEC system For: (9) where N is the flight period of the aerial STAR-RIS the number of discrete time intervals; K is the number of IoT devices k.

3. The method of claim 2, wherein the method is performed by an aerial STAR-RIS empowered mobile edge computing system. Step S3 sets eight constraint conditions expressed by formulas (10b) to (10i) for the purpose of minimizing the energy consumption of the STAR-RIS-enabled MEC system in the air expressed by formula (10a): (10a) (10b) (10c) (10d) (10e) (10f) (10g) (10h) (10i) where constraint (10b) limits the delay of task completion not to exceed the maximum tolerable delay T k max ; constraint (10c) limits the computing resource of the MEC server not to exceed the maximum computing resource F max ; constraint (10d) limits the transmission feasible phase shift value of the aerial STAR-RIS and the reflection feasible phase shift value are both within ; constraint (10e) limits the transmission amplitude coefficient β m t (n) of the aerial STAR-RIS m r (n) to be added up to 1; constraint (10f) limits the task offloading ratio λ k (n) to be within ; constraint (10g) limits the maximum power budget allowed to be used by each IoT device when transmitting data not to exceed its maximum power p k max ; constraint (10h) limits the speed of the aerial STAR-RIS not to exceed the maximum speed V max ; constraint (10i) is to ensure that the aerial STAR-RIS is able to return to its starting point at the end of the flight.

Citation Information

Patent Citations

  • Energy consumption optimization method in IRS and UAV auxiliary MEC system

    CN117826971A

  • Multi-unmanned aerial vehicle assisted MEC system safety communication energy efficiency optimization method enhanced by using air RIS

    CN118632298A