Average information age optimization method based on AARIS mobile edge calculation
Through air-ground collaborative mobile edge computing, combined with AARIS and NOMA technologies, DHPG algorithms are improved to optimize drone trajectory and beamforming, solving the problem of insufficient freshness of calculation results in real-time scenarios by traditional MEC networks, and achieving a significant reduction in information age and an improvement in resource utilization efficiency.
Patent Information
- Application Number
- CN202510495457.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Traditional mobile edge computing networks have limited service coverage in remote areas and user-intensive areas, and fixed deployment strategies are difficult to cope with fluctuations in user demand, resulting in insufficient freshness of computing results, especially in real-time scenarios such as autonomous driving.
Air-ground collaborative mobile edge computing is used, and communication models are constructed in combination with air-active intelligent reflective surfaces (AARIS) and non-orthogonal multiple access (NOMA). UAV trajectory and beamforming are optimized through improved deep hybrid strategy gradient (DHPG) algorithm, and the minimum system average information age objective function and constraints are constructed to improve information freshness and timeliness.
Effectively reduce the average information age of the system by 30%-50%, improve information freshness, enhance the adaptability and robustness of the algorithm in complex environments, optimize resource utilization efficiency, extend the battery life of the drone, reduce multi-user interference, and ensure policy stability.
Smart Images

Figure CN120456060A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile edge computing technology assisted by airborne active intelligent reflective surfaces (AARIS), and in particular to a method for optimizing average information age based on AARIS mobile edge computing. Background Art
[0002] Mobile Edge Computing (MEC), a cutting-edge computing model, focuses on embedding computing and caching capabilities directly at the edge of the network, close to user devices. This significantly improves service efficiency and reduces latency. While traditional computing servers deployed in ground base stations play an indispensable role in MEC networks, they face several key limitations. The primary issue is the relatively limited deployment density of ground base stations, making it difficult to fully cover remote areas or densely populated areas, thus limiting the effective range of services. Secondly, the construction and maintenance of ground base stations, especially in complex terrain and sparsely populated areas, require significant capital investment, making costs particularly prominent. Furthermore, the fixed deployment strategy of ground base stations significantly reduces their flexibility to respond to fluctuations in user demand. In particular, ground base stations struggle to quickly adapt to temporary or sudden changes in demand, preventing them from providing immediate responses. These limitations pose significant challenges to the performance and service quality of traditional mobile edge computing networks.
[0003] To address these issues, unmanned aerial vehicles (UAVs) have been introduced into mobile edge computing networks. With their high maneuverability, flexible deployment, and low cost, UAVs can quickly respond to dynamic demands, expand service coverage, and improve service quality. However, with the growing demand for high-performance computing services, relying solely on aerial mobile edge computing is no longer sufficient. To address this, Air-Ground Collaborative Mobile Edge Computing (AGC-MEC) was proposed. By integrating the computing resources of aerial UAVs and ground base stations, AGC-MEC provides more efficient and flexible computing support for user devices.
[0004] Performance optimization in air-ground collaborative mobile edge computing networks has attracted considerable attention. However, current research has largely focused on traditional performance metrics, such as system latency and energy consumption. However, in real-time applications such as autonomous driving and anomaly detection, the freshness of computational results is particularly important. Once information becomes outdated, it can become invalid and even pose a security risk. Therefore, the freshness of computational results has become a key factor in evaluating the performance of air-ground mobile edge computing. To effectively address this issue, research has introduced the concept of "information freshness," which is the time interval between the generation of the latest data by an IoT terminal and the receipt of that data by an information receiver.
[0005] For example, the invention application with application number 202411184537.8 discloses an AoI optimization method and system for a multi-UAV-supported communication and perception system, belonging to the field of mobile communications. This application scheme minimizes the average information age of ground equipment by jointly optimizing UAV trajectories, user association, target sensor selection, and communication and perception beamforming. However, its scheme also has the following problems: (1) It does not effectively construct a communication model that coordinates AARIS with non-orthogonal multiple access (NOMA), which is not conducive to ensuring data timeliness in real-time scenarios; (2) It does not adopt the DHPG algorithm, which is not conducive to communication adaptability in complex environments. Summary of the Invention
[0006] In view of the above problems, the purpose of the present invention is to provide an average information age optimization method based on AARIS mobile edge computing.
[0007] An embodiment of the present invention provides a method for optimizing average information age based on AARIS mobile edge computing, comprising the steps of:
[0008] S1. Improve the deep deterministic policy gradient DDPG algorithm to obtain the deep hybrid policy gradient DHPG algorithm;
[0009] S2. Build a mobile edge computing communication model, information age model, and system energy consumption model based on the aerial active intelligent reflective surface AARIS and non-orthogonal multiple access NOMA;
[0010] S3. Based on the mobile edge computing communication model, information age model and system energy consumption model, construct the minimum system average information age objective function and constraints to improve the freshness and timeliness of information in the mobile edge computing communication system.
[0011] Furthermore, the improvements to the DDPG algorithm in S1 include:
[0012] Introducing a hybrid strategy mechanism to improve the adaptability and robustness of the DDPG algorithm in complex environments by combining the advantages of deterministic and random strategies;
[0013] Optimize the network structure and enhance the learning efficiency and convergence performance of the algorithm by adjusting the number of neural network layers, number of nodes, and activation functions of the DDPG algorithm.
[0014] Furthermore, in S2, a mobile edge computing communication model is constructed, including that the incident signal sent by the ground user is sent to the base station after being directionally enhanced by AARIS; at the same time, the base station receives the incident signal directly sent from the ground user; the base station decodes the received mixed signal with the help of non-orthogonal multiple access technology and serial interference cancellation technology to obtain the mixed signal, and the steps are as follows:
[0015] S21. Obtain an equivalent channel gain based on the channel gain of the link between the user and AARIS, the channel gain of the link between the base station and AARIS, and the channel gain of the link between the user and the base station. The formula is:
[0016]
[0017] in, is the channel gain between the kth ground user and AARIS, h R (t) is the channel gain of the link between the base station and AARIS; is the direct channel gain between the ground user and the base station in the system model;
[0018] S22, calculating the mixed user input signal of AARIS and the mixed user amplified signal of AARIS;
[0019] The formula for the mixed user input signal of AARIS is:
[0020]
[0021] in, is the channel gain between the kth ground user and AARIS, p k is the constant transmission power of user k, U k (t) is randomly generated in each time slot through Poisson distribution, which is used to represent whether user k has a communication task in time slot t. k (t) transmit signals for k users;
[0022] The formula for AARIS's hybrid user amplification signal is:
[0023]
[0024] Where x(t) is the mixed user input signal of AARIS, A(t)Θ(t)x(t) is the expected reflected signal, and A(t)Θ(t)x d (t) is the dynamic noise of AARIS, n sIt is static noise, which can be ignored compared with dynamic noise;
[0025] S23. Obtain a mixed user signal after reflection enhancement based on the equivalent channel gain, the mixed user input signal of AARIS, the mixed user amplified signal of AARIS, and the user transmission signal. The formula is:
[0026]
[0027] Among them, A(t) is the power amplification factor matrix, Θ(t) is the phase offset matrix, n d (t) represents the additive white Gaussian noise at the base station.
[0028] Furthermore, the power amplification factor matrix is expressed as follows:
[0029]
[0030] Among them, the diagonal element a m (t) represents the magnification of each reflection unit;
[0031] The phase offset matrix is expressed as:
[0032]
[0033] Among them, the diagonal elements Indicates the phase offset of each reflector unit.
[0034] Furthermore, the information age model is constructed in S3, including the steps of:
[0035] S31. Calculate the interference power I between users suffered by user k. k (t), the formula is expressed as:
[0036]
[0037] Among them, η jk A 01 identifier, when the value is 1, it means that the signal strength of user j is stronger than that of user k, D j (t) = 1 means that user j decodes successfully, D j (t) = 0 means decoding failure or no decoding; ν∈(0,1) quantifies the information distortion caused by channel state uncertainty and hardware limitations, p j (t) represents the transmission power of other users other than user k.
[0038] S32. Calculate the current real rate R of user k using the Shannon formula k (t), the formula is expressed as:
[0039]
[0040] Among them, r k (t) represents the signal strength of user k, Indicates that when AARIS actively amplifies the signal, it will also amplify the noise power it receives; I k (t) represents the interference signal power received by user k; δ 2 is the noise power.
[0041] S33. Construct an information age model, the formula is expressed as:
[0042]
[0043] Among them, O k (t) represents the lifetime of the data packet of user k in time slot t. If there is a new data packet to be sent (i.e., U k (t)=1), O k (t) is reset to 0, otherwise O k (t) Add 1 to the original value; Δ k (t+1) represents the information age of user k in time slot t+1. When user k successfully transmits the task in time slot t (i.e., R k (t)>=R0, R0 represents the minimum communication rate requirement), S k (t) = 1, the information age of the t+1 time slot is the lifetime of the data packet of the t time slot + 1, otherwise, it is the information age of the t time slot + 1. It should be noted that Δ max Indicates an upper limit on the age of information to prevent it from growing forever.
[0044] Furthermore, the system energy consumption model is expressed as follows:
[0045]
[0046] E f (t) = τP U (v(t))
[0047]
[0048] E c (t) = E f (t)+E i (t)
[0049] Among them, P U (v) is the flight power of the UAV, P0 and P1 are the blade profile power and induced power when the UAV is hovering. tip is the rotor tip speed, v0 is the average rotor induced speed; E f(t) is the flight energy consumption of time slot t, τ is the length of a time slot; E i (t) is the energy consumption of active RIS, E c (t) is the sum of the two parts of energy consumption, that is, the total energy consumption of AARIS.
[0050] Furthermore, the minimum system average information age objective function and constraint conditions in S4 are expressed as follows:
[0051]
[0052] Where (P) is the objective function, is the average information age of the system, Satisfy the UAV's flight speed v(t) and the UAV's horizontal flight angle θ u (t); Signal amplification factor a of each reflection unit of AARIS m (t) and the phase shift angle θ of each reflector unit of AARIS m The minimum value of (t).
[0053] subject to: C1, C2, C3, C4 and C5 are constraints;
[0054] C1 is the current real rate R of user k k (t) is greater than the set value R0;
[0055] C2 is the current remaining energy E of the drone r (t) greater than 0;
[0056] C3 is the two-dimensional position coordinate q of the UAV u , flight speed v(t) and horizontal flight angle θ m (t) meet the set value;
[0057] C4 is the signal amplification factor of each reflection unit of AARIS m (t), phase shift angle θ m (t) meet the set value;
[0058] C5 is the current time slot t, the number of users k and the number of reflection units m of AARIS meeting the set values.
[0059] Furthermore, the average information age of the system is expressed as:
[0060]
[0061] Among them, Δ k (t) represents the information age of user k in time slot t.
[0062] Beneficial effects of the present invention:
[0063] 1. This invention builds a communication model that combines aerial active intelligent reflective surface (AARIS) with non-orthogonal multiple access (NOMA), and combines it with the deep hybrid policy gradient (DHPG) algorithm to optimize the trajectory and beamforming of UAVs, thereby reducing the average age of information (AoI) of the system by 30%-50% (e.g. Figure 3-4 It effectively ensures the timeliness of data in real-time scenarios such as autonomous driving and significantly improves the freshness of information.
[0064] 2. The improved DHPG algorithm of the present invention combines the advantages of the DDPG deterministic strategy and the PPO random strategy, and its convergence speed is increased by 40% compared with the traditional DDPG algorithm. It also enhances its adaptability to complex environments through hybrid action space design (formula: new action = 0.6 × deterministic action + 0.4 × probabilistic action), effectively enhancing the algorithm's convergence performance.
[0065] 3. The AARIS of the present invention actively amplifies the signal to achieve equivalent channel gain, and combines the NOMA technology to reduce multi-user interference. The user's actual rate meets the QoS constraint, improves communication efficiency, and optimizes resource utilization efficiency.
[0066] 4. The energy consumption model of the system of the present invention dynamically balances the flight energy consumption and the AARIS reflection energy consumption, thereby extending the flight endurance of the UAV under energy constraints and effectively controlling energy consumption.
[0067] 5. The active reflection unit (amplification factor A(t), phase shift Θ(t)) of the AARIS of the present invention improves signal strength compared to the passive RIS. At the same time, the air-ground collaborative model achieves interference suppression and noise elimination through joint optimization of equivalent channels.
[0068] 6. This invention uses a Markov decision process and a dual reward mechanism to ensure policy stability under constraints such as location and energy (C1-C5). Experiments show that AoI fluctuations are reduced by 25% in burst traffic scenarios, enhancing robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 Schematic diagram of the process of optimizing the average information age based on AARIS mobile edge computing of the present invention;
[0070] Figure 2 This is a schematic diagram of the structure of the mobile edge computing communication model of the present invention;
[0071] Figure 3 This is a simulation verification effect diagram of non-orthogonal multiple access and orthogonal multiple access using the method of the present invention;
[0072] Figure 4 These are simulation verification effect diagrams of the inventive method with and without AARIS. DETAILED DESCRIPTION
[0073] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings. The same or similar symbols throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0074] Performance optimization in air-ground collaborative mobile edge computing networks has attracted considerable attention. However, current research has largely focused on traditional performance metrics, such as system latency and energy consumption. However, in real-time applications such as autonomous driving and anomaly detection, the freshness of computational results is particularly important. Once information becomes outdated, it can become invalid and even pose a safety hazard.
[0075] In response to the above problems, the present invention provides an average information age optimization method based on AARIS mobile edge computing. Figure 1 A flowchart of a method for optimizing average information age based on AARIS mobile edge computing provided in an embodiment of the present invention includes:
[0076] S1. Improve the deep deterministic policy gradient DDPG algorithm and obtain the deep hybrid policy gradient DHPG algorithm.
[0077] Among them, the improvements to the DDPG algorithm include: introducing a hybrid strategy mechanism to improve the adaptability and robustness of the DDPG algorithm in complex environments by combining the advantages of deterministic strategies and random strategies; optimizing the network structure to enhance the learning efficiency and convergence performance of the algorithm by adjusting the number of neural network layers, number of nodes and activation functions of the DDPG algorithm.
[0078] The traditional DDPG algorithm outputs deterministic actions (a1, a2, a3, ... a1) with dimension N through the Actor network. N ), in order to increase the exploration efficiency of the algorithm, it is usually necessary to add random noise to the action to improve the exploration efficiency of the agent in the algorithm, such as: (a1+δ1, a2+δ2, a3+δ3, a N +δ N ), where δ is random noise. The required work is to introduce the output of the PPO algorithm based on the probability strategy strategy on the basis of deterministic actions. The way the PPO algorithm outputs actions is different from the DDPG algorithm. The PPO algorithm outputs the mean and standard deviation of N actions through the Actor network to form a probability distribution, and obtains the action (a'1, a2', a3', ..., N a)'. This randomness advantage is utilized, and the deterministic actions of DDPG and the probabilistic actions of PPO are weighted to form a hybrid strategy.
[0079] Specifically, two algorithms are used to generate actions simultaneously, and then new actions are generated for the agent to execute in the following way.
[0080] The new hybrid action = weight 1 × deterministic action + weight 2 × probabilistic action. The weights were determined to be 0.6 and 0.4 through extensive simulations. This new action is more purposeful than the original DDPG algorithm's exploration method, which introduced random noise, resulting in faster convergence. Therefore, the hybrid strategy mechanism, combining the advantages of deterministic and random strategies, improves the adaptability and robustness of the DDPG algorithm in complex environments.
[0081] S2. Based on the aerial active intelligent reflective surface AARIS and non-orthogonal multiple access NOMA, a mobile edge computing communication model, information age model and system energy consumption model are constructed.
[0082] Among them, the mobile edge computing communication model structure, such as Figure 2 As shown, the system includes a drone, an active intelligent reflective surface panel, a ground base station and K ground user terminals; the drone and the intelligent reflective panel are connected together through components to form an aerial active intelligent reflective surface AARIS, sharing the same control system and power supply; AARIS uses a controller to adjust the signal amplification factor and the phase offset of the reflective unit to directionally enhance the incident signal sent by the ground user and send it to the base station; the ground base station receives signals sent directly from the ground user and through the aerial active intelligent reflective surface; the ground user randomly generates a transmission task.
[0083] The channel gain of the link between the ground user and AARIS is obtained based on the position information of the ground user and AARIS. The formula is expressed as:
[0084]
[0085] in, represents the channel gain between the kth ground user and AARIS, Indicates the path loss of the link and the large-scale fading caused by shadowing; Indicates the small-scale fading factor of the link.
[0086] Obtain the channel gain of the link between the ground base station and AARIS based on the location information of the ground base station and AARIS;
[0087]
[0088] where h R (t) represents the channel gain of the link between AARIS and the base station, g R(t) represents the path loss of the link and the large-scale fading caused by shadowing; Indicates the small-scale fading factor of the link.
[0089] Obtain the channel gain of the link between the user and the base station based on the location information of the user and the ground base station;
[0090]
[0091] in, In the system model, the direct channel gain between the ground user and the base station is represented by f k (t) is the Rayleigh attenuation coefficient, β is the path loss constant, is the communication distance, α is the path loss exponent, represents log-normal shadow fading, and σ is the standard deviation of shadow fading.
[0092] The equivalent channel gain is obtained based on the channel gain of the link between the user and AARIS, the channel gain of the link between the base station and AARIS, and the channel gain of the link between the user and the base station. The formula is expressed as follows:
[0093]
[0094] in, represents the channel gain between the kth ground user and AARIS, h R (t) is the channel gain of the link between the base station and AARIS; Represents the direct channel gain between the ground user and the base station in the system model.
[0095] The power amplification factor matrix and phase shift matrix are obtained according to the amplitude coefficient and phase shift coefficient of the smart reflector element. The power amplification factor matrix is expressed as follows:
[0096]
[0097] Diagonal element a m (t) represents the magnification of each reflection unit.
[0098] The phase offset matrix is expressed as follows:
[0099]
[0100] Diagonal elements Indicates the phase offset of each reflector unit.
[0101] Then, the AARIS mixed user input signal and the AARIS mixed user amplified signal are calculated.
[0102] The formula for the mixed user input signal of AARIS is:
[0103]
[0104] in, is the channel gain between the kth ground user and AARIS, p k is the constant transmission power of user k, U k (t) is randomly generated in each time slot through Poisson distribution, which is used to represent whether user k has a communication task in time slot t. k (t) transmit signals for k users;
[0105] The formula for AARIS's hybrid user amplification signal is:
[0106]
[0107] Where x(t) is the mixed user input signal of AARIS, A(t)Θ(t)x(t) is the expected reflected signal, and A(t)Θ(t)x d (t) is the dynamic noise of AARIS, n s It is static noise, which can be ignored compared with dynamic noise;
[0108] Furthermore, based on the equivalent channel gain, the base station receives the signals from the air channel of AARIS and the signals transmitted via the direct link. Therefore, the mixed signal received at the base station is expressed as follows:
[0109]
[0110] Among them, A(t) is the power amplification factor matrix, Θ(t) is the phase offset matrix, n d (t) represents the additive white Gaussian noise at the base station, represents the interference signal from other users compared to user k (symbol j represents a user different from user k).
[0111] By calculating the above formula, we can construct Figure 2 The mobile edge computing communication model structure is shown.
[0112] Furthermore, an information age model is constructed.
[0113] First, calculate the interference power I between users affected by user k k (t), the formula is expressed as:
[0114]
[0115] Among them, η jkA 01 identifier, when the value is 1, it means that the signal strength of user j is stronger than that of user k, D j (t) = 1 means that user j decodes successfully, D j (t) = 0 means decoding failure or no decoding; ν∈(0,1) quantifies the information distortion caused by channel state uncertainty and hardware limitations, p j (t) represents the transmission power of other users other than user k.
[0116] Then, by combining the Shannon formula with the interference power I between users, user k is k (t), calculate the current real rate R of user k k (t), the formula is expressed as:
[0117]
[0118] Among them, r k (t) represents the signal strength of user k, Indicates that when AARIS actively amplifies the signal, it will also amplify the noise power it receives; I k (t) represents the interference signal power received by user k; δ 2 is the noise power.
[0119] Then, the information age model is constructed, and the formula is expressed as:
[0120]
[0121] Among them, O k (t) represents the lifetime of the data packet of user k in time slot t. If there is a new data packet to be sent (i.e., U k (t)=1), O k (t) is reset to 0, otherwise O k (t) Add 1 to the original amount.
[0122] Δ k (t+1) represents the information age of user k in time slot t+1. When user k successfully transmits the task in time slot t (i.e., R k (t)>=R0, R0 represents the minimum communication rate requirement), S k (t) = 1, the information age of the t+1 time slot is the lifetime of the data packet of the t time slot + 1, otherwise, it is the information age of the t time slot + 1. It should be noted that Δ max Indicates an upper limit on the age of information to prevent it from growing forever.
[0123] Finally, according to the information age Δ of user k in time slot t k (t), calculate the average information age of the system, the formula is expressed as:
[0124]
[0125] Among them, Δ k (t) represents the information age of user k in time slot t.
[0126] Furthermore, a system energy consumption model is constructed, and the formula is expressed as follows:
[0127]
[0128] E f (t) = τP U (v(t))
[0129]
[0130] E c (t) = E f (t)+E i (t)
[0131] Among them, P U (v) is the flight power of the UAV, P0 and P1 are the blade profile power and induced power when the UAV is hovering. tip is the rotor tip speed, v0 is the average rotor induced speed; E f (t) is the flight energy consumption of time slot t, τ is the length of a time slot; E i (t) is the energy consumption of active RIS, E c (t) is the sum of the two parts of energy consumption, that is, the total energy consumption of AARIS.
[0132] Among them, the formula for obtaining P0 and P1 is:
[0133]
[0134] The parameters δ, Ω, R, and W are the profile drag coefficient, blade angular velocity, rotor radius, and aircraft weight, respectively; d0, ρ, s, and A are the fuselage drag ratio, air density, rotor solidity, and rotor disk area, respectively.
[0135] S3. Based on the mobile edge computing communication model, information age model and system energy consumption model, construct the minimum system average information age objective function and constraints to improve the freshness and timeliness of information in the mobile edge computing communication system.
[0136] The expression of the minimum system average information age optimization problem is:
[0137]
[0138] Where (P) is the objective function, is the average information age of the system, Satisfy the UAV's flight speed v(t) and the UAV's horizontal flight angle θ u (t); Signal amplification factor a of each reflection unit of AARIS m (t) and the phase shift angle θ of each reflector unit of AARIS m The minimum value of (t).
[0139] subject to: C1, C2, C3, C4 and C5 are constraints;
[0140] C1 is the current real rate R of user k k (t) is greater than the set value R0, which defines the minimum real rate (QoS) requirement for each user.
[0141] C2 is the current remaining energy E of the drone r (t) is greater than 0, defining the energy consumption requirement of the UAV.
[0142] C3 is the constraint adjustment of the active RIS in the air, including the two-dimensional position coordinates q of the UAV u , satisfying x u ∈[x min ,x max ],y u ∈[y min ,y max ]; flight speed v(t) satisfies the set value s(t) = q u (t),E r (t), x(t), u(t), horizontal flight angle θ u (t) satisfies [0,2π].
[0143] C4 is the signal amplification factor of each reflection unit of AARIS m (t), phase shift angle θ m (t) satisfies the set value; where L and V max They are the maximum magnification and the maximum speed of the drone, respectively.
[0144] C5 is the current time slot t, the number of users k and the number of reflection units m of AARIS meeting the set values.
[0145] Furthermore, the drone and AARIS interact with the environment to obtain state information and find the optimal strategy to maximize the cumulative reward. The state space, action space, and reward function of the Markov decision process are defined as follows:
[0146] State space, at the current time slot t, the state space is defined by several key elements. Among them is the horizontal coordinate q of the UAV u (t), the current remaining energy E of the droner (t), in addition, it also includes the information age Δ(t) of each ground user and the mission arrival state u(t). The state space can be expressed as:
[0147] s(t)=q u (t),E r (t),Δ(t),u(t)
[0148] Action space, at each time slot t, the agent's action includes the drone's flight direction, drone's flight speed, the signal amplification coefficient matrix of the active intelligent reflective surface, and the phase offset matrix. The action space can be expressed as:
[0149] a(t)=v u (t),θ u (t),Α(t),Θ(t)
[0150] The reward function, the reward obtained by taking action a in the state of the t-th time slot can be defined as:
[0151]
[0152] Where r0 represents a fixed reward value, which is used to motivate the agent to explore in the early stages of learning. w and b are the weighting coefficients of the average AoI and energy consumption in the reward function, respectively. and The additional negative penalty and reward items can be expressed as:
[0153]
[0154] The algorithm execution process includes:
[0155] Initialize the relevant parameters of the six neural networks and the parameters in the simulation environment.
[0156] The agent obtains the current state information s(t) from the environment, inputs the DDPG main actor network and the PPO-based actor network to obtain two deterministic actions a d and the probability action a p , and then perform weighted fusion of the two actions.
[0157] The fused action is executed in the simulation environment to obtain the reward value r(t), the next state information s(t+1) and the state symbol done during training.
[0158] Store the current state, current action, next state, current reward and training terminator in the cache.
[0159] When the number of experiences in the cache is greater than the batch size, training begins and the network parameters are updated.
[0160] The target Q value in DDPG is calculated using the Bellman equation, which estimates the expected cumulative reward for the next state and action. The Bellman equation is expressed as follows:
[0161] y(t)=r(t)+γ1Q′(s(t+1),μ′(s(t+1))|θ Q′ )
[0162] where r(t) is the current reward obtained after executing action a(t) in state s(t), γ1∈[0,1] is the discount factor; Q′(s(t+1),μ′(s(t+1))|θ Q′ ) represents the target Q-value estimated by the target critic network, and Q(s(t),μ(s(t))|θ Q ) represents the Q value predicted by the current network in the DDPG module.
[0163] DDPG's critic network updates are achieved by minimizing the following loss function and optimizing the critic network parameters θ by gradient descent: Q .
[0164]
[0165] Among them, N b represents the batch size, that is, the number of experiences extracted from the replay buffer D at each update. d (t)|θ Q ) represents the value estimated by the current critic network. By minimizing this loss function, the parameters θ of the DDPG critic network Q is updated, thereby improving the accuracy of its predicted state-action pair values.
[0166] The loss function of the PPO critic network is defined as:
[0167]
[0168] Where V(s(t)|θ V ) is the state value estimated by the PPO critic network. By minimizing this loss function, the parameters θ of the PPO critic network V is adjusted to provide a more accurate value assessment for each state.
[0169] After the critic network is updated, the deterministic actor network (DDPG-actor) and the auxiliary actor network (PPO-actor) are also updated. The deterministic actor network is optimized using policy gradients as follows:
[0170]
[0171] in, is the Q value relative to the deterministic action a d (t). This gradient is used to update the parameters θ of the deterministic actor network μ .
[0172] The auxiliary actor network following the PPO framework is also updated according to the clip strategy to ensure the stability of learning. The gradient of the auxiliary actor network update is:
[0173]
[0174] in, represents the ratio of the probability of the new strategy to the old strategy, represents the estimated advantage function, and ∈ D Here, the time difference TD error is used to approximate the advantage function
[0175] The ratio of the new and old strategies is calculated as follows: o (π)=exp(lp(t)-lp old )
[0176] The advantage function in the PPO module can be defined as:
[0177]
[0178] In addition, the target network in the DDPG algorithm obtains updated parameters from the main network through soft updates.
[0179] θ Q′ ←τ D θ Q +(1-τ D )θ Q′ ,θ μ′ ←τ D θ μ +(1-τ D )θ μ′
[0180] Among them, τ D The soft update factor, typically a small value, ensures a gradual adjustment of the target network parameters. This approach reduces instability during training by promoting smoother updates.
[0181] The core of DDPG consists of an actor network μ(s(t)|θ) for policy optimization. μ ), a critic network Q(s(t),a(t)|θ for value estimation Q), and their respective target networks μ'(s(t)|θ μ' ) and Q'(s(t),a(t)|θ Q ) to stabilize the training process. The auxiliary module is based on PPO and includes an actor network π(s(t)θ for enhanced exploration π ) and a critic network V(s(t)θ for value function approximation V ).
[0182] In each time slot t, the agent perceives the current state S(t) of the MEC environment from NOMA and AARIS assistance, which is defined by the formula. S(t) is then input into μ(s(t)|θ μ ) and π(s(t)|θ π ) to generate a mixed action a(t). The agent then performs a h (t), obtain reward r(t) and transfer to the next state s(t+1). Experience tuple (s(t), a h (t),a d (t),a p (t), lp(t), r(t), s(t+1), done) are then stored into the playback buffer D.
[0183] Figure 3 We validated the advantages of using non-orthogonal multiple access (NMA) in a mobile edge computing network assisted by an aerial active smart reflective surface (ASR). We compared the performance of NMA and OMA in this network. We used the traditional DDPG algorithm and an adaptively optimized DHPG algorithm to jointly optimize the flight trajectory of the drone and the beamforming of the ASR. First, we observed that the application of NMA resulted in a lower average information age in the network than that of OMA using both the traditional DDPG and the improved DHPG algorithms. This demonstrates that NMA improves network performance in the current multi-user scenario, effectively reducing the average information age in the system. Second, we observed that the improved DHPG algorithm outperformed the traditional DDPG algorithm in both the NMA and OMA scenarios, demonstrating faster convergence. Therefore, NMA achieved superior performance in the ASR network, demonstrating its superior adaptability to ASR networks.
[0184] Figure 4This paper describes the impact of active and passive smart reflective surfaces on experimental results in a mobile edge computing network assisted by non-orthogonal multiple access technology. It can be seen that, using both the traditional and improved methods, active smart reflective surfaces achieve better results in optimizing the information age in the network compared to passive methods. This is because active smart reflective surfaces have the function of signal amplification, which can directionally amplify the incident signal, thereby improving data transmission efficiency and reducing the average information age. Similarly, for both active and passive smart reflective surfaces, the improved algorithm outperforms the traditional algorithm in terms of convergence speed. Therefore, compared to passive smart reflective surfaces, active smart reflective surfaces can achieve better performance in non-orthogonal multiple access-assisted air-to-ground mobile edge computing networks.
[0185] Reconfigurable Intelligent Surfaces (RIS), an innovative technology in future wireless communication systems, combined with unmanned aerial vehicles (UAVs) and mobile edge computing (MEC), are revolutionizing the wireless communications field. Through the precise control of RIS, electromagnetic wave propagation paths can be flexibly optimized, thereby enhancing signal strength, effectively suppressing interference, and significantly improving energy efficiency. Furthermore, by integrating RIS with unmanned aerial vehicle (UAV) technology, leveraging the high maneuverability of UAVs and the precise control of signal propagation paths by AARIS, an innovative aerial active intelligent reflective surface (AARIS) has been constructed. This combination enables the establishment of an aerial channel in mobile edge computing (MEC) systems, significantly improving the quality of service (QoS) of MEC systems and significantly enhancing the timeliness of computational results and the freshness of data. Due to the complex, ever-changing, and random uncertainty of the communication environment, traditional methods face the dual challenges of drone flight trajectory planning and AARIS beamforming adjustment when controlling aerial intelligent reflective surfaces (especially when combined with drones and AARIS). To address this, we introduced deep reinforcement learning (DRL) as a solution. Within the DRL framework, we use the system's real-time state (including drone location, ground equipment mission generation, energy consumption status, and the information age of each ground equipment) as input. This information is used to extract high-dimensional features through a deep neural network (DNN), which then predicts the long-term benefits of different action strategies (such as adjusting the drone's flight path and adjusting the beamforming of the active intelligent reflective surface).
[0186] Leveraging the trial-and-error learning mechanism of reinforcement learning, our system continuously collects feedback during operation, dynamically adjusts policy parameters, and gradually approaches the optimal control strategy. Furthermore, we innovatively integrate two mainstream deep reinforcement learning algorithms, significantly improving the performance of a single algorithm in jointly optimizing drone trajectories and beamforming, achieving more efficient and precise control of intelligent reflective surfaces in the air.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for optimizing average information age based on AARIS mobile edge computing, characterized in that: include: S1. Improve the deep deterministic policy gradient DDPG algorithm to obtain the deep hybrid policy gradient DHPG algorithm; S2. Build a mobile edge computing communication model, information age model, and system energy consumption model based on the aerial active intelligent reflective surface AARIS and non-orthogonal multiple access NOMA; S3. Based on the mobile edge computing communication model, information age model and system energy consumption model, construct the minimum system average information age objective function and constraints to improve the freshness and timeliness of information in the mobile edge computing communication system.
2. The average information age optimization method according to claim 1, characterized in that: The improvements to the DDPG algorithm in S1 include: Introducing a hybrid strategy mechanism to improve the adaptability and robustness of the DDPG algorithm in complex environments by combining the advantages of deterministic and random strategies; Optimize the network structure and enhance the learning efficiency and convergence performance of the algorithm by adjusting the number of neural network layers, number of nodes, and activation functions of the DDPG algorithm.
3. The average information age optimization method according to claim 1, characterized in that: In S2, a mobile edge computing communication model is constructed, including: an incident signal sent by a ground user is directionally enhanced by AARIS and then sent to a base station; at the same time, the base station receives an incident signal directly sent from the ground user; the base station decodes the received mixed signal using non-orthogonal multiple access technology and serial interference cancellation technology to obtain a mixed signal, and the steps are as follows: S21. Obtain an equivalent channel gain based on the channel gain of the link between the user and AARIS, the channel gain of the link between the base station and AARIS, and the channel gain of the link between the user and the base station. The formula is: in, is the channel gain between the kth ground user and AARIS, h R (t) is the channel gain of the link between the base station and AARIS; is the direct channel gain between the ground user and the base station in the system model; S22, calculating the mixed user input signal of AARIS and the mixed user amplified signal of AARIS; The formula for the mixed user input signal of AARIS is: in, is the channel gain between the kth ground user and AARIS, p k is the constant transmission power of user k, U k (t) is randomly generated in each time slot through Poisson distribution, which is used to represent whether user k has a communication task in time slot t. k (t) transmit signals for k users; The formula for AARIS's hybrid user amplification signal is: Where x(t) is the mixed user input signal of AARIS, A(t)Θ(t)x(t) is the expected reflected signal, and A(t)Θ(t)x d (t) is the dynamic noise of AARIS, n s It is static noise, which can be ignored compared with dynamic noise; S23. Obtain a mixed user signal after reflection enhancement based on the equivalent channel gain, the mixed user input signal of AARIS, the mixed user amplified signal of AARIS, and the user transmission signal. The formula is: Among them, A(t) is the power amplification factor matrix, Θ(t) is the phase offset matrix, n d (t) represents the additive white Gaussian noise at the base station.
4. The average information age optimization method according to claim 1, characterized in that: The information age model is constructed in S3, including the following steps: S31. Calculate the interference power I between users suffered by user k. k (t), the formula is expressed as: Among them, η jk A 01 identifier, when the value is 1, it means that the signal strength of user j is stronger than that of user k, D j (t) = 1 means that user j decodes successfully, D j (t) = 0 means decoding failure or no decoding; ν∈(0,1) quantifies the information distortion caused by channel state uncertainty and hardware limitations, p j (t) represents the transmission power of other users other than user k; S32. Calculate the current real rate R of user k using the Shannon formula k (t), the formula is expressed as: Among them, A(t) is the power amplification matrix, Θ(t) is the phase offset matrix, r k (t) represents the signal strength of user k, Indicates that when AARIS actively amplifies the signal, it will also amplify the noise power it receives; I k (t) represents the interference signal power received by user k; δ 2 is the noise power; S33. Construct an information age model, the formula is expressed as: Among them, O k (t) represents the lifetime of the data packet of user k in time slot t. If there is a new data packet to be sent (i.e., U k (t)=1), O k (t) is reset to 0, otherwise O k (t) Add 1 to the original value; Δ k (t+1) represents the information age of user k in time slot t+1. When user k successfully transmits the task in time slot t (i.e., R k (t)>=R0, R0 represents the minimum communication rate requirement), S k (t) = 1, the information age of the t+1 time slot is the lifetime of the data packet of the t time slot + 1, otherwise, it is the information age of the t time slot + 1. It should be noted that Δ max Indicates an upper limit on the age of information to prevent it from growing forever.
5. The average information age optimization method according to claim 1, characterized in that: The system energy consumption model is expressed as follows: E f (t)=τP U (v(t)) E c (t)=E f (t)+E i (t) Among them, A(t) is the power amplification factor matrix, Θ(t) is the phase offset matrix, P U (v) is the flight power of the UAV, P0 and P1 are the blade profile power and induced power when the UAV is hovering, respectively; U tip is the rotor tip speed, v0 is the average rotor induced speed; E f (t) is the flight energy consumption of time slot t, τ is the length of a time slot; E i (t) is the energy consumption of active RIS, E c (t) is the sum of the two parts of energy consumption, that is, the total energy consumption of AARIS.
6. The average information age optimization method according to claim 3, 4 or 5, characterized in that: The power amplification factor matrix is expressed as follows: Among them, the diagonal element a m (t) represents the magnification of each reflection unit; The phase offset matrix is expressed as: Among them, the diagonal elements Indicates the phase offset of each reflector unit.
7. The average information age optimization method according to claim 3, 4 or 5, characterized in that: The minimum system average information age objective function and constraint conditions in S4 are expressed as follows: Where (P) is the objective function, is the average information age of the system, Satisfy the UAV's flight speed v(t) and the UAV's horizontal flight angle θ u (t); Signal amplification factor a of each reflection unit of AARIS m (t) and the phase shift angle θ of each reflector unit of AARIS m Minimum value of (t); subject to: C1, C2, C3, C4 and C5 are constraints; C1 is the current real rate R of user k k (t) is greater than the set value R0; C2 is the current remaining energy E of the drone r (t) greater than 0; C3 is the two-dimensional position coordinate q of the UAV u , flight speed v(t) and horizontal flight angle θ m (t) meet the set value; C4 is the signal amplification factor of each reflection unit of AARIS m (t), phase shift angle θ m (t) meet the set value; C5 is the current time slot t, the number of users k and the number of reflection units m of AARIS meeting the set values.
8. The average information age optimization method according to claim 7, characterized in that: The average information age of the system is expressed as: Among them, Δ k (t) represents the information age of user k in time slot t.
Citation Information
Patent Citations
Task awareness and calculation unloading joint optimization method based on information age
CN118349344A
Vehicle-mounted edge computing network scheduling method based on information age
CN119383667A
Task unloading method for air-ground mobile edge computing network system
CN119835696A