An AARIS-based method for optimizing the average information age

By leveraging air-ground collaborative mobile edge computing and AARIS-NOMA technology, the DDPG algorithm was improved, and UAV trajectories and beamforming were optimized. This solved the problems of insufficient network coverage and lack of freshness of computing results in traditional mobile edge computing, resulting in a significant reduction in information age and an improvement in communication efficiency.

CN120456060BActive Publication Date: 2026-01-30ANQING NORMAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510495457.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2026-01-30
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Traditional mobile edge computing networks suffer from insufficient coverage in remote or densely populated areas, high costs, difficulty in responding quickly to changes in user needs, and the freshness of computing results is not effectively guaranteed in real-time application scenarios.

Method used

By introducing air-ground collaborative mobile edge computing, combining aerial active intelligent reflective surfaces (AARIS) and non-orthogonal multiple access (NOMA), improving the deep deterministic policy gradient (DDPG) algorithm, constructing communication models and information age models, and optimizing system performance.

Benefits of technology

It effectively reduces the average information age of the system by 30%-50%, improves information freshness, enhances the adaptability and robustness of the algorithm in complex environments, optimizes resource utilization, extends the drone's endurance, reduces multi-user interference, and improves communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456060B_ABST
    Figure CN120456060B_ABST
Patent Text Reader

Abstract

This invention discloses a method for optimizing the average information age (AoI) of mobile edge computing based on AARIS. The method includes: obtaining a deep hybrid strategy gradient (DHPG) algorithm; constructing a mobile edge computing communication model, an information age model, and a system energy consumption model based on the Airborne Active Intelligent Reflector (AARIS) and Non-Orthogonal Multiple Access (NOMA); and constructing a minimum system average information age objective function and constraints to improve the freshness and timeliness of information in the mobile edge computing communication system. This invention reduces the system's average information age (AoI) by constructing a collaborative communication model between the Airborne Active Intelligent Reflector (AARIS) and NOMA, combined with the deep hybrid strategy gradient (DHPG) algorithm to optimize UAV trajectories and beamforming. This effectively ensures data timeliness in real-time scenarios such as autonomous driving and significantly improves information freshness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of air active intelligent reflecting surface (AARIS) assisted mobile edge computing, and particularly relates to an average information age optimization method based on AARIS mobile edge computing. BACKGROUND

[0002] Mobile edge computing (MEC, Mobile Edge Computing) as a cutting-edge computing mode, its core lies in embedding computing and caching capabilities directly into the network edge adjacent to user equipment, thereby greatly improving service efficiency and reducing delay time. Although traditional computing servers are deployed in ground base stations, they play an indispensable role in MEC networks, but they face several key limitations. The first problem is that the deployment density of ground base stations is relatively limited, which is difficult to fully cover remote areas or user-intensive areas, thereby limiting the effective range of services. Secondly, the construction and maintenance of ground base stations, especially in complex terrain and sparsely populated areas, require huge capital investment, and the cost problem is particularly prominent. Furthermore, ground base stations adopt a fixed deployment strategy, which greatly weakens their flexibility in responding to fluctuations in user demand. In particular, when facing temporary or sudden changes in demand, ground base stations are difficult to quickly adjust to adapt, and cannot provide immediate response. The above limitations pose a serious challenge to the performance and quality of service of traditional mobile edge computing networks.

[0003] To solve the above problems, unmanned aerial vehicles (UAV) are introduced into mobile edge computing networks. Unmanned aerial vehicles can quickly respond to dynamic demand, expand service range and improve service quality due to their high mobility, flexible deployment and low cost. However, with the growing demand for high-performance computing services from users, relying solely on air mobile edge computing has been difficult to meet demand. Therefore, air-ground collaborative mobile edge computing (AGC-MEC) is proposed, which integrates the computing resources of air unmanned aerial vehicles and ground base stations to provide more efficient and flexible computing support for user equipment.

[0004] In air-ground collaborative mobile edge computing networks, performance optimization is of great concern, but current research mostly focuses on traditional performance indicators, such as system delay and energy consumption. However, in application scenarios such as autonomous driving and anomaly detection that require real-time performance, the freshness of the computing results is particularly important. Once the information is outdated, it may not only become invalid, but also may pose a safety hazard. Therefore, the freshness of the computing results has become a key factor in evaluating the performance of air-ground mobile edge computing. To effectively solve this problem, the concept of "information freshness" is introduced, which is the time interval from the generation of the latest data by the Internet of Things terminal to the reception of the data by the information receiver.

[0005] For example: the application number 202411184537.8, discloses an AoI optimization method and system of a multi-unmanned aerial vehicle supported communication and sensing system, belonging to the field of mobile communication. The application scheme minimizes the average information age of ground equipment by jointly optimizing the unmanned aerial vehicle trajectory, user association, target sensing selection, and communication and sensing beamforming. However, the scheme has the following problems: (1) it does not effectively construct a communication model of AARIS and non-orthogonal multiple access (NOMA) cooperation, which is not conducive to ensuring the timeliness of data in real-time scenarios; (2) it does not use the DHPG algorithm, which is not conducive to the adaptability of communication in complex environments. SUMMARY

[0006] In view of the above problems, the purpose of the present application is to provide an average information age optimization method based on AARIS mobile edge computing.

[0007] The present application provides an average information age optimization method based on AARIS mobile edge computing, comprising the following steps:

[0008] S1, improving the deep deterministic policy gradient (DDPG) algorithm to obtain a deep hybrid policy gradient (DHPG) algorithm;

[0009] S2, constructing a mobile edge computing communication model, an information age model and a system energy consumption model based on an air active intelligent reflecting surface (AARIS) and a non-orthogonal multiple access (NOMA);

[0010] S3, based on the mobile edge computing communication model, the information age model and the system energy consumption model, constructing a minimum system average information age objective function and constraint conditions to improve the freshness and timeliness of information in the mobile edge computing communication system.

[0011] Further, the improvement of the DDPG algorithm in S1 includes:

[0012] A hybrid policy mechanism is introduced to combine the advantages of deterministic policy and random policy, thereby improving the adaptability and robustness of the DDPG algorithm in complex environments;

[0013] Optimize the network structure, adjust the number of neural network layers, node number and activation function of DDPG algorithm, enhance the learning efficiency and convergence performance of the algorithm.

[0014] Further, the S2 constructs a mobile edge computing communication model, including the incident signal sent by the ground user is sent to the base station after being enhanced by AARIS; at the same time, the base station receives the incident signal directly sent by the ground user; the base station decodes the received mixed signal by means of non-orthogonal multiple access technology and serial interference cancellation technology, and obtains the mixed signal, the steps are:

[0015] S21, according to the channel gain between the user and AARIS, the channel gain between the base station and AARIS, and the channel gain between the user and the base station, the equivalent channel gain is obtained, the formula is:

[0016]

[0017] Wherein, is the channel gain between the kth ground user and AARIS, h R (t) is the channel gain between the base station and AARIS; is the direct channel gain between the ground user and the base station in the system model;

[0018] S22, calculate the mixed user input signal of AARIS and the mixed user amplification signal of AARIS;

[0019] The mixed user input signal of AARIS is:

[0020]

[0021] Wherein, is the channel gain between the kth ground user and AARIS, p k is the constant transmission power of user k, U k (t) is randomly generated by Poisson distribution in each time slot, which is used to represent whether user k has communication task at time slot t, x k (t) is the transmission signal of k user;

[0022] The mixed user amplification signal of AARIS is:

[0023]

[0024] Wherein, x(t) is the mixed user input signal of AARIS, A(t)Θ(t)x(t) is the expected reflection signal, A(t)Θ(t)x d (t) is the dynamic noise of AARIS, n sFor static noise, compared to dynamic noise can be ignored;

[0025] S23, according to the equivalent channel gain, AARIS mixed user input signal and AARIS mixed user amplification signal and user transmission signal, obtain the reflection enhanced mixed user signal, the formula is:

[0026]

[0027] Wherein, A(t) is the power amplification coefficient matrix, Θ(t) is the phase offset matrix, n d (t) represents the additive white Gaussian noise at the base station.

[0028] Further, the power amplification coefficient matrix, the formula is expressed as:

[0029]

[0030] Wherein, the diagonal element a m (t) represents the amplification of each reflection unit;

[0031] Phase offset matrix, the formula is expressed as:

[0032]

[0033] Wherein, the diagonal element Indicates the phase offset of each reflection unit.

[0034] Further, the S3 in the information age model, including the steps of:

[0035] S31, calculate the user k is affected by the interference power I k (t), the formula is expressed as:

[0036]

[0037] Wherein, η jk A 01 identifier, when the value is 1, it means that the signal strength of user j is stronger than k, D j (t) = 1 indicates that user j decodes successfully, D j (t) = 0 indicates that decoding fails or has not been decoded; ν ∈ (0, 1) quantifies the information distortion caused by channel state uncertainty and hardware limitation, p j (t) represents the transmission power of other users different from user k.

[0038] S32, calculate the current real rate R k (t) of user k by Shannon formula, the formula is expressed as:

[0039]

[0040] wherein r k (t) represents the signal strength of user k, represents that AARIS amplifies the noise power when it amplifies the signal; I k (t) represents the interference signal power received by user k; δ 2 is the noise power.

[0041] S33, constructing an information age model, the formula is represented as:

[0042]

[0043] wherein O k (t) represents the survival time of the data packet of user k at time slot t, if there is a new data packet to be sent (i.e. U k (t) = 1), O k (t) is reset to 0, otherwise O k (t) is increased by 1 on the basis of the original; Δ k (t+1) represents the information age of user k at time slot t+1, when the task of user k at time slot t is successfully transmitted (i.e. R k (t) >= R0, R0 represents the minimum communication rate requirement), S k (t) = 1, the information age at time slot t+1 is the survival time of the data packet at time slot t + 1, otherwise, it is increased by 1 on the basis of the information age at time slot t, it should be noted that Δ max represents the upper limit of the information age to prevent it from increasing all the time.

[0044] Further, the system energy consumption model, the formula is represented as:

[0045]

[0046] E f (t) = τP U (v(t))

[0047]

[0048] E c (t) = E f (t) + E i (t)

[0049] wherein P U (v) is the flight power of the UAV, P0 and P1 are respectively the blade profile power and induced power when the UAV hovers. U tip is the rotor tip speed, v0 is the average rotor induced speed; E f(t) is the flight energy consumption of time slot t, τ is the length of a time slot; E i (t) is the energy consumption representation of active RIS c (t) is the sum of the two parts of energy consumption, that is, the total energy consumption of AARIS.

[0050] Further, the minimum system average information age target function and constraint condition in S4 are expressed as:

[0051]

[0052] Where (P) is the target function, is the system average information age, satisfies the flight speed v(t) of the unmanned aerial vehicle, the horizontal flight angle θ u (t) of the unmanned aerial vehicle; m (t) and the phase shift angle θ m (t) of each reflection unit of AARIS.

[0053] subject to: C1, C2, C3, C4 and C5 are constraint conditions;

[0054] C1 is the real rate R k (t) of user k at present is greater than the set value R0;

[0055] C2 is the current residual energy E r (t) of the unmanned aerial vehicle is greater than 0;

[0056] C3 is the two-dimensional position coordinates q u , flight speed v(t) and horizontal flight angle θ m (t) of the unmanned aerial vehicle satisfy the set value;

[0057] C4 is the signal amplification coefficient a m (t) and the phase shift angle θ m (t) of each reflection unit of AARIS satisfy the set value;

[0058] C5 is that the current time slot t, the number of users k and the number of reflection units m of AARIS satisfy the set value.

[0059] Further, the system average information age is expressed as:

[0060]

[0061] Where Δ k (t) represents the information age of user k in time slot t.

[0062] The beneficial effects of the present application are:

[0063] 1、The present application constructs an air active intelligent reflecting surface (AARIS) and non-orthogonal multiple access (NOMA) cooperative communication model, combines a deep hybrid policy gradient (DHPG) algorithm to optimize the unmanned aerial vehicle trajectory and beamforming, and reduces the system average information age (AoI) by 30%-50% (as shown in Fig. 3-4 ), effectively guarantees the data timeliness in real-time scenarios such as automatic driving, and significantly improves the information freshness.

[0064] 2、The improved DHPG algorithm of the present application combines the advantages of DDPG deterministic policy and PPO random policy, and the convergence speed is improved by 40% compared with the traditional DDPG algorithm, and the mixed action space design (formula: new action = 0.6 x deterministic action + 0.4 x probability action) is used to strengthen the adaptability to complex environment, and effectively enhance the convergence performance of the algorithm.

[0065] 3、The AARIS of the present application actively amplifies the signal to increase the equivalent channel gain, combines the NOMA technology to reduce the multi-user interference, the user real rate meets the QoS constraint, improves the communication efficiency, and optimizes the resource utilization efficiency.

[0066] 4、The energy consumption model of the system of the present application dynamically balances the flight energy consumption and the AARIS reflection energy consumption, prolongs the endurance of the unmanned aerial vehicle under the energy constraint, and effectively controls the energy consumption.

[0067] 5、The active reflecting unit (amplification factor A(t), phase shift Θ(t)) of the AARIS of the present application improves the signal strength compared with the passive RIS, and the air-ground cooperative model realizes interference suppression and noise elimination through equivalent channel joint optimization.

[0068] 6、The present application ensures the stability of the policy under the constraints of position, energy, etc. (C1-C5) through Markov decision process and double reward mechanism. The experiment shows that the AoI fluctuation is reduced by 25% in the burst traffic scenario, and the robustness is enhanced. DETAILED DESCRIPTION

[0069] Fig. 1 It is a flowchart of the average information age optimization method based on AARIS mobile edge computing of the present application;

[0070] Fig. 2 It is a structural diagram of the mobile edge computing communication model of the present application;

[0071] Fig. 3 It is a simulation verification effect diagram of non-orthogonal multiple access and orthogonal multiple access using the method of the present application;

[0072] Fig. 4 It is a simulation verification effect diagram of the present application method with AARIS and without AARIS. DETAILED DESCRIPTION

[0073] Embodiments of the present application are described below in detail with reference to the accompanying drawings, in which the same or similar symbols or signs refer to the same or similar elements or elements having the same or similar functions throughout the drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only for explanation of the present application, and should not be understood as limiting the present application.

[0074] In the air-ground collaborative mobile edge computing network, performance optimization is of great concern, but current researches are mostly focused on traditional performance indicators, such as system delay and energy consumption. However, in application scenarios such as autonomous driving and anomaly detection, which have high real-time requirements, the freshness of the computing results is particularly important. Once the information is outdated, it may not only become invalid, but also may cause safety hazards.

[0075] To solve the above problems, the present application provides an average information age optimization method based on AARIS mobile edge computing, Fig. 1 The flowchart of the average information age optimization method based on AARIS mobile edge computing provided by the embodiments of the present application is shown in the figure, and the method comprises:

[0076] S1, improve the deep deterministic policy gradient (DDPG) algorithm to obtain a deep hybrid policy gradient (DHPG) algorithm.

[0077] Among them, the improvement of DDPG algorithm includes: introducing a hybrid policy mechanism, by combining the advantages of deterministic policy and random policy, to improve the adaptability and robustness of DDPG algorithm in complex environment; optimizing the network structure, by adjusting the number of layers, the number of nodes and the activation function of DDPG algorithm neural network, to enhance the learning efficiency and convergence performance of the algorithm.

[0078] The traditional DDPG algorithm outputs a deterministic action (a1, a2, a3,...a N ), in order to increase the exploration efficiency of the algorithm, it is usually necessary to add random noise to the action to improve the exploration efficiency of the agent in the algorithm, such as (a1+δ1, a2+δ2, a3+δ3, a N +δ N ), where δ is a random noise. The demand work is to introduce the output of the PPO algorithm based on the probability policy strategy on the basis of the deterministic action. The PPO algorithm outputs actions in a different way from the DDPG algorithm. The PPO algorithm outputs the mean and standard deviation of N actions through the Actor network to form a probability distribution, and obtains the action (a'1, a'2, a'3,..., N a)' by random sampling. The advantage of randomness is utilized, and the deterministic action of DDPG and the probability action of PPO are weighted to form a hybrid strategy.

[0079] Specifically, two algorithms are used to generate actions simultaneously, and then new actions are generated for the agent to execute in the following way.

[0080] The new mixed action = weight 1 x deterministic action + weight 2 x probability action, and the selection of the weight is determined as 0.6 and 0.4 through a large number of simulations. In this way, the new action has more purpose than the original DDPG algorithm which introduces random noise exploration, so as to reflect faster convergence speed in the performance of the algorithm. Therefore, the mixed strategy mechanism is introduced to combine the advantages of deterministic strategy and random strategy, and to improve the adaptability and robustness of DDPG algorithm in complex environment.

[0081] S2, based on air active intelligent reflecting surface AARIS and non-orthogonal multiple access NOMA, a mobile edge computing communication model, an information age model and a system energy consumption model are constructed.

[0082] Among them, the mobile edge computing communication model structure is as shown in Fig. 2 The UAV and the intelligent reflecting panel are connected together through components to form an air active intelligent reflecting surface AARIS, share the same control system and power supply; the AARIS adjusts the signal amplification coefficient and the phase offset of the reflecting unit through the controller to directionally enhance the incident signal sent by the ground user and send it to the base station; the ground base station receives the signal sent directly from the ground user and through the air active intelligent reflecting surface; the ground user randomly generates a sending task.

[0083] According to the position information of the ground user and the AARIS, the channel gain of the link between the user and the AARIS is obtained, which is expressed by the formula:

[0084]

[0085] Among them, h k represents the channel gain between the kth ground user and the AARIS, L represents the large-scale fading caused by the path loss and shadowing of the link; n represents the small-scale fading factor of the link.

[0086] According to the position information of the ground base station and the AARIS, the channel gain of the link between the base station and the AARIS is obtained;

[0087]

[0088] Among them, h R (t) represents the channel gain of the link between the AARIS and the base station, g R(t) represents the large-scale fading caused by path loss and shadowing of the link; represents the small-scale fading factor of the link.

[0089] The channel gain of the link between the user and the base station is obtained according to the position information of the user and the ground base station;

[0090]

[0091] wherein, represents the direct channel gain between the ground user and the base station in the system model, f k (t) is the Rayleigh attenuation coefficient, and β represents a path loss constant, is a communication distance, and α is a path loss exponent, represents a lognormal shadow fading, and σ is a standard deviation of the shadow fading.

[0092] The equivalent channel gain is obtained according to the channel gain of the link between the user and the AARIS, the channel gain of the link between the base station and the AARIS, and the channel gain of the link between the user and the base station, and the formula is represented as:

[0093]

[0094] wherein, represents the channel gain between the kth ground user and the AARIS, h R (t) is the channel gain of the link between the base station and the AARIS; represents the direct channel gain between the ground user and the base station in the system model.

[0095] The power amplification coefficient matrix and the phase offset matrix are obtained according to the amplitude coefficient and the phase shift coefficient of the intelligent reflecting surface element; wherein the power amplification coefficient matrix is represented by the formula:

[0096]

[0097] The diagonal element a m (t) represents the amplification multiple of each reflecting unit.

[0098] The phase offset matrix is represented by the formula:

[0099]

[0100] The diagonal element represents the phase offset of each reflecting unit.

[0101] Then, the mixed user input signal of the AARIS and the mixed user amplified signal of the AARIS are calculated.

[0102] The mixed user input signal of AARIS is formula:

[0103]

[0104] Wherein, is the channel gain between the kth ground user and AARIS, p k is the constant transmit power of user k, U k (t) is randomly generated by Poisson distribution in each time slot, which is used to represent whether user k has a communication task at time slot t, x k (t) is the transmission signal of k user;

[0105] The mixed user amplification signal of AARIS is formula:

[0106]

[0107] Wherein, x(t) is the mixed user input signal of AARIS, A(t)Θ(t)x(t) is the expected reflection signal, A(t)Θ(t)x d (t) is the dynamic noise of AARIS, n s is static noise, which can be ignored compared with dynamic noise;

[0108] Further, according to the equivalent channel gain, the base station receives the air channel and the signal transmitted by the direct connection from AARIS, so the mixed signal received at the base station is formula:

[0109]

[0110] Wherein, A(t) is the power amplification coefficient matrix, Θ(t) is the phase offset matrix, n d (t) represents the additive white Gaussian noise at the base station, represents the interference signal from other users compared with user k (represented by symbol j).

[0111] Through the calculation of the above formula, the mobile edge computing communication model structure can be constructed as shown in Fig. 2 .

[0112] Further, the information age model is constructed.

[0113] Firstly, the interference power I k (t) suffered by user k is calculated, which is formula:

[0114]

[0115] Wherein, η jkAn 01 identifier, when the value is 1, indicates that the signal strength of user j is stronger than k, D j (t) = 1 indicates that user j decodes successfully, D j (t) = 0 indicates decoding failure or has not been decoded; v e (0, 1) quantifies the information distortion caused by channel state uncertainty and hardware limitations, p j (t) represents the transmission power of other users different from user k.

[0116] Then, the current real rate R k (t) of user k is calculated by Shannon formula combined with the interference power I k (t) received by user k from other users, the formula is:

[0117]

[0118] Where, r k (t) represents the signal strength of user k, represents that when AARIS actively amplifies the signal, it will amplify the noise power received by itself; I k (t) represents the interference signal power received by user k; δ 2 is the noise power.

[0119] Then, the information age model is constructed, the formula is:

[0120]

[0121] Where, O k (t) represents the survival time of the data packet of user k at time slot t, if there is currently a new data packet to be sent (i.e. U k (t) = 1), O k (t) is reset to 0, otherwise O k (t) is increased by 1 based on the original value.

[0122] Δ k (t+1) represents the information age of user k at time slot t+1, when the task of user k at time slot t is successfully transmitted (i.e. R k (t) >= R0, R0 represents the minimum communication rate requirement), S k (t) = 1, the information age at time slot t+1 is the survival time of the data packet at time slot t + 1, otherwise, it is + 1 based on the information age at time slot t, it should be noted that Δ max represents the upper limit of information age to prevent it from growing indefinitely.

[0123] Finally, the system average information age is calculated according to the information age Δ k (t) of user k at time slot t, the formula is:

[0124]

[0125] where, Δ k (t) denotes the information age of user k at time slot t.

[0126] Further, the system energy consumption model is constructed, and the formula is expressed as:

[0127]

[0128] E f (t) = τP U (v(t))

[0129]

[0130] E c (t) = E f (t) + E i (t)

[0131] where, P U (v) is the flight power of the UAV, P0 and P1 are the blade profile power and induced power when the UAV hovers, respectively. tip is the rotor tip speed, v0 is the average rotor induced speed; E f (t) is the flight energy consumption at time slot t, τ is the length of a time slot; E i (t) is expressed as the energy consumption of the active RIS, E c (t) is the sum of the two parts of energy consumption, that is, the total energy consumption of the AARIS.

[0132] where, the formulas for obtaining P0 and P1 are:

[0133]

[0134] The parameters δ, Ω, R, W are the profile drag coefficient, blade angular velocity, rotor radius and aircraft weight, respectively. d0, ρ, s and A are the fuselage drag ratio, air density, rotor solidity and rotor disc area, respectively.

[0135] S3, based on the mobile edge computing communication model, information age model and system energy consumption model, the minimum system average information age objective function and constraint condition are constructed, and the freshness and timeliness of information in the mobile edge computing communication system are improved.

[0136] The expression of the minimum system average information age optimization problem is:

[0137]

[0138] where, (P) is the objective function, to system average information age, subject to flight speed v(t) of UAV, horizontal flight angle θ u (t) of AARIS, signal amplification coefficient a m (t) of each reflection unit of AARIS, and phase shift angle θ m (t) of each reflection unit of AARIS.

[0139] subject to: C1, C2, C3, C4 and C5 are constraints;

[0140] C1 is the current real rate R k (t) of user k, which is greater than the set value R0, defining the requirement of the minimum real rate (QoS) of each user.

[0141] C2 is the current residual energy E r (t) of UAV, which is greater than 0, defining the energy consumption requirement of UAV.

[0142] C3 is the constraint adjustment of aerial active RIS, including two-dimensional position coordinates q u of UAV, satisfying x u ∈[x min ,x max ], y u ∈[y min ,y max ]; flight speed v(t) satisfies the set value s(t) = q u (t), E r (t), x(t), u(t), horizontal flight angle θ u (t) satisfies [0, 2π].

[0143] C4 is the signal amplification coefficient a m (t) and phase shift angle θ m (t) of each reflection unit of AARIS, which satisfies the set value; wherein L and V max are the maximum amplification and the maximum speed of UAV, respectively.

[0144] C5 is the set value of current time slot t, user number k and reflection unit number m of AARIS.

[0145] Further, the UAV and AARIS interact with the environment to obtain state information and find the optimal strategy to maximize the cumulative reward; the state space, action space and reward function of the formulated Markov decision process are defined as follows:

[0146] State space, at the current time slot t, the state space is defined by several key elements. Which includes the horizontal coordinates q u (t) of UAV, the current residual energy Er (t), in addition, it also includes the information age of each ground user Δ(t) and the task arrival state u(t). The state space can be represented as:

[0147] s(t) = q u (t), E r (t), Δ(t), u(t)

[0148] Action space, at each time slot t, the action of the agent includes the flight direction of the UAV, the flight speed of the UAV, the signal amplification coefficient matrix and the phase shift matrix of the active smart reflecting surface. The action space can be represented as:

[0149] a(t) = v u (t), θ u (t), A(t), Θ(t)

[0150] Reward function, the reward obtained by making action a in state s(t) at time t can be defined as:

[0151]

[0152] where r0 represents a fixed reward value, used to encourage the agent to explore actions in the early stage of learning. w and b are the weighting coefficients of the average AoI and energy consumption in the reward function, respectively. and The additional negative penalty term and the reward term can be represented as:

[0153]

[0154] The algorithm execution process includes:

[0155] Initialize the related parameters of the six neural networks and the parameters in the simulation environment.

[0156] The agent obtains the current state information s(t) from the environment, inputs the DDPG main actor network and the PPO-based actor network to obtain two actions deterministic action a d and probability action a p Then the two actions are weighted and fused.

[0157] The fused action is executed in the simulation environment to obtain the reward value r(t), the next state information s(t+1) and the training termination symbol done.

[0158] The current state, current action, next state, current reward and training termination symbol are stored in the buffer.

[0159] When the number of experiences in the buffer is greater than a batch size, start training and update the network parameters.

[0160] The target Q-value in DDPG is calculated using the Bellman equation, which estimates the expected cumulative reward for the next state and action. The Bellman equation is expressed as follows:

[0161] y(t) = r(t) + γ1Q'(s(t+1), μ'(s(t+1)) | θ Q′ )

[0162] where r(t) is the current reward obtained after performing action a(t) in state s(t), γ1∈[0,1] is the discount factor; Q'(s(t+1), μ'(s(t+1)) | θ Q′ ) represents the target Q-value estimated by the target critic network, while Q(s(t), μ(s(t)) | θ Q ) represents the Q-value predicted by the current network in the DDPG module.

[0163] The critic network update of DDPG is achieved by minimizing the following loss function, which optimizes the critic network parameters θ Q through gradient descent.

[0164]

[0165] where N b represents the size of the batch, i.e., the number of experiences extracted from the replay buffer D at each update. Q(s(t), a d (t) | θ Q ) represents the value estimated by the current critic network. By minimizing this loss function, the parameters θ Q of the DDPG critic network are updated, thereby improving its accuracy in predicting the value of state-action pairs.

[0166] The loss function loss of the critic network of PPO is defined as:

[0167]

[0168] where V(s(t) | θ V ) is the state value estimated by the PPO critic network. By minimizing this loss function, the parameters θ V of the PPO critic network are adjusted, thereby providing a more accurate value assessment for each state.

[0169] After the critic network update, the deterministic actor network (DDPG-actor) and the auxiliary actor network (PPO-actor) are also updated. The deterministic actor network is optimized through policy gradient, as follows:

[0170]

[0171] where, is the gradient of Q value with respect to deterministic action a d (t). This gradient is used to update the parameters of the deterministic actor network θ μ .

[0172] The auxiliary actor network following the PPO framework is also updated according to the clip policy to ensure the stability of learning. The gradient of the auxiliary actor network update is:

[0173]

[0174] where, represents the ratio of the probability of the new policy and the old policy, represents the estimated advantage function, and ∈ D represents the truncation parameter. Here, the time difference TD error is used to approximate the advantage function

[0175] where the ratio of the new and old policies is calculated as follows: r o (π) = exp(lp(t) - lp old )

[0176] The advantage function in the PPO module can be defined as:

[0177]

[0178] In addition, the target network in the DDPG algorithm obtains updated parameters from the main network through soft updating.

[0179] θ Q′ ← τ D θ Q + (1 - τ D ) θ Q′ , θ μ′ ← τ D θ μ + (1 - τ D ) θ μ′

[0180] where τ D is the soft update factor, usually a small value, ensuring gradual adjustment of the target network parameters. This method reduces instability in the training process by promoting smoother updates.

[0181] The core part of DDPG includes an actor network μ(s(t)|θ μ for policy optimization and a critic network Q(s(t),a(t)|θ Q), and their respective target networks μ'(s(t)|θ μ' ) and Q'(s(t),a(t)|θ Q ) to stabilize the training process. The auxiliary module is based on PPO, including an actor network π(s(t)|θ π ) for enhanced exploration and a critic network V(s(t)|θ V ) for value function approximation.

[0182] At each time slot t, the agent perceives the current state S(t) of the NOMA and AARIS-aided MEC environment, which is defined by the formula S(t) = (s(t), a(t-1), a(t-2),..., a(t-n), lp(t-1), lp(t-2),..., lp(t-n), r(t-1), r(t-2),..., r(t-n), s(t-1), s(t-2),..., s(t-n), done). Then S(t) is input into μ(s(t)|θ μ ) and π(s(t)|θ π ) to generate the hybrid action a(t). Subsequently, the agent performs a h (t) in the current environment, obtains the reward r(t) and moves to the next state s(t+1). The experience tuple (s(t), a h (t), a d (t), a p (t), lp(t), r(t), s(t+1), done) is then stored into the replay buffer D.

[0183] Fig. 3 The advantage of using non-orthogonal multiple access technology in the mobile edge computing network assisted by aerial active intelligent reflecting surfaces is verified, and we compare the performance of non-orthogonal multiple access and orthogonal multiple access in this network. The traditional DDPG algorithm and the improved DHPG algorithm are used to jointly optimize the flight trajectory of the UAV and the beamforming of the active intelligent reflecting surface. First of all, it can be observed that the application of non-orthogonal multiple access makes the average information age in the network lower than that of orthogonal multiple access, whether using the traditional DDPG algorithm or the improved DHPG algorithm, which shows that non-orthogonal multiple access technology can improve the network performance in the current multi-user situation, thereby effectively reducing the average information age in the system. Secondly, we can see that the performance of the improved DHPG algorithm is better than that of the traditional DDPG algorithm, whether using non-orthogonal multiple access or orthogonal multiple access, showing faster convergence effect. Therefore, non-orthogonal multiple access achieves better performance in the active intelligent reflecting surface network, demonstrating the superior adaptability of non-orthogonal multiple access to aerial active intelligent reflecting surface networks.

[0184] Fig. 4It is illustrated that the influence of active intelligent reflecting surface and passive intelligent reflecting surface on the experimental results in the mobile edge computing network assisted by non-orthogonal multiple access technology. It can be seen that whether it is a traditional method or using an improved method, the active intelligent reflecting surface has better effect on the optimization of information age in the network compared with the passive case. This is because the active intelligent reflecting surface has the function of signal amplification, which can amplify the incident signal in a directional manner, thereby improving the efficiency of data transmission and achieving the purpose of reducing the average information age. Similarly, whether it is an active intelligent reflecting surface or a passive intelligent reflecting surface, the improved algorithm is superior to the traditional algorithm in terms of convergence speed. Therefore, compared with the passive intelligent reflecting surface, the active intelligent reflecting surface can achieve better performance in the non-orthogonal multiple access assisted air-ground mobile edge computing network.

[0185] Reconfigurable Intelligent Surface (RIS) as an innovative technology in future wireless communication systems, combined with Unmanned Aerial Vehicles (UAV) and Mobile Edge Computing (MEC), brings revolutionary changes to the field of wireless communication. Through the fine control of intelligent reflecting surface, the propagation path of electromagnetic wave is flexibly optimized, thereby realizing the enhancement of signal strength, effective suppression of interference and significant improvement of energy efficiency. Further, the active intelligent reflecting surface (AARIS) is combined with unmanned aerial vehicle (UAV) technology, and the high mobility of unmanned aerial vehicle and the precise control ability of AARIS on signal propagation path are utilized to innovatively construct an aerial active intelligent reflecting surface (AARIS). This combination can build an aerial channel in the mobile edge computing (MEC) system, bringing significant improvement of service quality (QoS) to the MEC system, thereby greatly enhancing the timeliness of computing results and the freshness of data. The complexity and randomness of the communication environment and the traditional method face the double challenges of unmanned aerial vehicle flight trajectory planning and AARIS beamforming adjustment when facing the control of the aerial intelligent reflecting surface (especially combined with unmanned aerial vehicle and AARIS). Therefore, we introduce deep reinforcement learning (DRL) as a solution. Under the framework of deep reinforcement learning, we take the real-time state of the system (including the position of the unmanned aerial vehicle, the task generation of the ground equipment, the energy consumption state, and the information age of each ground equipment) as the input information. These information are extracted by deep neural network (DNN) for high-dimensional feature extraction, and then the long-term benefits brought by different action strategies (such as adjustment of unmanned aerial vehicle flight path and adjustment of active intelligent reflecting surface beamforming) are predicted.

[0186] With the trial-and-error learning mechanism of reinforcement learning, our system can continuously collect feedback during actual operation, dynamically adjust policy parameters, and gradually approach the optimal control strategy. In addition, we innovatively integrate two mainstream deep reinforcement learning algorithms, significantly improving the performance of a single algorithm in jointly optimizing the trajectory and beamforming of the unmanned aerial vehicle, achieving more efficient and accurate control of the intelligent reflecting surface in the air.

[0187] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not limiting; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions described in the foregoing examples can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for average information age optimization based on AARIS mobile edge computing, characterized in that, The method comprises the following steps: S1, improving a deep deterministic policy gradient (DDPG) algorithm to obtain a deep hybrid policy gradient (DHPG) algorithm; The improvement of the DDPG algorithm comprises the following steps: introducing a hybrid policy mechanism to combine the advantages of deterministic policy and stochastic policy, thereby improving the adaptability and robustness of the DDPG algorithm in a complex environment; optimizing the network structure by adjusting the number of neural network layers, the number of nodes and the activation function of the DDPG algorithm, thereby enhancing the learning efficiency and convergence performance of the algorithm; S2, constructing a mobile edge computing communication model, an information age model and a system energy consumption model based on an air-based active intelligent reflecting surface (AARIS) and a non-orthogonal multiple access (NOMA); The mobile edge computing communication model is constructed, including that the incident signal transmitted by the ground user is transmitted to the base station after being directionally enhanced by the AARIS; meanwhile, the base station receives the incident signal directly transmitted by the ground user; the base station decodes the mixed signal received by means of the non-orthogonal multiple access technology and the serial interference cancellation technology to obtain the mixed signal, and the steps are as follows: S21, obtaining the equivalent channel gain according to the channel gain between the user and the AARIS, the channel gain between the base station and the AARIS and the channel gain between the user and the base station; S22, calculating the mixed user incident signal of the AARIS and the mixed user amplification signal of the AARIS; S23, obtaining the mixed user signal after reflection enhancement according to the equivalent channel gain, the mixed user incident signal of the AARIS, the mixed user amplification signal of the AARIS and the user transmission signal; S3, constructing a minimum system average information age target function and a constraint condition based on the mobile edge computing communication model, the information age model and the system energy consumption model, thereby improving the freshness and timeliness of information in the mobile edge computing communication system.

2. The average information age optimization method of claim 1, wherein In the S21, the equivalent channel gain is obtained, and the formula is as follows: ; wherein, is the channel gain between the kth ground user and the AARIS, is the channel gain of the link between the base station and the AARIS; is the direct channel gain between the ground users and the base station in the system model; In the step S22, the mixed user incident signal of the AARIS is calculated, and the formula is as follows: ; wherein, is the channel gain between the kth ground user and the AARIS, is the constant transmit power of user k, is randomly generated by a Poisson distribution in each time slot, used to represent whether user k has a communication task at time slot t, is the kth user transmission signal; In the step S22, the mixed user amplification signal of the AARIS is calculated, and the formula is as follows: ; wherein, is a mixed user-in signal for AARIS, is a desired reflected signal, is a dynamic noise for AARIS, is a static noise, which is negligible compared to the dynamic noise; In the step S23, the mixed user signal after reflection enhancement is obtained, and the formula is as follows: ; wherein is a power amplification coefficient matrix, is a phase offset matrix, denotes additive white Gaussian noise at the base station.

3. The average information age optimization method of claim 1, wherein In the S3, the information age model is constructed, and the steps are as follows: S31, calculate the interference power between users suffered by user k The formula is: ; wherein, a 01 identifier, when the value is 1, indicates that the signal strength of user j is stronger than k, indicates that user j decodes successfully, indicates that decoding fails or has not been decoded; quantifies the information distortion caused by channel state uncertainty and hardware limitations, indicates the transmission power of other users different from user k; S32, calculate the real rate of user k at present by Shannon formula The formula is represented as: ; wherein, is a power amplification coefficient matrix, is a phase offset matrix, denotes the signal strength of user k, denotes the noise power amplified by the AARIS active amplification of the signal; denotes the interference signal power experienced by user k; is the noise power; S33, constructing the information age model, and the formula is as follows: ; ; Wherein, represents the survival time of the data packet of user k at time slot t, if there is a new data packet to be sent (i.e. ), is reset to 0, otherwise is increased by 1 on the basis of the original value; represents the information age of user k at time slot t+1, when the task of user k at time slot t is successfully transmitted (i.e. > , represents the minimum communication rate requirement, the information age at time slot t+1 is the survival time of the data packet at time slot t+1, otherwise, it is the basis of the information age at time slot t+1+1, it should be noted that represents the upper limit of the information age to prevent it from always increasing.

4. The average information age optimization method of claim 1, wherein The system energy consumption model is as follows: ; ; ; ; where, is the power amplification coefficient matrix, is the phase offset matrix, is the UAV flight power, and are the blade profile power and induced power of the UAV in hover, respectively; is the rotor tip speed, is the average rotor induced speed; is is the flight energy consumption of a time slot, is the length of a time slot; is the energy consumption representation of the active RIS, is the sum of the two parts, i.e., the total energy consumption cost of the AARIS.

5. The average information age optimization method of any one of claims 2-4, wherein, The power amplification coefficient matrix is as follows: ; where the diagonal elements represent the magnification of each reflecting unit; The phase offset matrix is as follows: ; where the diagonal elements represent the phase shift of each reflection unit.

6. The average information age optimization method of any one of claims 2-4, wherein, In the S3, the minimum system average information age target function and the constraint condition are as follows: ; Wherein, is the target function, is the system average information age, satisfies the flight speed of the unmanned aerial vehicle , the horizontal flight angle of the unmanned aerial vehicle ; the signal amplification coefficient of each reflection unit of the AARIS and the phase offset angle of each reflection unit of the AARIS the minimum value of , , , and are the constraints; the current true rate for user k greater than a set value ; current remaining energy for the drone greater than 0; two-dimensional position coordinates of the drone flight speed and horizontal flight angle satisfy the set value; signal amplification coefficient for each reflection unit of the AARIS , phase offset angle satisfies a set value; For the current time slot t, the number of users k, and the number of reflection units m of AARIS, the set values are satisfied.

7. The average information age optimization method of claim 6, wherein, The system average information age is as follows: ; wherein, denotes the information age of user k at time slot t.

Citation Information

Patent Citations

  • AoI optimization method and system of communication and perception system supported by multiple unmanned aerial vehicles

    CN118945693A

  • Vehicle-mounted edge computing network scheduling method based on information age

    CN119383667A

  • Task unloading method for air-ground mobile edge computing network system

    CN119835696A