Information age optimization method based on intelligent reflecting surface

By introducing drones equipped with intelligent reflection surfaces and deep reinforcement learning algorithms in the Internet of Things communication system, the scheduling and transmission power of IoT devices are optimized, and the problem of delay in information update of IoT devices is solved, thus minimizing information age and improving information freshness is achieved.

CN120165727APending Publication Date: 2025-06-17SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201401.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When IoT devices use drone relay to transmit update information, there is a delay in update information, resulting in poor information freshness. How to minimize the information age of the communication system has become an urgent problem.

Method used

By introducing a drone equipped with intelligent reflection surfaces as a relay node, combined with deep reinforcement learning algorithms, the scheduling, transmission power and intelligent reflection surface phase shift of IoT devices are jointly optimized, and an information age optimization model is constructed to minimize the long-term average information age.

Benefits of technology

It effectively reduces the information age of the Internet of Things communication system, improves the freshness and availability of information, and improves the overall performance and communication efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165727A_ABST
    Figure CN120165727A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Internet of Things communication, in particular to an information age optimization method based on an intelligent reflecting surface, which comprises the following steps: constructing a system model which comprises an unmanned aerial vehicle carrying the intelligent reflecting surface, a base station, K Internet of Things devices and a ground control station, wherein a virtual link is established between the Internet of Things equipment and the base station through an intelligent reflecting surface carried by the unmanned aerial vehicle, and the ground control station is used for dispatching the Internet of Things equipment to transmit information and controlling the displacement of the intelligent reflecting surface and the transmitting power of the Internet of Things equipment; constructing an information age optimization model, wherein the optimization objective is to minimize the long-term average information age; and solving the constructed information age optimization model by using a deep reinforcement learning algorithm based on an SD3 algorithm to obtain an optimal Internet of Things equipment scheduling strategy, transmitting power and intelligent reflecting surface shift. According to the invention, the information age of the Internet of Things communication system is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things communications, and in particular to an information age optimization method based on an intelligent reflective surface. Background Art

[0002] Today is an era of the Internet of Everything. With the rapid development of smart cities and IoT technologies, higher requirements are placed on the reliability and timeliness of device status updates. In applications such as intelligent transportation, industrial control systems, and environmental monitoring, the updated data generated by IoT devices is an important basis for decision-making. However, information transmission delays and the uncontrollability of the network environment will affect the timeliness of data updates, leading to wrong decisions. Therefore, in order to ensure the efficient operation of the IoT system, a key performance indicator, information age, is proposed to measure the freshness of information. For IoT applications with real-time transmission requirements, the freshness of information is crucial.

[0003] When it is difficult to obtain a strong direct line-of-sight communication link due to the limited capabilities of IoT devices and environmental obstacles, aerial unit drones are often used as relays to transmit information. However, there is a delay in updating the status information of IoT devices using drone relays, which makes the information update not timely and the information freshness is poor. How to minimize the information age of the communication system has become an urgent problem to be solved. Summary of the invention

[0004] The purpose of the present invention is to provide an information age optimization method based on an intelligent reflective surface, which minimizes the long-term average information age of the Internet of Things communication system and ensures the freshness of the information by reasonably planning the scheduling of Internet of Things devices, combining the phase shift of the intelligent reflective surface elements and the transmission power control of the Internet of Things devices.

[0005] The first aspect of the present invention provides an information age optimization method based on a smart reflective surface, comprising the following steps:

[0006] Construct a system model, including a UAV equipped with a smart reflective surface, a base station, K IoT devices, and a ground control station. The smart reflective surface has a total of F reflective elements. A virtual link is established between the IoT device and the base station through the smart reflective surface carried by the UAV. The ground control station is used to schedule the IoT device to transmit information and control the phase shift of the smart reflective surface and the transmission power of the IoT device.

[0007] Construct an information age optimization model, the optimization goal is to minimize the long-term average information age;

[0008] The constructed information age optimization model is solved using a deep reinforcement learning algorithm based on the SD3 algorithm to obtain the optimal IoT device scheduling strategy, transmission power, and smart reflector phase shift.

[0009] A further improvement lies in that the constructed information age optimization model has an optimization objective of minimizing the long-term average information age, specifically:

[0010]

[0011] The constraint conditions are:

[0012] φ F [n] ∈ [0, 2π)

[0013] α k [n] ∈ [0, 1]

[0014] 0 < P k [n] < P max

[0015] Among them, D(n) represents the transmit power control of the Internet of Things device, C(n) represents the information transmission scheduling of the Internet of Things device, Θ(n) represents the reflection phase shift of the intelligent reflecting surface, K represents the number of Internet of Things devices, N represents the service duration and is evenly divided into N time slots, n ∈ [0, N], A k [n represents the information age of Internet of Things device k at time slot n, φ F [n] represents the phase shift of the reflection element of the intelligent reflecting surface, α k [n] represents the scheduling situation of the Internet of Things device. When α k [n] = 1, it means that Internet of Things device k is scheduled at time slot n. When α k [n] = 0, it means that Internet of Things device k is not scheduled at time slot n. P k [n] represents the transmit power of the Internet of Things device, and P max represents the maximum transmit power of the Internet of Things device.

[0016] A further improvement lies in that the activation mode of the Internet of Things device follows a uniform distribution.

[0017] A further improvement lies in that the constructed information age optimization model is solved using a deep reinforcement learning algorithm based on the SD3 algorithm, specifically:

[0018] The optimization problem is formulated as a Markov decision process, which consists of the tuple <s, a, r>. s represents the state, a represents the action, and r represents the reward function. In each training set, the agent observes the current state s(t), then selects an action a(t) to execute. Once the action is selected, the agent will receive the corresponding reward r(t) and continue to observe the state s(t + 1) in the next time slot, where:

[0019] State space: The state of the system consists of the state of the UAV and the state of the IoT devices, denoted as s[n] = (A[n], γ[n]), where A[n] represents the age of information of each IoT device at time slot n, and γ[n] represents the signal-to-noise ratio when the signals transmitted by the IoT devices at time slot n are coherently combined through the phase of the intelligent reflecting surface elements;

[0020] Action space: The actions of the system include two aspects, namely the scheduling of IoT devices and the transmit power control of IoT devices, denoted as a[n] = (ξ[n], Ρ[n]), where ξ[n] represents the scheduling vector of the k IoT devices by the ground control station at time slot n, and Ρ[n] represents the transmit power of the scheduled IoT devices at time slot n;

[0021] Reward function: The reward function is defined as the negative of minimizing the sum of the ages of information of all IoT devices:

[0022]

[0023] where AoI k(n) represents the age of information of the k-th IoT device at time slot n.

[0024] Furthermore, the improvement lies in that the age of information A k [n] of IoT device k at time slot n evolves in the next time slot as:

[0025]

[0026] where, G k [n] is a binary variable indicating whether the k-th IoT device is active in time slot n. When G k [n] = 1, it means the IoT device is active, and when G k [n] = 0, it means the IoT device is in sleep; γ th is the minimum threshold to ensure reliable decoding, and γ k is the signal-to-noise ratio of the base station at time slot n.

[0027] Furthermore, the improvement lies in that each IoT device and the base station are equipped with a transmit antenna and a receive antenna, and the non-payload information exchanged between the IoT device and the base station via the UAV uplink with an intelligent reflecting surface is only the channel state information.

[0028] Furthermore, the improvement also includes the following steps: According to the optimal IoT device scheduling strategy, transmit power, and intelligent reflecting surface phase shift, the ground control station arranges the IoT devices to perform data packet transmission in each time slot, adjusts the intelligent reflecting surface element phase shift, and sets the IoT device transmit power.

[0029] In the second aspect of the present invention, an information age optimization system based on intelligent reflecting surface is proposed, including the following modules:

[0030] A system model construction module for constructing a system model, which includes a drone carrying an intelligent reflecting surface, a base station, K Internet of Things (IoT) devices, and a ground control station. The intelligent reflecting surface has a total of F reflecting elements. Among them, a virtual link is established between the IoT devices and the base station through the intelligent reflecting surface carried by the drone. The ground control station is used to schedule the IoT devices to transmit information and control the phase shift of the intelligent reflecting surface and the transmission power of the IoT devices.

[0031] An optimization model construction module for constructing an information age optimization model, with the optimization objective of minimizing the long-term average information age.

[0032] A solving module for solving the constructed information age optimization model by using a deep reinforcement learning algorithm based on the SD3 algorithm to obtain the optimal IoT device scheduling strategy, transmission power, and phase shift of the intelligent reflecting surface.

[0033] In the third aspect of the present invention, an electronic device is proposed, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements an information age optimization method based on intelligent reflecting surface as described in any one of the first aspects.

[0034] In the fourth aspect of the present invention, a computer-readable storage medium is proposed. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute an information age optimization method based on intelligent reflecting surface as described in any one of the first aspects.

[0035] The beneficial effects of the present invention are as follows:

[0036] By introducing a drone carrying an intelligent reflecting surface as a relay node and combining a deep reinforcement learning algorithm, the present invention jointly optimizes the scheduling, transmission power, and phase shift of the intelligent reflecting surface of the IoT devices, effectively reducing the information age of the IoT communication system, improving the freshness and availability of information, and enhancing the overall performance and communication efficiency of the system. At the same time, the deep reinforcement learning method based on the SD3 algorithm can adapt to the randomness of the activation mode of IoT devices and the complexity of the optimization problem, and find a near-optimal control strategy. Description of the Drawings

[0037] Figure 1 It is a schematic diagram of the system model;

[0038] Figure 2 It is a flowchart of an information age optimization method based on intelligent reflecting surface;

[0039] Figure 3 It is the convergence graph of the SD3 algorithm;

[0040] Figure 4 It is the comparison graph of the long-term average age of information under different numbers of Internet of Things devices;

[0041] Figure 5 It is the comparison graph of the long-term average age of information under different sizes of distribution areas of Internet of Things devices;

[0042] Figure 6 It is the comparison graph of the long-term average age of information under different numbers of intelligent reflecting surface elements;

[0043] Figure 7 It is the comparison graph of the long-term average age of information under different activation probabilities;

[0044] Figure 8 It is a schematic diagram of an electronic device. Specific implementation manners

[0045] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0047] Please refer to the attached Figure 1 - attached Figure 8 , the first aspect of the embodiment of the present invention proposes an information age optimization method based on an intelligent reflecting surface, including the following steps:

[0048] Step S1: Construct a system model, which includes a drone equipped with an intelligent reflecting surface, a base station, K Internet of Things (IoT) devices, and a ground control station. The intelligent reflecting surface has a total of F reflecting elements. Among them, a virtual link is established between the IoT devices and the base station through the intelligent reflecting surface carried by the drone. The ground control station is used to schedule the IoT devices to transmit information and control the phase shift of the intelligent reflecting surface and the transmission power of the IoT devices.

[0049] Step S2: Construct an age-of-information optimization model, and the optimization objective is to minimize the long-term average age of information.

[0050] Step S3: Use a deep reinforcement learning algorithm based on the SD3 algorithm to solve the constructed age-of-information optimization model, and obtain the optimal IoT device scheduling strategy, transmission power, and phase shift of the intelligent reflecting surface.

[0051] The embodiments of the present invention are introduced in more detail below:

[0052] Specifically, in the system model, the size of the scenario is X×Y square kilometers, and the coordinates of the IoT device I k =(X k , Y K , 0), where k∈{1,...,K}. The position where the drone equipped with the intelligent reflecting surface is deployed is represented by the coordinates (X U , Y U , H U ). The position of the base station is represented by the coordinates X B , Y B , H B ). The three-dimensional Euclidean distance, elevation angle, and azimuth angle formed between the IoT device and the intelligent reflecting surface are respectively denoted as d IOT,RIS , θ k→RIS , and ξ k→RIS . Similarly, the three-dimensional Euclidean distance, elevation angle, and azimuth angle formed between the intelligent reflecting surface and the base station are respectively expressed as d RIS,BS , θ RIS→BS , ξ RIS→BS .

[0053] Each IoT device and the base station are equipped with a transmitting antenna and a receiving antenna. The non-payload information exchanged between the IoT device and the base station through the uplink of the drone carrying the intelligent reflecting surface is only the channel state information. The ground control station is a separate node, which interacts with the drone of the intelligent reflecting surface and obtains the spatial position and activation information from the IoT devices. It is responsible for collecting and monitoring the positions of the drone and the positions and activations of the IoT devices.

[0054] Specifically, it further includes the following steps:

[0055] Step S4: Based on the optimal IoT device scheduling strategy, transmission power, and intelligent reflecting surface phase shift, the ground control station arranges for the IoT devices to perform data packet transmission in each time slot, adjusts the phase shift of the intelligent reflecting surface elements, and sets the transmission power of the IoT devices. It can be understood that the adjustment of these parameters is to reduce the overall age of information of the IoT devices and improve their received signal-to-noise ratio.

[0056] In addition, according to different application scenarios, the activation modes of IoT devices are also different. They are generally divided into the activation of IoT devices following the beta distribution and the uniform distribution. The activation model following the beta distribution can be considered an extreme scenario where a large number of devices are activated in a highly synchronized manner, which is common in some applications such as intelligent traffic congestion control. The activation model following the uniform distribution is considered a scenario where devices access the network evenly, that is, in an asynchronous manner, over a period of time. Considering that the IoT devices in the scenario of the present invention are not highly synchronized activated, the activation mode of the IoT devices in the present invention follows the uniform distribution.

[0057] The following is an explanation of the channel model of the system:

[0058] Assume that the base station is equipped with one antenna and receives data from a predetermined IoT device through an intelligent reflecting surface. The intelligent reflecting surface has a total of F reflecting units, and the horizontal and vertical intervals between the reflecting units are d H = λ / 2 and d V = λ / 2, where λ is the wavelength of the carrier.

[0059] According to the communication environment, the IoT devices are located on the ground, while the intelligent reflecting surface is carried on a drone and located in the air. Therefore, it is assumed that the communication link between the IoT device and the intelligent reflecting surface follows the ground-to-air model. According to the free space path loss model, the path loss between the IoT device and the reflecting surface can be described as:

[0060]

[0061] where fc is the carrier frequency (Hz), c is the speed of light (m / s), η LOS and η NLOS respectively represent the average additional path loss in the line-of-sight and non-line-of-sight scenarios. FSPL IOT,RIS represents the free space path loss between the IoT device and the intelligent reflecting surface, and d IOT,RIS represents the three-dimensional Euclidean distance formed between the IoT device and the intelligent reflecting surface. and respectively represent the path loss between the IoT device and the intelligent reflecting surface, which includes the free space path loss and the additional loss caused by non-line-of-sight propagation. This parameter is determined by the specific communication environment.

[0062] Assume h IOT→RIS If [n] is the channel model between the Internet of Things (IoT) device and the intelligent reflecting surface, it can be specifically expressed as:

[0063]

[0064] The small-scale fading of the communication link between the IoT device and the unmanned aerial vehicle (UAV) is modeled according to the Rice distribution, where represents the line-of-sight component of the small-scale fading, represents the non-line-of-sight component of the small-scale fading, then:

[0065]

[0066] obeys a complex Gaussian distribution with a mean of 0 and a variance of 1; where K1 is the Rice factor, which defines the ratio of the line-of-sight component power to the non-line-of-sight component power; is a fixed component vector, and its specific value is determined by the elevation angle θ and azimuth angle ξ between the IoT device and the intelligent reflecting surface, as well as the number and spacing of the reflecting elements in the intelligent reflecting surface.

[0067]

[0068] where U represents the position of the intelligent reflecting surface element, is the phase factor, representing the phase change of the signal during propagation, Ψ is the phase angle vector, and ψ k,F is the phase angle, determined by the elevation angle and azimuth angle between the IoT device and the intelligent reflecting surface.

[0069] U = [u1,..., u F

[0070]

[0071] Considering that in the communication environment, the intelligent reflecting surface is carried on the UAV in the sky, and the base station antenna is located at the top of the base station with a certain height. Therefore, it is assumed that the communication link between the intelligent reflecting surface and the base station follows the air-to-air model. Then the path loss between the intelligent reflecting surface and the base station can be described as:

[0072]

[0073] where β0 is the average channel gain of the path loss at the reference distance d0 = 1m, α B is the path loss exponent, d RIS,BS is the distance between the intelligent reflecting surface and the base station, represents the relationship between the path loss and the distance.

[0074] Assume h​RIS→BS [n] is the channel model between the intelligent reflecting surface and the base station, and can be specifically expressed as:

[0075]

[0076] Assume that the small-scale fading of the communication link between the intelligent reflecting surface and the base station is modeled according to the Rice distribution, where represents the line-of-sight component of the small-scale fading, represents the non-line-of-sight component of the small-scale fading, Δ RIS→BS [n] is the path loss, then:

[0077]

[0078] obeys a complex Gaussian distribution with a mean of 0 and a variance of 1. Among them, K2 is the Rice factor, which defines the ratio of the line-of-sight component power to the non-line-of-sight component power. is a fixed component vector, and its specific value is determined by the elevation angle θ and azimuth angle ξ between the intelligent reflecting surface and the base station antenna, as well as the number and spacing of the reflecting elements in the intelligent reflecting surface.

[0079]

[0080] where U represents the position of the intelligent reflecting surface element, is the phase factor, which represents the phase change of the signal during propagation. ω is the phase angle vector, ω i,F is the phase angle, which is determined by the elevation angle and azimuth angle between the intelligent reflecting surface and the base station.

[0081] According to the channel model parameters calculated above, the signal-to-noise ratio of the base station at time slot n is expressed as:

[0082]

[0083] Since and both follow a complex Gaussian distribution with a mean of 0 and a variance of 1, and their matrix components are random, so the expectation calculation is considered. Among them, Φ is the reflecting element matrix of the intelligent reflecting surface, and is specifically expressed as follows:

[0084]

[0085] where g0 represents the gain of the intelligent reflecting surface element, β F represents the amplitude coefficient, φ FDenote the phase shift. According to the signal-to-noise ratio formula (16), it can be seen that by adjusting the phase shift of the reflection elements of the intelligent reflecting surface, the signal-to-noise ratio can be improved. To obtain the optimal phase shift matrix, we can consider starting from the signal-to-noise ratio formula (16). Since the transmit power P N , the noise are both constants. To maximize the signal-to-noise ratio, that is, to maximize Ε[||h k→RIS [n]Φh RIS→BS [n]|| 2 .

[0086]

[0087] Since and follow a complex Gaussian distribution with a mean of 0 and a variance of 1 and are random variables, the expectation operator E is used; where C is a constant; ψ is the phase angle from the Internet of Things device to the intelligent reflecting surface, and ω is the phase angle between the intelligent reflecting surface and the base station.

[0088] Therefore, the optimal phase shift φ of the intelligent reflecting surface only depends on the elevation angle θ and azimuth angle ξ between the Internet of Things device and the intelligent reflecting surface, between the intelligent reflecting surface and the base station, as well as the number F of reflection elements in the intelligent reflecting surface and the horizontal spacing d between the reflection elements H and the vertical spacing d V , so there is:

[0089] φ = [φ1,.., φ F = Ψ - ω(19)

[0090] The following details the age-of-information model of the system:

[0091] The data sampling model of the Internet of Things device uses a random generation model. Specifically, each Internet of Things device can only generate updated data in the time slot when it is activated, that is, the Internet of Things device only collects data when the drone requests the Internet of Things device to upload data. For the convenience of analysis, the present invention assumes that the wake-up time and information sampling time of the Internet of Things device can be ignored relative to the information transmission time. Assume that the Internet of Things device adopts a single-packet queuing strategy. The data packet will be placed in the cache space of the Internet of Things device without a request from the drone. When the Internet of Things device is activated again, the older state update data packet will be discarded, and the newly arrived data packet will be put into the cache space.

[0092] To achieve successful transmission, the signal-to-noise ratio γ k of the base station in time slot n should be strictly greater than or equal to γ th , where γ this the minimum threshold to ensure reliable decoding. Therefore, the ground control station must appropriately control the scheduling of IoT devices, the transmission power of IoT devices, and the phase shift of the reflecting elements of the intelligent reflecting surface to relay the status update information, while considering the activation status of the IoT devices. Obviously, the age of information depends on the communication scheduling, the phase shift of the intelligent reflecting surface unit, and the activation status of the IoT devices. Therefore, the age of information A k of IoT device k at time slot n evolves in the next time slot as follows:

[0093]

[0094] where α k [n] is a binary variable indicating whether the k-th IoT device is scheduled in time slot n. When α k [n]=1, it means the IoT device is scheduled; when α k [n]=0, it means the IoT device is not scheduled; G k [n] is a binary variable indicating whether the k-th IoT device is active in time slot n. When G k [n]=1, it means the IoT device is active; when G k [n]=0, it means the IoT device is in sleep mode; γ th is the minimum threshold to ensure reliable decoding, and γ k is the signal-to-noise ratio of the base station in time slot n.

[0095] The following details the optimization problem:

[0096] The present invention uses the total age of information of the system as an index to evaluate the freshness of data.

[0097] In step S2, when constructing the age-of-information optimization model, the optimization objective is to minimize the long-term average age of information, specifically:

[0098]

[0099] The constraint conditions are:

[0100] φ F [n]∈[0,2π)(21a)

[0101] α k [n]∈[0,1](21b)

[0102] 0<P k [n]<P max (21c)

[0103] Among them, D(n) represents the transmit power control of Internet of Things (IoT) devices, C(n) represents the information transmission scheduling of IoT devices, Θ(n) represents the reflection phase shift of intelligent reflecting surfaces, K represents the number of IoT devices, N represents the service duration and is evenly divided into N time slots, n ∈ [0, N], A k [n] represents the age of information of IoT device k at time slot n, φ F [n] represents the phase shift of the reflection elements of the intelligent reflecting surface, α k [n] represents the scheduling situation of IoT devices. When α k [n] = 1, it means that IoT device k is scheduled at time slot n, α k [n] = 0 means that IoT device k is not scheduled at time slot n, P k [n] represents the transmit power of IoT devices, P max represents the maximum transmit power of IoT devices.

[0104] Since the activation pattern of IoT devices follows a uniform distribution and has a certain degree of randomness, the optimization problem is a stochastic optimization problem for the service time N. In fact, before dispatching the drone to the target area, it is crucial to obtain the activation pattern of IoT devices. This is because the formulated problem aims to find a control strategy that minimizes the age of information of active IoT devices within the service time N. However, obtaining complete information about the activation pattern requires a large number of measurements, especially in remote areas where it is difficult to obtain. The optimization problem of the present invention is a mixed-integer non-convex optimization problem and is difficult to solve because the optimization problem contains binary variables α i [n] and continuous variables φ F [n] as well as P i [n]. Therefore, the present invention constructs the optimization problem into a Markov decision process and uses a deep reinforcement learning algorithm based on the SD3 algorithm to find an effective control strategy that minimizes the total expected age of information.

[0105] The following details the deep reinforcement learning algorithm based on the SD3 algorithm:

[0106] There is an overestimation problem in the DDPG algorithm of traditional deep reinforcement learning algorithms, that is, the algorithm may overestimate the Q-values of some state-action pairs, thus reducing the performance in complex wireless communication environments.

[0107] One reason for the overestimation problem in DDPG is that the estimation of the Q-value may be overly high, especially when using a single Q-network. Additionally, research has shown that the TD3 algorithm can effectively improve the overestimation problem in the DDPG algorithm. In the TD3 algorithm, by introducing two Q-networks, namely Q1 and Q2, when updating the Q-value, TD3 selects the minimum value of the two Q-networks for update, and at the same time delays the update of the target network. In this way, the TD3 algorithm can effectively reduce the probability of overestimation. However, precisely because the minimum value of the two Q-networks is adopted in the TD3 algorithm, an underestimation bias inevitably occurs, which will significantly reduce its performance.

[0108] The SD3 algorithm improves the TD3 algorithm by introducing the Softmax operator and clipping the action space. The Softmax operator makes the selection of the policy probabilistic, avoiding the agent overly conservatively choosing actions with underestimated Q-values. Even if the Q-values of some actions are relatively low, the agent may still choose these actions, thus preventing the overestimation of Q-values. The clipped action space, on the other hand, restricts the range of actions to ensure that the generated actions are not too extreme, making the Q-value estimation more stable and reducing the underestimation or overestimation bias caused by extreme actions.

[0109] Based on the above considerations, the present invention proposes a reinforcement learning algorithm based on the SD3 algorithm for learning the optimal coordination between the scheduling of Internet of Things devices and the transmission power control of Internet of Things devices.

[0110] Specifically, in step S3, the information age optimization model constructed is solved using the deep reinforcement learning algorithm based on the SD3 algorithm, specifically as follows:

[0111] The optimization problem is formulated as a Markov decision process, which consists of the tuple <s, a, r>. Here, s represents the state, a represents the action, and r represents the reward function. In each training set, the agent observes the current state s(t), and then selects an action a(t) to execute. Once the action is selected, the agent will obtain the corresponding reward r(t), and continue to observe the state s(t + 1) in the next time slot. In this way, the optimal scheduling strategy and device transmission power can be obtained when the training converges, where:

[0112] State space: The state of the system consists of the state of the drone and the state of the Internet of Things devices, denoted as s[n] = (A[n], γ[n]), where A[n] represents the age of information of each Internet of Things device at time slot n, and γ[n] represents the signal-to-noise ratio when the signals sent by the Internet of Things devices at time slot n are coherently combined through the phase of the intelligent reflecting surface elements.

[0113] Action space: The actions of the system include two aspects, namely the scheduling of IoT devices and the transmit power control of IoT devices, denoted as a[n] = (ξ[n], Ρ[n]), where ξ[n] represents the scheduling vector of k IoT devices by the ground control station at time slot n, and Ρ[n] represents the transmit power of the scheduled IoT devices at time slot n.

[0114] Reward function: The reward function is defined as the negative of minimizing the sum of the ages of information of all IoT devices:

[0115]

[0116] where AoI k(n) represents the age of information of the k-th IoT device at time slot n.

[0117] The deep reinforcement learning algorithm adopted is as follows:

[0118]

[0119] The simulation results of the present invention are described below:

[0120] The present invention considers a square area of 0.5 km × 0.5 km with the center of the area at (0, 0, 0), where IoT devices are randomly distributed. A drone equipped with an intelligent reflecting surface is placed at the center of the area to relay the status update information of IoT devices to a base station located at (2000, 500, 25). The experimental time N = 120 and the number of IoT devices K = 10. The maximum transmit power P of the IoT devices max = 20 dBm. The communication parameters are taken as: path loss exponent α B = 2.3, channel gain β0 = -20 dBm, noise power signal-to-noise ratio threshold γ th = 0 dB and K1 = K2 = 8 dB. The number of reflecting elements of the intelligent reflecting surface F = 16. The activation mode of each IoTD follows a uniform distribution, where the activation probability p = 0.6.

[0121] To evaluate the effectiveness of the proposed algorithm, the present invention designs three benchmark strategies for comparison, which are specifically described as follows:

[0122] (1) Random scheduling strategy for IoT devices: This strategy randomly selects an IoT device for relaying the status update. At the same time, the phase setting of the intelligent reflecting surface elements and the transmit power control of the IoT devices are consistent with the algorithm of the present invention.

[0123] (2) Random Phase Shift Strategy of IRS Elements: This strategy randomly adjusts the phase shift of IRS elements to optimize the signal reflection direction. In addition, the scheduling of IoT devices and the control of transmission power are consistent with the algorithm of the present invention.

[0124] (3) Greedy Strategy: This strategy selects IoT devices for status update relaying based on the maximum current age of information value. The phase shift control of the intelligent reflecting surface is consistent with the algorithm of the present invention, and the transmission power is a fixed value.

[0125] Figure 3 It shows that the SD3 algorithm proposed by the present invention can effectively converge in dealing with the optimization problem. From Figure 3 it can be seen that the algorithm basically converges to a stable point after about 2000 rounds of training.

[0126] Figure 4 It shows the comparison of the long-term average age of information of different algorithms under different numbers of IoT devices. When the number of IoT devices increases, the long-term average age of information will increase to varying degrees. This is because when the number of IoT devices increases, the scheduling ability of the ground control station is limited and it cannot schedule more IoT devices. At the same time, from Figure 4 it can be observed that the SD3 algorithm proposed by the present invention is superior to the benchmark algorithm. This is because the SD3 algorithm proposed by the present invention obtains better phase shifts, device scheduling, and appropriate transmission power through exploration and training, thereby improving the received signal-to-noise ratio of the base station and ensuring that the received signal-to-noise ratio of the device to the base station is greater than the threshold when scheduling the device as much as possible, thus ensuring the reliability of transmission. Therefore, the SD3 algorithm proposed by the present invention has better results compared to the random phase shift algorithm, the random scheduling algorithm, and the greedy algorithm.

[0127] Figure 5 It shows the comparison of the long-term average age of information of different algorithms under different sizes of the distribution area of IoT devices. The distribution of IoT devices follows a uniform distribution. Therefore, when the distribution area of IoT devices is larger, the average distance between IoT devices and the intelligent reflecting surface is larger, resulting in an increase in the signal-to-noise ratio. The increase in the total age of information of the method proposed by the present invention is relatively small, indicating the importance of optimizing the phase shift and adjusting the transmission power of IoT devices.

[0128] Figure 6 It is the comparison of the long-term average age of information under different numbers of intelligent reflecting surface elements. Since the intelligent reflecting surface can optimize the communication link between IoT devices and the base station, increasing the number of reflecting elements of the intelligent reflecting surface can effectively improve the channel gain between IoT devices and the base station, thereby improving the signal-to-noise ratio and effectively reducing the age of information.

[0129] Figure 7Comparison of the long - term average age of information for different activation probabilities of IoT devices. Obviously, when the activation probability of IoT devices increases, the total age of information will decrease. This is because the higher the activation probability of IoT devices, the more activation times within a period of time. If the received signal - to - noise ratio is greater than the threshold during its activation period, it is more likely to reduce information transmission, thus reducing the age of information.

[0130] The present invention optimizes the long - term average age of information of a communication system by scheduling IoT devices, jointly adjusting the phase shift of intelligent reflecting surface elements and the transmission power control of IoT devices, and proposes a deep reinforcement learning method based on the SD3 algorithm to solve the optimization problem. The simulation experiment results show that, compared with benchmark algorithms such as the random algorithm and the greedy algorithm, the deep reinforcement learning method based on the SD3 algorithm proposed by the present invention can significantly reduce the age of information and ensure the freshness of information.

[0131] Generally speaking, the present invention introduces a drone equipped with an intelligent reflecting surface as a relay node, combines with a deep reinforcement learning algorithm, jointly optimizes the scheduling, transmission power of IoT devices and the phase shift of the intelligent reflecting surface, effectively reduces the age of information in the IoT communication system, improves the freshness and availability of information, and enhances the overall performance and communication efficiency of the system. At the same time, the deep reinforcement learning method based on the SD3 algorithm can adapt to the randomness of the activation mode of IoT devices and the complexity of the optimization problem, and find a near - optimal control strategy.

[0132] In the second aspect of the embodiments of the present invention, an information age optimization system based on an intelligent reflecting surface is proposed, corresponding to the information age optimization method based on an intelligent reflecting surface provided in the above - mentioned embodiments of the present invention. Since the information age optimization system based on an intelligent reflecting surface provided in the embodiments of the present invention corresponds to the information age optimization method based on an intelligent reflecting surface provided in the above - mentioned embodiments of the present invention, the implementation manners of the foregoing information age optimization method based on an intelligent reflecting surface are also applicable to the information age optimization system provided in this embodiment.

[0133] Specifically, the system includes the following modules:

[0134] A system model construction module, used to construct a system model, including a drone equipped with an intelligent reflecting surface, a base station, K IoT devices, and a ground control station. The intelligent reflecting surface has a total of F reflecting elements. Among them, a virtual link is established between the IoT devices and the base station through the intelligent reflecting surface carried by the drone. The ground control station is used to schedule the IoT devices to transmit information and control the phase shift of the intelligent reflecting surface and the transmission power of the IoT devices.

[0135] An optimization model construction module, used to construct an information age optimization model, and the optimization goal is to minimize the long - term average age of information.

[0136] A solution module, which is used to solve the constructed information age optimization model by using a deep reinforcement learning algorithm based on the SD3 algorithm, so as to obtain the optimal Internet of Things device scheduling strategy, transmission power, and intelligent reflecting surface phase shift.

[0137] See Figure 8 , and the embodiment of the present invention also correspondingly provides an electronic device and a computer-readable storage medium.

[0138] As Figure 8 shown is a schematic diagram of an electronic device provided by an embodiment of the present invention. The electronic device of this embodiment includes: a processor 11, a memory 12, and a computer program stored in the memory and executable on the processor 11. When the processor 11 executes the computer program, the steps in the above-mentioned embodiment of the information age optimization method based on an intelligent reflecting surface are implemented. Alternatively, when the processor 11 executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0139] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor 11 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0140] The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the electronic device, and does not constitute a limitation on the electronic device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, buses, etc.

[0141] The so-called processor 11 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects all parts of the entire electronic device through various interfaces and circuits.

[0142] The memory 12 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by invoking the data stored in the memory, the processor realizes various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system 121, application programs 122 required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0143] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0144] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative effort.

[0145] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

Claims

1. A method for optimizing information age based on intelligent reflective surface, characterized in that: The following steps are involved: Construct a system model, including a UAV equipped with a smart reflective surface, a base station, K IoT devices, and a ground control station. The smart reflective surface has a total of F reflective elements. A virtual link is established between the IoT device and the base station through the smart reflective surface carried by the UAV. The ground control station is used to schedule the IoT device to transmit information and control the phase shift of the smart reflective surface and the transmission power of the IoT device. Construct an information age optimization model, the optimization goal is to minimize the long-term average information age; The constructed information age optimization model is solved using a deep reinforcement learning algorithm based on the SD3 algorithm to obtain the optimal IoT device scheduling strategy, transmission power, and smart reflector phase shift.

2. The information age optimization method based on intelligent reflective surface according to claim 1 is characterized in that: The information age optimization model is constructed, and the optimization goal is to minimize the long-term average information age, specifically: The constraints are: f F [n]∈[0,2π) α k [n]∈[0,1] 0<P k [n]<P max Where D(n) represents the transmission power control of IoT devices, C(n) represents the information transmission scheduling of IoT devices, Θ(n) represents the reflection phase shift of the smart reflector, K represents the number of IoT devices, N represents the service duration and is evenly divided into N time slots, n∈[0,N], A k [n] represents the information age of IoT device k in time slot n, φ F [n] represents the phase shift of the reflective element of the smart reflective surface, α k [n] represents the scheduling of IoT devices. k When [n] = 1, it means that IoT device k is scheduled in time slot n, α k When [n] = 0, it means that IoT device k is not scheduled in time slot n. k [n] represents the transmission power of IoT devices, P max Indicates the maximum transmit power of the IoT device.

3. The information age optimization method based on intelligent reflective surface according to claim 2 is characterized in that: The activation pattern of IoT devices follows a uniform distribution.

4. The information age optimization method based on intelligent reflective surface according to claim 2 is characterized in that: The deep reinforcement learning algorithm based on the SD3 algorithm is used to solve the constructed information age optimization model, specifically: The optimization problem is formulated as a Markov decision process, which consists of<s,a,r> This tuple consists of s representing the state, a representing the action, and r representing the reward function. In each training set, the agent observes the current state s(t) and then selects an action a(t) to perform. Once the action is selected, the agent will receive the corresponding reward r(t) and continue to observe the state s(t+1) in the next time slot, where: State space: The state of the system consists of the state of the drone and the state of the IoT device, expressed as s[n] = (A[n], Υ[n]), where A[n] represents the information age of each IoT device at time slot n, and Υ[n] represents the signal-to-noise ratio when the signal sent by the IoT device at time slot n is coherently combined through the phase of the smart reflective surface element; Action space: The action of the system includes two aspects, namely, the scheduling of IoT devices and the control of the transmission power of IoT devices, which is expressed as a[n] = (ξ[n], P[n]), where ξ[n] represents the scheduling vector of k IoT devices by the ground control station at time slot n, and P[n] represents the transmission power of the scheduled IoT device at time slot n; Reward function: The reward function is defined to minimize the negative sum of the information age of all IoT devices: AoI k(n) represents the information age of the kth IoT device at time slot n.

5. The information age optimization method based on intelligent reflective surface according to claim 2 is characterized in that: The information age A of IoT device k in time slot n k [n] evolves in the next time slot to: Among them, G k [n] is a binary variable indicating whether the kth IoT device is activated in time slot n. k [n] = 1 indicates that the IoT device is activated. k [n] = 0 means the IoT device is in sleep mode; th is the minimum threshold to ensure reliable decoding, k is the signal-to-noise ratio of the base station in time slot n.

6. The information age optimization method based on intelligent reflective surface according to claim 1 is characterized in that: Each IoT device and base station is equipped with a transmitting antenna and a receiving antenna. The only non-payload information exchanged between the IoT device and the base station via the uplink of the drone equipped with a smart reflective surface is the channel state information.

7. The information age optimization method based on intelligent reflective surface according to claim 1 is characterized in that: The following steps are also included: According to the optimal IoT device scheduling strategy, transmission power and smart reflector phase shift, the ground control station schedules IoT devices for data packet transmission in each time slot, adjusts the phase shift of smart reflector elements and sets the transmission power of IoT devices.

8. An information age optimization system based on intelligent reflective surface, characterized in that: Includes the following modules: The system model building module is used to build a system model, which includes a UAV equipped with an intelligent reflective surface, a base station, K IoT devices, and a ground control station. The intelligent reflective surface has a total of F reflective elements. A virtual link is established between the IoT device and the base station through the intelligent reflective surface carried by the UAV. The ground control station is used to schedule the IoT device to transmit information and control the phase shift of the intelligent reflective surface and the transmission power of the IoT device. The optimization model building module is used to build an information age optimization model, and the optimization goal is to minimize the long-term average information age; The solution module is used to solve the constructed information age optimization model using a deep reinforcement learning algorithm based on the SD3 algorithm to obtain the optimal IoT device scheduling strategy, transmission power, and smart reflector phase shift.

9. An electronic device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements an information age optimization method based on a smart reflective surface as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the information age optimization method based on the smart reflective surface as described in any one of claims 1 to 7.