A post-disaster emergency communication system
By combining reconfigurable smart surfaces and drones, utilizing signal modulation and energy transfer modules, and incorporating MDP and DRL optimization strategies, the problems of obstruction and dynamic changes in post-disaster communication were solved, achieving efficient communication restoration and energy supply in disaster areas, and improving emergency response efficiency and system robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-05
AI Technical Summary
Existing UAV-assisted mobile edge computing emergency communication architectures do not fully consider complex wireless propagation characteristics such as obstruction, multipath fading, and link interruption in post-disaster scenarios. They are difficult to adapt to highly dynamic changes, lack energy self-sustaining capabilities and cross-layer joint optimization, resulting in insufficient robustness and adaptability. They fail to simultaneously meet the comprehensive needs of mission latency, UAV energy consumption, and stable operation of end users.
By combining a reconfigurable smart surface (STAR-RIS) with a drone, the system modulates the diffraction of obstacles through the transmission and reflection of signals in the air-to-ground link. Combined with a wireless information and energy transmission module, a control module is introduced for adaptive optimization. Taking into account the drone's flight trajectory, resource allocation, and communication computing decisions, Markov Decision Process (MDP) and Reinforcement Learning (DRL) optimization strategies are used to achieve real-time adjustment of the post-disaster environment.
It significantly improved the communication recovery capability and emergency response efficiency in disaster areas, achieved near-full space coverage, enhanced link reliability and controllability, ensured data communication needs and provided continuous energy replenishment, extended the flight time of UAVs, and improved the system's sustainable operation capability under energy-constrained conditions.
Smart Images

Figure CN121585983B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emergency communications, and more specifically to a post-disaster emergency communications system. Background Technology
[0002] Given the urgency of post-disaster missions and the complexity of the environment, emergency communication networks need to be rapidly deployable, flexible, and highly reliable. Drones, due to their advantages such as maneuverability, rapid deployment, low cost, and strong line-of-sight links, are considered a key solution for post-disaster communication recovery. Meanwhile, mobile edge computing can enhance computing and communication capabilities at the network edge. Combining mobile edge computing technology with drone communication technology to construct a drone-assisted mobile edge computing (U-MEC) emergency architecture can rapidly restore communication and provide computing support in disaster scenarios where infrastructure has been severely damaged.
[0003] Existing technology 1 (Shah Z, Javed U, Naeem M, et al. Mobile edge computing (MEC)-enabled UAV placement and computation efficiency maximization in disaster scenario[J]. IEEE Transactions on Vehicular Technology, 2023, 72(10): 13406-13416) adopts an emergency architecture for UAV-assisted mobile edge computing, aiming to maximize computational efficiency. However, it only considers the direct link between the user and the UAV, and does not adequately consider the obstacles, complex electromagnetic environment, and channel degradation that are common in post-disaster environments. Its optimization method relies on solution frameworks such as k-means clustering and phased static interior point method, which are difficult to adapt to the rapid time-varying characteristics of user distribution, channel conditions, and energy status in post-disaster scenarios. It lacks real-time performance and adaptive capabilities, and mainly focuses on the joint allocation of communication resources and computing resources, without further exploring the mechanism to improve the sustainable operation capability of the system under energy shortage conditions.
[0004] Existing technology two (Luan Q, Cui H, Zhang L, et al. A hierarchical hybrid subtask scheduling algorithm in UAV-assisted MEC emergency network[J]. IEEE Internet of Things Journal, 2021, 9(14): 12737-12753) mainly focuses on topology reconstruction and subtask scheduling to reduce the average completion latency of computing tasks. However, its system model is still based on the traditional UAV-assisted mobile edge computing architecture and does not consider the wireless environment enhancement mechanism under the conditions of severe obstruction, link interruption and energy limitation that are common in post-disaster communication. At the same time, the H-HSS algorithm proposed by this technology depends on the given channel conditions and the preset network topology, and lacks the ability to perform end-to-end wireless propagation enhancement, energy self-sustaining and cross-layer intelligent optimization in highly dynamic post-disaster environments.
[0005] Existing technology three (Khalid R, Shah Z, Naeem M, et al. Computational efficiencymaximization for UAV-assisted MEC networks with energy harvesting in disasterscenarios[J]. IEEE Internet of Things Journal, 2023, 11(5): 9004-9018) constructs an emergency architecture for UAV-assisted mobile edge computing with radio frequency energy harvesting capabilities, aiming to maximize computational efficiency. It adopts a three-stage static solution framework of "k-means clustering + phased linearization + interior point method". However, this work is still based on direct links with probabilistic Loss of Spectrum (LoS), and does not fully characterize the complex factors in post-disaster scenarios such as severe occlusion, multipath fading, link interruption, and dynamic user arrival. Its offline phased optimization is difficult to achieve real-time adaptive adjustment of UAV position, energy status, and offloading strategy, and has limited support for system robustness and multi-dimensional resource joint intelligent optimization.
[0006] Existing technology four ([Li J, He Q, Wang X, et al. UAV-assisted Microservice Mobile Edge Computing Architecture: Addressing Post-Disaster Emergency Medical Rescue[J]. IEEE Transactions on Computers, 2025) optimizes the UAV-assisted mobile edge computing architecture through microservices and Transformer resource management to improve task scheduling and energy efficiency in post-disaster medical rescue. However, its overall design still relies on traditional air-to-ground links and the coverage of the UAV itself, and lacks in-depth consideration of the uncertainties in wireless propagation such as multipath fading and link instability that are common in post-disaster environments. In addition, the optimization framework in this paper mainly revolves around the resource allocation of Lyapunov + Transformer, focusing more on computing and energy consumption, while lacking sufficient exploration of air-to-ground link quality improvement, energy acquisition mechanisms, and intelligent environmental perception-decision coupling.
[0007] Existing technology five (Sun L, Liu Z, Ning Z, et al. Multi-Agent Q-Net Enhanced Coevolutionary Algorithm for Resource Allocation in Emergency Human-MachineFusion UAV-MEC System[J]. IEEE Transactions on Automation Science and Engineering, 2024, 22: 4473-4489) proposes an optimization method combining multi-agent Q-networks and co-evolutionary algorithms. This method utilizes Q-networks for local agent policy learning and leverages the co-evolutionary mechanism to jointly optimize global resource allocation strategies, thereby improving the task processing efficiency and resource utilization of UAV-assisted mobile edge computing emergency systems. However, this optimization method still relies on a pre-defined air-space-ground link model and static resource parameters, and primarily focuses on policy evolution at the task scheduling and resource allocation levels. It does not consider factors such as channel randomness, energy constraints, and environmental reconfigurability in complex post-disaster environments, making its adaptability and robustness insufficient in highly dynamic emergency scenarios.
[0008] Existing technology 6 (Wang B, Sun Y, Jung H, et al. Digital twin-enabled computation offloading in UAV-assisted MEC emergency networks[J]. IEEE Wireless Communications Letters, 2023, 12(9): 1588-1592) constructs a digital twin-driven UAV-assisted mobile edge computing emergency communication architecture. In disaster scenarios, it utilizes the DT layer of the front-end command center to map the UAV and user status in real time, and organizes UAVs to provide edge computing and offloading services for front-line rescue and survivors. At the same time, it models the offloading problem in uncertain environments as online matching on an uncertain bipartite graph, and designs a SUOM online stable matching algorithm based on MAB–UCB to adaptively select the matching window to balance latency and stability. However, this method still relies on idealized Loss channels and static deviation interval assumptions, and only optimizes at the task offloading and matching level. It does not fully consider factors such as severe obstruction, multipath fading, link interruption and energy constraints after disasters, and its robustness and multi-dimensional resource joint optimization capabilities in complex and highly dynamic emergency scenarios are still insufficient.
[0009] Based on the above analysis, the existing UAV-assisted mobile edge computing emergency communication architecture and deployment optimization algorithms have the following shortcomings:
[0010] a. It still relies on idealized air-to-ground direct links and does not adequately consider the complex wireless propagation characteristics common in post-disaster environments, such as obstruction, multipath fading, and link interruption.
[0011] b. Most optimization strategies use static or semi-static solution frameworks, which are difficult to adapt to the highly dynamic changes in user distribution, channel conditions and energy status in post-disaster scenarios.
[0012] c. Insufficient consideration of network sustainability under energy-constrained conditions, reconfigurable wireless environment mechanisms, and cross-layer joint intelligent optimization of communication-computing-energy leads to significant limitations in robustness, adaptability, and resource coordination efficiency in real post-disaster emergency scenarios.
[0013] d. It has not yet fully taken into account the comprehensive needs of emergency communication scenarios, such as mission latency, drone energy consumption, and the continuous and stable operation of end users. Summary of the Invention
[0014] In view of the above-mentioned shortcomings in the prior art, the present invention provides a post-disaster emergency communication system that solves the problem that the prior art is unable to improve the communication recovery capability and emergency response efficiency of the disaster area while meeting multiple needs such as post-disaster communication, computing and energy supply.
[0015] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0016] A post-disaster emergency communication system is provided, comprising a reconfigurable smart surface, a drone, and a control module; wherein:
[0017] The drone is equipped with an edge server for adjusting the spatial location of the edge server; the edge server is used to communicate with the reconfigurable smart surface and / or the user.
[0018] Reconfigurable smart surfaces can be deployed on the ground and respond to the communication needs of users and / or drones. By controlling the transmission and reflection of incident signals in the air-to-ground link, they can diffract or avoid obstacles in post-disaster scenarios and improve signal coverage.
[0019] The control module is used to acquire and generate control actions for one or more devices based on the operating status of each device in the system, with the goal of minimizing the maximum processing latency of all user tasks and the energy consumption of the drone, so as to achieve the ability to adapt to changes in the post-disaster environment; the energy consumption of the drone includes the energy consumption of the edge server.
[0020] The beneficial effects of this invention are as follows:
[0021] 1. This system comprehensively considers multiple decision variables such as UAV flight trajectory planning, STAR-RIS beamforming, user task offloading ratio, and communication and computing resource allocation. It proposes a joint optimization strategy to minimize the maximum processing latency of all user tasks and UAV energy consumption, thereby extending the UAV's endurance while meeting the timeliness requirements of emergency services. This system significantly improves communication recovery capabilities and emergency response efficiency in disaster areas while meeting multiple needs such as post-disaster communication, computing, and energy supply. It also possesses advantages such as rapid deployment, flexibility, and high reliability.
[0022] 2. This system achieves near-full 360° coverage by finely controlling the transmission and reflection of incident signals in the air-to-ground link, effectively diffracting or avoiding obstacles and obstructions in post-disaster scenarios, and significantly improving the reliability and controllability of the link in complex environments.
[0023] 3. This system introduces a wireless information and energy simultaneous transmission module, which unifies the data transmission and energy harvesting in radio frequency signals into the same physical process. This not only ensures the data communication needs of emergency services, but also provides continuous energy supply for terminal equipment and some network nodes, thereby enhancing the system's sustainable operation capability under energy-constrained conditions. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the system architecture;
[0025] Figure 2 This is a logical diagram of a recurrent neural network (GRU).
[0026] Figure 3 This is a structural diagram of the decision-making model;
[0027] Figure 4 The logic diagram of the power allocation mechanism used in the wireless information and energy simultaneous transmission module. Detailed Implementation
[0028] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0029] like Figure 1 As shown, the post-disaster emergency communication system includes a reconfigurable smart surface, a drone, and a control module; wherein:
[0030] The drone is equipped with an edge server for adjusting the spatial location of the edge server; the edge server is used to communicate with the reconfigurable smart surface and / or the user.
[0031] Reconfigurable smart surfaces can be deployed on the ground and respond to the communication needs of users and / or drones. By controlling the transmission and reflection of incident signals in the air-to-ground link, they can diffract or avoid obstacles in post-disaster scenarios and improve signal coverage.
[0032] The control module is used to acquire and generate control actions for one or more devices based on the operating status of each device in the system, with the goal of minimizing the maximum processing latency of all user tasks and the energy consumption of the drone, so as to achieve the ability to adapt to changes in the post-disaster environment; the energy consumption of the drone includes the energy consumption of the edge server.
[0033] The wireless information and energy co-transmission module is used to unify the data transmission and energy harvesting in radio frequency signals into the same physical process, providing wireless energy replenishment for one or more devices in the system.
[0034] exist Figure 1In this scenario, one user (UE) is randomly distributed across the disaster area. Since ground base stations (BS) are almost entirely disabled, this embodiment deploys a drone (UAV) equipped with an edge server. By combining Wireless Information and Power Transfer (SWIPT) technology, the UAV provides communication restoration and power supply while assisting ground users in completing their tasks. To address insufficient communication coverage and potential signal obstruction, a reconfigurable smart surface (STAR-RIS) capable of simultaneous transmission and reflection is also installed on the ground. Therefore, data transmission relies on the signal superposition of the direct link (UE-UAV) and the STAR-RIS auxiliary link (UE-RIS-UAV). The auxiliary link can be further divided into reflection and transmission links. When the UAV and UE are on the same side of the STAR-RIS, data is uploaded via the reflection link. Otherwise, the signal is uploaded to the UAV via the transmission link.
[0035] Assuming the UE travels at a speed within a specified area Randomly move, its position is UAVs at high altitudes Above speed Flying, its position is The location of STAR-RIS is The entire communication period T is divided into K time slots. Within each time slot, the UAV hovers at its current location and establishes communication with the user. The UE can choose to compute some tasks locally and offload others to the UAV for processing. The UAV receives the tasks, processes them, and returns the results to the UE. These signals simultaneously transmit information and energy to the UE. The UE then separates the signals to perform information decoding (ID) and energy harvesting (EH) simultaneously.
[0036] Assume STAR-RIS consists of a uniform linear array (ULA) with M cells. The reflection and transmission characteristics of each cell are determined by a complex phase factor. and Decision. Among them... and These represent the phase shifts of reflection and transmission, respectively, with values ranging from... The matrices of the reflected and transmitted waves of STAR-RIS can be represented as angle matrices, respectively. and .
[0037] In this embodiment, the simulation environment is composed of multiple mathematical models to approximate real post-disaster communication scenarios, including a communication channel model, a nonlinear energy harvesting model, a task latency model, and a system energy consumption model. The key parameters of these models are calibrated based on real test data to ensure the simulation environment's real-world transferability.
[0038] For the communication channel model, this embodiment uses the Rician fading model to describe the channel characteristics in the wireless communication system. This model combines large-scale fading and small-scale fading, and can effectively simulate the attenuation and fluctuation of signal propagation in a post-disaster emergency environment.
[0039] In the STAR-RIS auxiliary link, the channel is divided into two segments: UE to STAR-RIS and STAR-RIS to UAV. The channel vector for the former is represented as:
[0040]
[0041] in Represents the large-scale fading coefficient. This represents the corresponding Rician factor. This is the channel gain at a reference distance of 1m. It is the loss factor between UE and STAR-RIS. The line-of-sight (LoS) component is given by the following formula:
[0042]
[0043] in For carrier wavelength, The azimuth angle from the i-th UE to STAR-RIS is expressed as: . This refers to the cell spacing of STAR-RIS. The corresponding non-line-of-sight channel components still follow... .
[0044] The channel vector from STAR-RIS to UAV can be represented as:
[0045]
[0046] in This represents the corresponding Rician factor. It is the loss factor between STAR-RIS and UAV. The line-of-sight (LoS) component is given by the following formula:
[0047]
[0048] in This represents the angle between STAR-RIS and the UAV. The corresponding non-line-of-sight channel component still follows... .
[0049] The channel vector for the direct link (UE-UAV) is:
[0050]
[0051] Where the distance , It is the loss factor between UE and UAV. This represents the Rician factor. The line-of-sight (LoS) component is given by the following formula:
[0052] .
[0053] The corresponding non-line-of-sight channel components still follow .
[0054] Based on the positional relationship between the UE and UAV relative to STAR-RIS, the UE is determined to belong to either the reflecting or refraction user group. Reflecting users use a reflection coefficient matrix to optimize the signal, while refraction users use a transmission coefficient matrix. Finally, the total channel vector is obtained by superimposing all signals, expressed as:
[0055]
[0056]
[0057] in , These represent the superimposed channel vectors of the reflected and transmitted signals, respectively.
[0058] For the nonlinear energy harvesting model, this embodiment employs a nonlinear energy harvesting (EH) model based on logic functions to accurately calculate the energy harvesting rate of the UE. Compared to the traditional linear model, this model can more effectively describe the energy transfer process in the system and truly reflect the characteristics of the actual circuit. Its mathematical expression is as follows:
[0059]
[0060] in This represents the energy that the circuit can collect when the input power P is [value missing]. This is a constant, representing the maximum collected power that the UE can obtain when the EH circuit is saturated. Parameter and These are used to describe the specific characteristics of a circuit, including resistance, capacitance, and circuit sensitivity.
[0061] In the SWIPT communication framework, we employ power distribution (PS) technology. This enables the UE to perform ID and EH simultaneously on the same received signal. The UE splits the received RF signal according to a power factor, where... Partially used for EH, Used for ID. Therefore, the EH rate can be expressed as:
[0062]
[0063] Where Pdown represents the downlink transmission power of the UAV. According to Shannon's theorem, the downlink transmission rate can be expressed as:
[0064]
[0065] Where B is the system bandwidth; σ² represents the noise power. Similarly, the uplink transmission rate sent by the i-th UE to the UAV is:
[0066]
[0067] in This indicates the UE's transmit power in the uplink.
[0068] For example, such as Figure 4 As shown, in the SWIPT communication framework, the receiver uses power distribution technology to process the received radio frequency signal. That is, the received signal is divided into two paths by a power divider according to a preset power distribution coefficient. One path is sent to the energy harvester to realize the conversion of radio frequency energy to DC energy, and the other path is sent to the information decoder to complete the demodulation and decoding of the baseband signal. Thus, energy harvesting and information transmission can be carried out in parallel under the same receiving architecture.
[0069] For the task latency model, the latency in this embodiment mainly includes four parts: task offloading latency, edge computing latency, local processing latency, and result feedback latency. For each UE's task, we adopted a partial offloading strategy. Assume the proportion of tasks offloaded to the UAV is... Then the task offloading delay of the i-th UE in time slot k can be expressed as:
[0070]
[0071] in This represents the total workload of the i-th UE. This represents the portion of the task that needs to be processed locally. Therefore, the local computation latency can be expressed as:
[0072]
[0073] in This represents the CPU processing speed of the UE, where C is the number of CPU cycles required per bit of task. Meanwhile, the computation latency of the edge server can be expressed as:
[0074]
[0075] Where f UAVThis represents the computation frequency of the edge server. Furthermore, the result feedback latency can be expressed as:
[0076]
[0077] Therefore, the task processing delay in time slot k can be expressed as: .
[0078] For the system energy consumption model, the total energy consumption of the UAV in this embodiment mainly includes flight energy consumption, transmission energy consumption, and computing energy consumption. During flight, the UAV's position information update is represented by the following formula:
[0079]
[0080] in This refers to the flight speed of the UAV. For flight angle, the range is ;and Let be the flight time. During this process, the UAV's flight energy consumption can be expressed as:
[0081]
[0082] in For UAV quality. UAV transmission power consumption. Transmission power P down The time consumption is determined by the product of the downlink transmission time and the time spent in the transmission. Since ID and EH are performed simultaneously during transmission, these parallel processes affect the transmission time of each time slot. Furthermore, when computation is performed on the MEC server, its energy consumption can be expressed as:
[0083]
[0084] in This represents the chip's energy consumption coefficient. Therefore, the total energy consumption in time slot k is... .
[0085] The UE's total energy consumption is divided into transmission energy consumption and computing energy consumption. When some tasks are offloaded, the UE will utilize local energy for data transmission. The offload time is... Therefore, the UE's transmission power consumption This can be expressed as UE transmit power P i The product of the unloading time and the UE's computational energy consumption is calculated based on the task's computational requirements and processing capabilities, as shown below:
[0086] .
[0087] Based on the energy consumption generated by task unloading and local calculations, as well as the charging amount obtained through SWIPT, the UE's energy level at the next moment is represented as follows:
[0088]
[0089] in The amount of electricity consumed, The amount of electricity obtained. These are the minimum and maximum battery limits for the UE, ensuring that the battery level does not become too low or too high. (Function) This is used to ensure that the battery level is always within a reasonable range.
[0090] In post-disaster communication scenarios, emergency tasks often require rapid processing, while UAV energy resources are limited. Therefore, our system addresses two objectives: minimizing user task processing latency and minimizing the total energy consumption of the UAV. Simultaneously, various practical constraints must be considered. For example, UE battery energy is limited, and the system must ensure that the energy requirements of each UE are met. By jointly optimizing the user UAV flight trajectory, task offloading ratio, STAR-RIS phase shift matrix, and resource allocation strategy, the following optimization objectives are simultaneously addressed:
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101] in The proportion of tasks offloaded to edge servers during time slot k; This represents the location of the drone in time slot k, which is also the location of the edge server. ; Let k be the flight speed of the drone, which is also the flight speed of the edge server. Let K be the flight angle of the UAV in time slot k, which is also the flight angle of the edge server. and These are the reflection phase matrix and transmission phase matrix of the reconfigurable smart surface, respectively, used to characterize the phase modulation applied to the reflected and transmitted signals by each reconfigurable smart surface; K is the total number of time slots in the entire communication period T, and k is the time slot index; I is the total number of users. Index for users; For time slot k, the first Maximum processing latency for a single user task; Processing the kth time slot Drone energy consumption per user task; Indicates constraints; Indicates the initial position of the drone; For flight time; This refers to the maximum flight speed of the drone; and respectively drones in shaft and The farthest flight distance of the axis; For time slot k, the first The location of each user ; and The first Individual users shaft and The furthest distance the axis can travel; Let k be the flight energy consumption of the UAV in time slot k. Let k be the data transmission energy consumption of the UAV in time slot k. Let k be the computing power consumption of the edge server in time slot k. To maximize the battery power of the drone; and These are the user's minimum and maximum battery limits, respectively. For time slot k, the first Battery level of each user; This indicates the reflection phase shift at time slot k; This indicates the transmission phase shift at time slot k; Indicates the first The total number of tasks for each user; D represents all tasks.
[0102] Of the constraints mentioned above, the first specifies the initial position of the UAV. The second constraint represents the maximum flight speed of the UAV. The third and fourth constraints state that the UAV and the user can only move within the designated area. The fifth constraint ensures that the UAV's energy consumption does not exceed its battery's maximum capacity. The sixth constraint ensures that the UE's battery power is within a reasonable range. The seventh constraint represents the range of task offload ratio values. The eighth constraint represents reflection and transmission phase shifts. The ninth constraint ensures that all tasks are completed on time.
[0103] The above problem has three key characteristics: 1) The optimization objectives are conflicting. On the one hand, to reduce UAV energy consumption, UAV communication and movement need to be minimized, but this leads to increased task latency and reduced charging efficiency; on the other hand, to reduce task processing latency, UAV communication and flight intensity need to be increased, which will inevitably increase energy consumption. 2) The scenario involved involves mixed actions, encompassing discrete and continuous variables, making the problem a non-convex optimization problem. 3) The problem is strongly coupled and long-term. The tight coupling between optimization variables makes simply decomposing the problem into multiple sub-problems and solving them one by one extremely complex. Furthermore, current decisions not only affect the current outcome but also influence future states, resulting in dynamic changes and temporal dependencies.
[0104] Therefore, this problem can be categorized as a mixed-integer nonconvex programming problem (MINLP). To solve this multi-objective, complex, and dynamic optimization problem, it is necessary to obtain feedback information from the environment in real time and adjust the decision-making strategy based on this information. The DRL method can continuously adjust the agent's action strategy through interaction with the environment, thereby achieving a global long-term optimization goal.
[0105] To solve the above problem, we model it as a Markov Decision Process (MDP). An MDP can be represented as a tuple. .in , , , and These are represented as state space, action space, reward function, state transition probability, and discount factor, respectively. In this embodiment, these key elements are defined as follows in the constructed system:
[0106] state space Specifically, this includes the location of the UE. Corresponding workload and the current location of the UAV. and battery level Total number of remaining tasks .
[0107] Action space Our MDP action space directly corresponds to the decision variables described in Section 2.6. This mainly includes the UAV direction. and speed Data offloading ratio per UE The scheduling strategy for each UE, and the beamforming matrix. , And the phase change of each unit.
[0108] reward function Our goal is to minimize the processing latency of all tasks while reducing the energy consumption of the UAV and ensuring the power requirements of the UE. Since these two optimization objectives have different physical dimensions, we first unify their orders of magnitude and then transform the multi-objective problem into a single scalar reward through weighted summation. Accordingly, the reward function is defined as the weighted negative of the computational latency generated by all task executions and the UAV's energy consumption. ,in and This represents the weighting coefficient. The penalty term introduced for UE power constraints is expressed as follows: .in The penalty coefficient is... This is a battery power threshold. The lower the battery level of a UE, the greater the number of UEs below the threshold. The larger the value, the lower the total reward the system receives.
[0109] State transition probability The system's state transitions depend on the current state. and the actions performed This process is influenced not only by the UAV's position and flight trajectory, but also by the combined effects of multiple factors such as the UE's position, energy consumption, and mission execution progress. These factors collectively determine the system's state at the next moment.
[0110] Discount factor : A parameter used to measure the importance of future rewards relative to current rewards; its value ranges from 100 to 100. .
[0111] To solve the above problem, such as Figure 3 As shown, the specific method (TD3-SPGK) for generating control actions for one or more devices in this embodiment, with the goal of minimizing the maximum processing latency of all user tasks and the power consumption of the drone, includes the following steps:
[0112] S1. Construct a decision model based on Actor networks and dual-Q networks; where the dual-Q network includes two sets of Critic networks;
[0113] S2. Obtain physical constraints and set decision model hyperparameters; where physical constraints include the spatial locations of all users, the location of the reconfigurable smart surface, the initial location of the drone and the total time duration, the maximum available energy of the drone, and the user's battery threshold; decision model hyperparameters include the reward function, discount factor, learning rate of the Actor network and two Critic networks, soft update coefficient of the target network, noise level, pruning boundary, and policy delay steps; the target network includes the target Actor network and the target Critic network;
[0114] S3. Initialize the reinforcement learning subject: Construct an experience pool to store the samples generated by the interaction, and empty or set the experience pool and its priority array to a uniform initial value; randomly initialize the parameters of the two sets of Critic network and Actor network, and copy their respective parameters to the target network as the initial target network;
[0115] S4. Initialize the system state, including initializing the user position, task queue, battery level, channel status, and resetting the drone to its starting position. This step is the entry point for the outer training loop. During the outer training loop, subsequent steps are executed sequentially for each episode. The input to this step is the maximum number of training episodes and the initialized network and experience pool obtained in the previous step. The algorithm starts from episode = 0 and continues until it reaches the preset maximum value. Within each episode, the algorithm repeatedly interacts with the environment, collects data, and updates network parameters, opening a new interaction and learning phase for each episode. At the beginning of each episode, the environment's reset function is called. The environment regenerates or sets the user position, task queue, battery level, channel status, etc., according to preset rules, and resets the drone's position to its starting position. The input to this step is the current episode number and the environment object, and the output is the initial environment state of the episode, which contains all the environmental information required for this round of decision-making. For the current episode, execution proceeds sequentially from time step k = 0 to k = T-1, forming the inner time step loop. The input here is the initial state. Given a total number of steps T, the algorithm selects an action based on the current strategy at each time step, interacts with the environment, and records samples. The loop ends when the maximum number of steps is reached or a termination condition is encountered.
[0116] S5. Normalize the operating status of each device in the system according to the pre-given minimum and maximum values using a pre-designed linear scaling formula to obtain a normalized state, so that all features fall within a uniform numerical range; then, stitch together the normalized states into a window state.
[0117] S6. Sample the action at the current moment from the window state through the Actor network, and use the action to interact with the environment to obtain the next window state and the reward at the current moment. The input in this step is the sequence state and the current Actor network parameters. The Actor network first maps the long vector to a high-dimensional feature space through the embedding layer, and then the GRU unit extracts the temporal features. After passing through the KAN layer and several fully connected layers, a normalized action vector is output. Then, Gaussian exploration noise is added to the action space and clipped to the allowable range to obtain the final control action a[k]. Input a[k] into the environment dynamic model to obtain the next moment's environment state s[k+1], the immediate reward r(s[k], a[k]), and whether it is terminated. Finally, the triple (s[k+1], r[k], done) and the actual action executed are output.
[0118] S7. Store the samples obtained at the current time step into the experience pool, including the current window state, the current action, the reward at the current moment, and the next window state. This step packs the obtained samples into a quadruple (s, a, r, s'), writes it into the corresponding position in the experience pool, and assigns an initial priority to the sample. Finally, the updated experience pool and its priority array are output.
[0119] S8. Determine whether the number of samples in the experience pool has reached the set value. If yes, proceed to step S9; otherwise, return to step S5.
[0120] S9. Sample from the experience pool according to priority to obtain a training batch. The input of this step is all samples in the experience pool, their priorities, and the batch size. According to the priority experience replay rule, the probability of each sample being selected is set to P(n). A probability distribution is obtained through normalization. Then, samples that meet the batch size are randomly selected according to this distribution to form a training batch. The final output is a batch containing several samples. The sample set and its corresponding index;
[0121] S10. Calculate the importance sampling weights for samples in the same training batch, and normalize these weights to obtain a set of normalized importance weights. The inputs in this step are the sampling probability P(n) of the batch samples, the total number of valid samples in the current experience pool, and the hyperparameters of PER. The formula is used to calculate the importance sampling weights. Calculate the weight of each sample, then normalize it by dividing it by the maximum value of all sample weights to stabilize the weight values. Finally, output a set of normalized importance weights, which will be used to weight the TD error of each sample in the Critic loss.
[0122] S11. Based on the samples in the training batch, construct the target Q value of TD3 through the target network. In this step, the target action of the window state at each next time step in the current training batch is calculated through the target Actor network. Pruned Gaussian noise is added to the target action to obtain a smooth target action. The smooth target action is fed into two target Critic networks to calculate the corresponding Q value. The element-wise minimum of the two corresponding Q values is taken, and the target Q value of each sample is obtained through the Bellman equation, which is the target Q value of TD3.
[0123] S12. Based on the target Q value of TD3, perform weighted MSE training on the two Critic networks to obtain the TD error;
[0124] S13. Update the priority of samples in the experience pool using TD error; in this step, according to... The rules rewrite the values of these samples in the priority array and output the updated priority experience pool.
[0125] S14. Based on the policy delay mechanism, determine whether to update the Actor network in the current learning step. If yes, proceed to step S15; otherwise, return to step S9.
[0126] S15. Use the current first Critic network to update the parameters of the Actor network and output the updated Actor network.
[0127] S16. Perform soft updates on the target Actor network and the target Critic network, and output a set of updated target Actor network and target Critic network. In this step, the control action generated by the Actor network in the current window state is first obtained, and the Q value of the control action is calculated through the current first Critic network. The negative average Q value is used as the loss of the Actor network, and the Actor network parameters are updated through gradient descent to output the updated Actor network.
[0128] S17. Determine whether the target Actor network and the target Critic network have reached the end of training conditions. If so, use the latest control action obtained by the target Actor network as the control action of one or more devices with the goal of minimizing the maximum processing latency of all user tasks and the power consumption of the UAV. Otherwise, return to step S4.
[0129] The corresponding pseudocode is as follows:
[0130] .
[0131] In this embodiment, the Actor network and the target Actor network have the same structure, both including an embedding layer, a recurrent neural network, a modified KAN network, and a fully connected layer connected in sequence; wherein:
[0132] Embedding layers are used to map input objects to a high-dimensional feature space;
[0133] Recurrent neural networks are used to obtain temporal features of the output of the embedding layer;
[0134] Improve the KAN network to enhance the temporal features of the output of recurrent neural networks;
[0135] Fully connected layers are used to map the enhanced temporal features into control actions;
[0136] Once the controlled action interacts with the environment, it obtains the next window state and the reward for the current moment.
[0137] Recurrent Neural Networks (GRUs) significantly simplify the structure of LSTMs while retaining the ability to model long-term dependencies in sequences. Due to their more compact gating structure, they drastically reduce the number of parameters and computational overhead, making them particularly suitable for dynamic scenarios with stringent real-time requirements. For example... Figure 2 As shown, GRU mainly consists of two parts: the update gate and the reset gate. By adjusting the degree to which historical states participate in the current state update, the network can highlight key information relevant to decision-making in complex time-varying environments, thereby improving the stability and convergence efficiency of policy learning.
[0138] During the operation of GRU, the input at each time step consists of the current observation s[t] and the previous hidden state h[t-1]. The network first determines how much historical information to retain and how much new information to accept through an update gate. The update gate is a sigmoid function. Activation generation, its output is located in [0,1]. The larger the value, the stronger the preservation of past states. Its calculation form is:
[0139]
[0140] in This is the weight matrix. The term "biased" refers to the term that is biased.
[0141] Subsequently, GRU introduces a reset gate to control the influence of the previous hidden state on the current candidate hidden state. When the reset gate outputs a small value, historical information is significantly weakened, achieving a "reset" effect. Its calculation is as follows:
[0142] .
[0143] After obtaining the two types of gating, the GRU calculates candidate hidden states, which generate new activations by integrating the current input with the historical state modulated by the reset gate:
[0144]
[0145] in This indicates element-wise multiplication, used to reflect the reset gate's dimension-wise filtering of historical states. Finally, GRU uses the update gate to perform a weighted fusion between the old state and the candidate state to obtain the current hidden state.
[0146] .
[0147] The improved KAN network is a learnable function network based on the Kolmogorov–Arnold representation theorem, which can significantly improve the accuracy of mathematical fitting. A KAN with L layers will input... The output space is mapped layer by layer using learnable univariate functions. Let the first... The layer has 1 neuron, and defined as For the first The first layer neurons ( (e.g., 1, ..., L), its output is:
[0148]
[0149] in From the first Layer The first neuron to the second Layer A learnable unary activation function for each neuron. The output of the entire KAN is:
[0150] .
[0151] The activation function for each edge consists of a linear combination of a simple basis function and a B-spline function, in the form of:
[0152]
[0153] in These are trainable coefficients and basis functions. The spline part is defined as:
[0154]
[0155] in For learnable parameters, For the first The k-th B-spline basis function of the layer.
[0156] It is worth noting that directly introducing KAN often results in a more complex network structure and relatively higher computational and storage costs. Therefore, we appropriately reduced the hidden layer dimension and the number of intermediate network layers, thereby significantly reducing computational costs. Even so, thanks to the strong function approximation ability of KAN itself and the synergistic effect of the GRU components, the model still maintains good mathematical fitting performance.
[0157] In this embodiment, to improve sample utilization efficiency and accelerate the convergence speed of the value function, a Priority Experience Playback (PER) mechanism is adopted. Sampling weights are assigned based on the temporal difference error of the samples. The priority of the nth sample is defined as follows:
[0158]
[0159] in For TD error, To avoid smoothing terms with a priority of zero, PER employs a proportional priority strategy to construct a non-uniform sampling distribution based on this priority.
[0160]
[0161] in The impact of control priority on sampling probability. Since non-uniform sampling introduces estimation bias, PER uses importance sampling weights to correct for this bias.
[0162]
[0163] Where N is the number of valid samples. Control the strength of bias correction. To avoid numerical instability during training, further normalization is performed:
[0164]
[0165] In the critic update, these weights are used to adjust the contribution of TD error. As β gradually increases to 1, PER maintains unbiasedness while highlighting samples with high TD error, thereby improving the stability and convergence speed of learning.
[0166] This embodiment uses a dual-Q network structure to approximate the state-action value function, mitigating Q-value overestimation and improving training stability. The Critic network employs a two-way parallel feedforward network. and In the current state and actions Given the input, output the corresponding Q-value estimate.
[0167] In each round of training, a batch of interaction samples is sampled from the experience pool according to the PER mechanism described in the previous section:
[0168]
[0169] in Indicates batch size. For the sample index. For the next state. Using a delayed update policy network Generate target action:
[0170] ;
[0171] Then truncated Gaussian noise was added. To obtain the final target action:
[0172] ;
[0173] The target Q-network uses the action obtained in the previous step to calculate the target Q-value and constructs the TD target value:
[0174] ;
[0175] To update the parameters of Critic, we construct a weighted mean squared error loss based on priority sampling weights:
[0176] ;
[0177] Subsequently, the Q-network parameters are updated by performing gradient descent on the loss:
[0178]
[0179] in The learning rate is used. The target Q-network uses Polyak smooth updates, gradually approximating the current Q-network parameters with small step sizes.
[0180] ,
[0181] in These are the parameters for the updated target network; To update the step size; The parameters of the target network before the update; These are the parameters of the network in the decision model that correspond to the target network to be updated.
[0182] In this embodiment, the Actor network is based on the current state. Output deterministic actions It is used to execute policies in a continuous action space. Specifically, the temporal state of the input is first processed by the GRU layer to extract dynamic information of the environment and capture long-term temporal correlations, thereby generating a hidden vector h[t] that can comprehensively reflect the multi-step state features.
[0183] Subsequently, h[t] is fed into the KAN module to improve its ability to fit complex nonlinear mapping relationships. The features processed by KAN are then further mapped and compressed through fully connected layers. Finally, the output is scaled using the tanh activation function and action boundaries to generate the deterministic action at the current time step.
[0184]
[0185] in For action boundaries, For connection layer parameters, Let Q be the activation function. Based on this, the performance of the current policy is evaluated using a Critic network. The Actor is updated by maximizing the Q-value, and its loss is:
[0186] ;
[0187] The corresponding parameter update is:
[0188] .
[0189] In addition, to reduce strategy oscillations, an Actor update is performed only once after every two Critic updates.
[0190] In this embodiment, the environmental state consists of multidimensional heterogeneous variables, with vastly different value ranges for each dimension. If these raw states are input into a deep neural network without processing, problems such as certain dimensions dominating gradient updates, oscillations during training, and even difficulty in convergence can easily occur. To mitigate the impact of this numerical scale inconsistency, this paper performs linear normalization on the state vector, compressing each dimension to the same numerical range. Specifically, the following transformation is applied to the j-th state component:
[0191]
[0192] in The result after normalization. This represents the j-th dimension of the original state. These represent the theoretical minimum and maximum values, respectively. This preprocessing weakens state scale differences and reduces the interference of feature imbalance, thereby improving the stability and learning efficiency of the policy network.
[0193] In practice, the control module inputs the operating status of each device in the system into the pre-trained TD3-SPGK policy network, generating multi-dimensional control actions through forward inference. These actions include: UAV 3D position, flight direction and speed commands, optional task load switching commands; STAR-RIS reflection / transmission coefficient configuration, phase shift matrix or cell grouping settings; SWIPT information signal power and energy signal power allocation; task offloading ratio; and edge server computing resource allocation.
[0194] Since the strategy has already learned the optimal value judgment during the training phase, there is no need to calculate the reward function or cost function in real time during this phase, enabling the control process to have a millisecond-level response speed.
[0195] The control module sends the generated control actions to the drone, STAR-RIS, and terminal devices, respectively. Each device then performs the following actions based on the received actions:
[0196] a. The UAV executes trajectory and energy control commands:
[0197] Adjust flight path to maintain communication coverage and avoid energy depletion.
[0198] b. STAR-RIS Dynamic Channel Reconfiguration:
[0199] The reflected and transmitted beams are updated in real time according to the phase / amplitude parameters output by the strategy, thereby enhancing the performance of the communication link.
[0200] c. Task scheduling and execution on terminals and edge servers:
[0201] Tasks are uploaded, energy is collected, data is calculated, or cached according to the unloading strategy.
[0202] The network operates continuously according to policy instructions, and the output control commands are sent directly from the central control unit to the drones and each terminal node, realizing dynamic optimization and continuous operation of the network.
[0203] When the environment changes, such as adjusting drone power or redeploying STAR-RIS, operators can reset key parameters in the simulation environment to generate simulation instances that match the current state. In the modified environment, pre-trained models can be quickly calibrated with a small number of samples, including: fine-tuning the parameters of the policy and value networks; updating the PER cache using incremental empirical data; and performing lightweight training via edge servers to avoid large-scale model backhaul.
[0204] In summary, this invention constructs a UAV-assisted mobile edge computing post-disaster emergency communication system integrating STAR-RIS, SWIPT, and TD3-SPGK intelligent strategies. It achieves intelligent collaborative optimization across the entire process from "offline training to online deployment to environmental adaptation," offering significant advantages over existing technologies: First, offline training based on a high-fidelity simulation environment, combined with GRU temporal feature extraction and improved KAN structure nonlinear representation, enables the TD3-SPGK agent to accurately perceive the dynamic coupling of communication, computing, and energy in complex post-disaster scenarios, obtaining a stable and highly generalizable control strategy under multi-objective constraints. Second, during the online deployment phase, the central control unit provides real-time perception of UAV position and energy, terminal task and energy status, STAR-RIS configuration, and channel quality. The policy network directly outputs continuous control actions such as UAV trajectory control, STAR-RIS transmission / reflection coefficients, SWIPT power allocation, and task offloading ratios. This eliminates the need for real-time solving of complex optimization problems, achieving dynamic optimal network scheduling within milliseconds, significantly reducing terminal task processing latency and UAV energy consumption, and improving system coverage performance and resource utilization efficiency. Furthermore, by utilizing a dual-Q network, delayed updates, target policy smoothing, and a sample reuse mechanism combining SN and PER, the algorithm effectively suppresses value estimation bias, improves training stability, and accelerates convergence, ensuring good robustness even in high-dimensional continuous action spaces. Finally, through a reconfigurable simulation environment and an incremental online fine-tuning mechanism, the system can quickly complete policy migration and seamless updates under environmental changes, achieving long-term adaptive capability to highly dynamic and uncertain post-disaster scenarios. This effectively reduces terminal task processing latency and system energy consumption while simultaneously ensuring terminal energy sustainability.
Claims
1. A post-disaster emergency communication system, characterized in that, This includes reconfigurable smart surfaces, drones, and control modules; among which: The drone is equipped with an edge server for adjusting the spatial location of the edge server; the edge server is used to communicate with the reconfigurable smart surface and / or the user. Reconfigurable smart surfaces can be deployed on the ground and respond to the communication needs of users and / or drones. By controlling the transmission and reflection of incident signals in the air-to-ground link, they can diffract or avoid obstacles in post-disaster scenarios and improve signal coverage. The control module is used to acquire and, based on the operating status of each device in the system, generate control actions for one or more devices with the goal of minimizing the maximum processing latency of all user tasks and the energy consumption of the drone while ensuring the user's power needs, thereby achieving the ability to adapt to changes in the post-disaster environment; where the drone's energy consumption includes the energy consumption of the edge server. The expression aimed at minimizing the maximum processing latency of all user tasks and the power consumption of the drone is: in The proportion of tasks offloaded to edge servers during time slot k; This represents the location of the drone in time slot k, which is also the location of the edge server. ; Let k be the flight speed of the drone, which is also the flight speed of the edge server. Let K be the flight angle of the UAV in time slot k, which is also the flight angle of the edge server. and These are the reflection phase matrix and transmission phase matrix of the reconfigurable smart surface, respectively, used to characterize the phase modulation applied to the reflected and transmitted signals by each reconfigurable smart surface; K is the total number of time slots in the entire communication period T, and k is the time slot index; I is the total number of users. Index for users; For time slot k, the first Maximum processing latency for a single user task; Processing the kth time slot Drone energy consumption per user task; Indicates constraints; Indicates the initial position of the drone; For flight time; This refers to the maximum flight speed of the drone; and respectively drones in shaft and The farthest flight distance of the axis; For time slot k, the first The location of each user. ; and The first Individual users shaft and The furthest distance the axis can travel; Let k be the flight energy consumption of the UAV in time slot k. Let k be the data transmission energy consumption of the UAV in time slot k. The computing power consumption of the edge server is k timeslot k. To maximize the battery power of the drone; and These are the user's minimum and maximum battery limits, respectively. For time slot k, the first Battery level of each user; This indicates the reflection phase shift at time slot k; This indicates the transmission phase shift at time slot k; Indicates the first The total number of tasks for each user; D represents all tasks; A specific method for generating control actions for one or more devices with the goal of minimizing the maximum processing latency of all user tasks and the power consumption of the drone includes the following steps: S1. Construct a decision model based on Actor networks and dual-Q networks; where the dual-Q network includes two sets of Critic networks; S2. Obtain physical constraints and set decision model hyperparameters; where physical constraints include the spatial locations of all users, the location of the reconfigurable smart surface, the initial location of the drone and the total time duration, the maximum available energy of the drone, and the user's battery threshold; decision model hyperparameters include the reward function, discount factor, learning rate of the Actor network and two Critic networks, soft update coefficient of the target network, noise level, pruning boundary, and policy delay steps; the target network includes the target Actor network and the target Critic network; S3. Initialize the reinforcement learning subject: Construct an experience pool to store the samples generated by the interaction, and empty or set the experience pool and its priority array to a uniform initial value; randomly initialize the parameters of the two sets of Critic network and Actor network, and copy their respective parameters to the target network as the initial target network; S4. Initialize system status, including initializing user location, task queue, battery level, channel status, and resetting the drone to its starting position; S5. Normalize the operating status of each device in the system according to the pre-given minimum and maximum values to obtain the normalized state; and combine the normalized states into a window state. S6. Sample the action at the current moment from the window state through the Actor network, and use the action to interact with the environment to obtain the next window state and the reward at the current moment; S7. Store the samples obtained at the current time step into the experience pool, including the current window state, current action, current reward, and next window state. S8. Determine whether the number of samples in the experience pool has reached the set value. If yes, proceed to step S9; otherwise, return to step S5. S9. Sample from the experience pool according to priority to obtain a training batch; S10. Calculate the importance sampling weights for samples in the same training batch, and normalize the importance sampling weights to obtain a set of normalized importance weights. S11. Based on the samples in the training batch, construct the target Q value of TD3 through the target network; S12. Based on the target Q value of TD3, perform weighted MSE training on the two Critic networks to obtain the TD error; S13. Update the priority of samples in the experience pool using TD error; S14. Based on the policy delay mechanism, determine whether to update the Actor network in the current learning step. If yes, proceed to step S15; otherwise, return to step S9. S15. Use the current first Critic network to update the parameters of the Actor network and output the updated Actor network. S16. Perform soft updates on the target Actor network and the target Critic network, and output a set of updated target Actor network and target Critic network. S17. Determine whether the target Actor network and the target Critic network have reached the end of training conditions. If so, use the latest control action obtained by the target Actor network as the control action of one or more devices with the goal of minimizing the maximum processing latency of all user tasks and the power consumption of the UAV. Otherwise, return to step S4. The expression for the reward is: , and These are the weighting coefficients. A penalty is introduced to limit user battery usage. , A penalty coefficient greater than 0 This is the power threshold.
2. The post-disaster emergency communication system according to claim 1, characterized in that, The link formed when a user transmits data directly to the edge server is called a direct link. The link formed when a user transmits data to the edge server through the reconfigurable smart surface is called an auxiliary link. Auxiliary links are divided into reflection links and transmission links. When the edge server and the user are on the same side of the reconfigurable smart surface, the auxiliary link formed is a reflection link. When the edge server and the user are on opposite sides of the reconfigurable smart surface, the auxiliary link formed is a transmission link.
3. The post-disaster emergency communication system according to claim 1, characterized in that, It also includes a wireless information and energy transmission module, which integrates data transmission and energy harvesting in radio frequency signals into the same physical process, providing wireless power supply to one or more devices in the system.
4. The post-disaster emergency communication system according to claim 1, characterized in that, The operating status of each device in the system includes: Drone status: 3D position information, remaining energy, current payload, and flight speed; Ground terminal status: coordinates, energy level, amount of computational tasks to be processed; Reconfigurable smart surface status: configuration status and controllable unit activation status; Environmental parameters: link quality, obstacle distribution, and wireless interference level.
5. The post-disaster emergency communication system according to claim 1, characterized in that, The Actor network and the target Actor network have the same structure, both including sequentially connected embedding layers, recurrent neural networks, improved KAN networks, and fully connected layers; where: Embedding layers are used to map input objects to a high-dimensional feature space; Recurrent neural networks are used to obtain temporal features of the output of the embedding layer; Improve the KAN network to enhance the temporal features of the output of recurrent neural networks; Fully connected layers are used to map the enhanced temporal features into control actions; Once the controlled action interacts with the environment, it obtains the next window state and the reward for the current moment.
6. The post-disaster emergency communication system according to claim 1, characterized in that, Specific methods for constructing the target Q value of TD3 through the target network include: The target action of the window state at each next time step in the current training batch is calculated by the target Actor network, and clipped Gaussian noise is added to the target action to obtain a smooth target action; The smoothed target action is fed into two target Critic networks to calculate the corresponding Q value. The element-wise minimum of the two corresponding Q values is taken, and the target Q value of each sample is obtained through the Bellman equation, which is the target Q value of TD3.
7. The post-disaster emergency communication system according to claim 1, characterized in that, The specific methods for updating the parameters of the Actor network using the current first Critic network and outputting the updated Actor network include: Obtain the control action generated by the Actor network in the current window state, calculate the Q value of the control action through the first Critic network, use the negative average Q value as the loss of the Actor network, update the Actor network parameters through gradient descent, and output the updated Actor network.
8. The post-disaster emergency communication system according to claim 1, characterized in that, Specific methods for soft updating the target Actor network and the target Critic network include: For the target Actor network and the target Critic network, soft updates are performed according to the following formula: in These are the parameters for the updated target network; To update the step size; The parameters of the target network before the update; These are the parameters of the network in the decision model that correspond to the target network to be updated.
Citation Information
Patent Citations
Mobile edge computing task unloading optimization method based on unmanned aerial vehicle and RIS assistance
CN120201497A