Digital twinborn enhanced unknown environment unmanned aerial vehicle three-dimensional trajectory optimization method

By constructing a digital twin-enhanced UAV 3D trajectory optimization method, combined with a system model and the TD3 algorithm, the efficiency and security issues of UAV communication services in unknown environments were solved, achieving efficient and secure communication services in complex environments.

CN121165764APending Publication Date: 2025-12-19BEIJING INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511525636.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing UAV trajectory optimization methods struggle to provide efficient and safe communication services to ground users in unknown environments, facing challenges such as large discrepancies between simulation and reality, high computational costs, and high collision risks.

Method used

We construct a digital twin-enhanced method for optimizing the 3D trajectory of unmanned aerial vehicles (UAVs) in unknown environments. Through system models and joint optimization problems, we combine a digital twin-driven trajectory framework, simulated annealing user scheduling, and Markov decision processes. We use the TD3 algorithm to train a path optimization model, generate a high-fidelity virtual environment for training and real-time perception, and optimize the UAV trajectory.

Benefits of technology

In unknown environments, the UAVs were able to provide communication services to ground users efficiently and safely, dynamically plan trajectories to avoid obstacles and shorten the total mission time, and improve the robustness and safety of learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121165764A_ABST
    Figure CN121165764A_ABST
Patent Text Reader

Abstract

The invention provides a digital twinborn enhanced unknown environment unmanned aerial vehicle three-dimensional trajectory optimization method, which relates to the field of wireless communication, and comprises the following steps: constructing a system model and a joint optimization problem; a digital twinning driven trajectory framework is constructed, the framework comprising a physical entity layer, a digital twinning layer and a connection between the physical entity layer and the digital twinning layer, the digital twinning layer running on a digital twinning server, the digital twinning layer running on the digital twinning server during task execution, and the digital twinning layer running on the digital twinning server. The digital twin server uses the constructed virtual environment to continuously monitor actions of the unmanned aerial vehicle and assess potential safety risks; user scheduling based on simulated annealing is designed, a joint optimization problem is modeled as a Markov decision process, and an unmanned aerial vehicle is used as an intelligent agent to maximize accumulated rewards related to task completion time and collision avoidance; and training the path optimization model by using a TD3 algorithm. According to the invention, the unmanned aerial vehicle can efficiently and safely provide communication service for ground users in an unknown environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of wireless communication, in particular to a method for three-dimensional trajectory optimization of unmanned aerial vehicles in unknown environments based on digital twinning enhancement. BACKGROUND

[0002] Unmanned aerial vehicles (UAVs) are transforming the current low-altitude economic ecosystem by driving low-altitude wireless networks (LAWN) to provide seamless and flexible connectivity for ground users. With their three-dimensional (3-D) maneuverability, UAVs can rapidly expand the connectivity range in areas with limited infrastructure, during emergencies, or under dynamic network demand. However, the deployment of UAV-based low-altitude wireless networks faces key challenges: limited onboard computing power, collision risks, and fluctuating service demand. To overcome these challenges, efficient UAV trajectory design is crucial for ensuring reliable communication and safe operation in complex environments.

[0003] In the prior art, research on optimizing UAV trajectories mainly focuses on convex optimization frameworks, but they face scalability bottlenecks and high computational overhead, which hinders their applicability in large-scale and dynamic network environments. To overcome these limitations, deep reinforcement learning (DRL) methods that can achieve adaptive decisions through trial-and-error interactions with the environment have been proposed. However, most DRL schemes rely on idealized simulations that fail to capture real-world uncertainties in low-altitude wireless network deployments, especially when highly accurate reconstructions of physical conditions are not available. Furthermore, due to model instability, slow convergence speed, and potential collision risks, training directly in the physical environment remains costly and risky. Therefore, how to narrow the gap between simulation and reality has become a key challenge for DRL-driven UAV trajectory optimization to move towards practical applications.

[0004] To bridge this gap, the Digital Twin (DT) framework was proposed, which builds dynamic Virtual Environments (VEs) that are synchronized with the physical world by integrating historical and real-time data. Open-source platforms such as Mago3D facilitate the creation of digital twins, enabling DRL agents to train in high-fidelity virtual environments to gain a large amount of experience, thereby accelerating convergence and improving robustness. Research on UAVs has demonstrated the effectiveness of digital twins in energy-efficient trajectory design, spectrum allocation, dynamic task allocation, and computation offloading optimization. However, deploying digital twins for UAVs in unknown environments remains challenging due to the lack of reliable prior data to initialize the digital twin. SUMMARY

[0005] The main purpose of the present application is to propose a digital twin enhanced three-dimensional trajectory optimization method for unmanned aerial vehicles in unknown environments, enabling unmanned aerial vehicles to efficiently and safely provide communication services for ground users in unknown environments.

[0006] The present application is implemented by the following technical solutions: The digital twin enhanced three-dimensional trajectory optimization method for unmanned aerial vehicles in unknown environments comprises the following steps: Step S1, constructing a system model and a joint optimization problem, the system model including a scenario for path planning of the unmanned aerial vehicle in the unknown environment, a time model for path planning of the unmanned aerial vehicle based on the scenario, a flight model for path planning of the unmanned aerial vehicle based on the scenario and the time model, and a signal and communication model between the unmanned aerial vehicle, the ground user and the base station based on the scenario, the joint optimization problem being constructed based on the system model with the objective of minimizing the task completion time; Step S2, based on the system model and the joint optimization problem, constructing a digital twin driven trajectory framework, the framework including a physical entity layer, a digital twin layer and a connection between the physical entity layer and the digital twin layer, the digital twin layer running on a digital twin server, the digital twin server continuously monitoring the actions of the unmanned aerial vehicle and evaluating potential safety risks using the constructed virtual environment during task execution; Step S3, based on the system model and the joint optimization problem, designing a user scheduling based on simulated annealing, and modeling the joint optimization problem as a Markov decision process, the unmanned aerial vehicle as an agent aiming to maximize the cumulative reward related to the task completion time and collision avoidance; Step S4, based on the system model, the user scheduling and the Markov decision process, training the path optimization model using the TD3 algorithm.

[0007] Further, in step S1, constructing the system model specifically comprises the following steps: Step S11, construct the scene of path planning of UAV in unknown environment: in the task area The ground has buildings and K ground users randomly distributed, the building occupies an area , the base station transmits communication data to the UAV, the UAV relays the data to the ground user using the ISAC signal, the communication coverage of the UAV is greater than its perception coverage; Step S12, construct the time model of UAV path planning: the entire task duration is divided into T time slots of equal length, the motion decision of the UAV is made at the beginning of each time slot, the UAV provides services to the ground users in TDMA mode, only the closest and unserved ground users are served in each time slot; Step S13, construct the UAV flight model of UAV path planning: in time slot t+1, the position coordinates of the UAV are According to the formula , , Update, where is the position coordinates of the UAV in time slot t, is the flight speed of the UAV in time slot t, is the vertical flight angle of the UAV in time slot t, is the horizontal flight angle of the UAV in time slot t, is the duration of each time slot; Step S14, construct the signal and communication model between the UAV, the ground user and the base station: the signal and communication model includes the base station-UAV link and the UAV-ground user link, the transmission rate of the base station-UAV link is , the transmission rate of the UAV-ground user link is , the effective communication rate of the UAV to the ground user k in time slot t is , where B represents the available frequency bandwidth, represents the constant transmission power of the base station, represents the power of zero-mean additive Gaussian white noise, , represents the large-scale fading of the base station-UAV link in time slot t, represents the small-scale fading of the base station-UAV link in time slot t, represents the transmission power of the UAV, , represents the large-scale fading of the UAV-ground user link in time slot t, represents the small-scale fading of the UAV-ground user link in time slot t.

[0008] Further, in the step S1, the joint optimization function is represented as wherein, is a set of speeds of the UAV throughout the mission duration, , is a maximum speed of the UAV, is a set of vertical flight angles of the UAV throughout the mission duration, is a set of horizontal flight angles of the UAV throughout the mission duration, is a minimum number of time slots required by the UAV to serve the ground user k, , is an amount of data transmitted by the UAV to the ground user k at time slot t, is a throughput requirement of the ground user k, is a set of ground users.

[0009] Further, in the step S2, in the digital twin driven trajectory framework, the physical entity layer includes the physical UAV, the task execution area and the ground users waiting for service, the UAV as an agent is responsible for observing the environment, executing the control actions received from the base station, providing communication services for the ground users, and uploading the perception data to the digital twin server; the virtual environment constructed by the digital twin server captures the mobility of the UAV, the distribution of the ground users and the environmental features including buildings, based on the virtual environment, the three-dimensional trajectory planning process is carried out in the digital twin layer; the physical entity layer and the digital twin layer are connected through a bidirectional link to ensure synchronization and coordination; during the task execution, before each action is transmitted to the UAV, the digital twin server checks whether the execution action will collide with obstacles or there is a boundary crossing behavior, if so, the action is cancelled and the UAV is instructed to remain stationary in that time slot.

[0010] Further, in the step S3, the user scheduling based on simulated annealing specifically includes the following steps: Step S31, initialization based on the nearest greedy strategy: starting from the initial position of the UAV, the nearest unserved ground user is selected from the remaining ground user set to construct a greedy scheduling vector ; Step S32, scheduling optimization: taking the scheduling vector as the initial scheduling, in each iteration, a candidate solution is generated by applying a neighborhood operator to the current scheduling, the objective function of iteration is defined as the total flight distance of the UAV serving all ground users, if the candidate solution can shorten the total distance, the candidate solution is received, otherwise, according to the Metropolis criterion, a certain probability is accepted, the probability decreases gradually with the cooling of the temperature, and the known best scheduling vector is updated whenever a shorter distance is found, wherein the neighborhood operator set is .

[0011] Further, the step S3 of modeling the joint optimization problem as a Markov decision process specifically comprises the following steps: Step S33, constructing state space: State space including every time slot all possible states , state is defined as the combination of two parts , matrix and matrix are both matrix, m and n correspond to the grid cell indices used to discretize the task area, matrix facilitates the communication service provision of the UAVs to the ground users according to the user dispatch vector, in matrix, the UAV is represented as a circular area with its position as the center, as the radius, and the corresponding value is , the ground user k is attracted by the UAV according to the user dispatch vector with a normal number in descending order, matrix realizes the obstacle avoidance of the UAV by integrating its position and the detected building information, in matrix, the UAV is represented as a circular area with its position as the center, as the radius, and the corresponding value is , the detected building is discretized in matrix according to its position perceived by the UAV, and the discretized value corresponds to its perceived height; Step S34, constructing action space: Action space contains all feasible actions of every time slot t , each action represents the flight decision of the UAV, and the flight decision includes the speed, vertical flight angle and horizontal flight angle of the UAV, represented as ; Step S35, constructing reward space: Reward is defined as , where , , is the collision penalty, , is the speed penalty, , , is the time penalty, , is the service reward, , , is the distance reward, , The distance reward threshold, Let be the change in distance between the UAV and the target ground user at time slot t. To explore rewards, , The change in the number of newly covered grid cells within the UAV communication range during time slot t. To complete the reward, , The moment the task is completed; Step S36: Construct state transition probabilities: State transition probabilities represent the probability of an agent transitioning from state S36 to state S36. Execute action When, transition to the next state. The possibility is expressed as .

[0012] Furthermore, step S4 specifically includes the following steps: Step S41: Initialize the online Actor network for the TD3 algorithm. Target Actor Network Two online Critic networks , and two target Critic networks and ; Step S42: Based on initialization, the digital twin server generates a state according to the created virtual environment. and pass it on to the online Actor network. The network deterministically maps this state to an action. To balance exploration and exploitation, random noise is added to the action, which is then represented as follows: ,in, The mean is zero and the variance is Additive white Gaussian noise, For the range of action constraints, , , This is the clipping function; Step S43: The drone agent performs actions. Receive rewards Observe the next state , experience tuple Stored in the playback buffer; Step S44: Determine whether the number of experience tuples stored in the buffer has reached the set threshold. If yes, proceed to step S45; otherwise, proceed to step S41. Step S45: For each empirical tuple, compute a smooth target action using the target Actor network. , represented as , is an additive white Gaussian noise with zero mean, is the maximum exploration noise; Step S46, calculating the target Q value , ; Step S47, updating the two online Critic networks by minimizing the time difference error between the predicted Q value and the target Q value, to obtain the updated online Critic networks and , wherein the loss function used for updating is , represents the number of experience tuples uniformly sampled from the buffer; Step S48, judging whether the number of updates of the two online Critic networks reaches a set threshold, if yes, entering step S49, otherwise, entering step S50; Step S49, updating the online Actor network with the maximum expected Q value of its action as the update target; Step S50, updating the target Actor network and the target Critic network in a soft update manner , is the update rate; Step S511, judging whether the time slot index reaches T, if yes, entering step S512, otherwise, entering step S51; Step S512, ending the current round of training, resetting the environment, and observing the initial state , judging whether the number of training rounds reaches a set threshold, if yes, completing the training, otherwise, entering step S51.

[0013] From the above description of the present application, compared with the prior art, the present application has the following beneficial effects: The present application firstly constructs a system model and a joint optimization problem, the system model constructed includes a scene of path planning of the unmanned aerial vehicle in an unknown environment, a time model of path planning of the unmanned aerial vehicle based on the scene, a flight model of path planning of the unmanned aerial vehicle based on the scene and the time model, and a signal and communication model between the unmanned aerial vehicle, the ground user and the base station based on the scene, the joint optimization problem takes minimizing the task completion time as the target, then based on the system model and the joint optimization problem, a digital twin driven trajectory framework is constructed and a user scheduling based on simulated annealing is designed, the joint optimization problem is modeled as a Markov decision process, finally based on the system model, the user scheduling and the Markov decision process, a path optimization model is trained by using a TD3 algorithm, after the training is completed, the path optimization model can be used for three-dimensional trajectory optimization of the unmanned aerial vehicle. In the process, the digital twin is used to construct a virtual environment highly consistent with the real environment, in the training stage, the digital twin can generate diversified and realistic virtual scenes, accelerate the convergence of the model and improve the robustness of learning, in the deployment stage, the unmanned aerial vehicle transmits an ISAC signal and receives its echo, perceives the environmental information in real time and dynamically updates the virtual environment in the digital twin, thereby supporting the unmanned aerial vehicle to make continuous trajectory optimization decisions, by using a continuous action space to make decisions, dynamic challenges in the unknown environment can be accurately coped with, and then an optimal trajectory that can effectively avoid obstacles, ensure flight safety and significantly shorten the total task time is dynamically planned, so that the unmanned aerial vehicle can efficiently and safely provide communication services for the ground user in the unknown environment. BRIEF DESCRIPTION OF DRAWINGS

[0014] The present application will be further described below in combination with the drawings and specific embodiments.

[0015] Figure 1 The flowchart of the present application.

[0016] Figure 2 The flowchart of the user scheduling based on simulated annealing of the present application.

[0017] Figure 3 The flowchart of the path planning training of the present application by using the TD3 algorithm.

[0018] Figure 4 The simulation comparison diagram of the training reward changing with the training round of the present application and the comparison method.

[0019] Figure 5 The simulation comparison diagram of the UAV providing time for all users changing with the user quantity of the present application and the comparison method. DETAILED DESCRIPTION

[0020] The present application will be further described below in combination with the drawings and specific embodiments.

[0021] As Figure 1 shown, the digital twin enhanced unknown environment unmanned aerial vehicle three-dimensional trajectory optimization method includes the following steps: Step S1, constructing a system model and a joint optimization problem, the system model including a scenario of path planning of an unmanned aerial vehicle in an unknown environment, a time model of path planning of the unmanned aerial vehicle based on the scenario, a flight model of path planning of the unmanned aerial vehicle based on the scenario and the time model, and a signal and communication model between the unmanned aerial vehicle, a ground user and a base station based on the scenario, the joint optimization problem being constructed based on the system model, with the goal of minimizing task completion time; Specifically, constructing the system model specifically includes the following steps: Step S11, constructing a scenario of path planning of an unmanned aerial vehicle in an unknown environment: in a task area The ground has buildings and K=10 randomly distributed ground users, whose set is represented as ={1, 2,..., K}, the position of the kth ground user being represented as a vector , and being a fixed horizontal coordinate. There are buildings in the task area, which occupy an area In this embodiment, the base station or the unmanned aerial vehicle is pre-aware of the position information of the ground users, so that the digital twin server can perform more efficient scheduling and trajectory involvement.

[0022] The base station transmits communication data to the unmanned aerial vehicle, and the unmanned aerial vehicle relays the data to the ground user using ISAC signals. To improve energy efficiency, the unmanned aerial vehicle uses a directional antenna with a beam width of , the angle and the height of the unmanned aerial vehicle at time slot t together determining the coverage range of the unmanned aerial vehicle. Because the communication signal experiences one-way propagation, while the perception echo needs to experience two-way propagation, the communication coverage range of the unmanned aerial vehicle is greater than its perception coverage range, i.e. In this embodiment, is π / 4, is π / 6; Step S12, constructing a time model of path planning of the unmanned aerial vehicle: the entire task duration is divided into T=500 time slots of equal length, the motion decision of the unmanned aerial vehicle being made at the beginning of each time slot, the unmanned aerial vehicle using TDMA to provide services for the ground users, and only serving the closest ground user that has not completed service in each time slot; The index of the divided time slot is , and the duration of each time slot is The motion decision of the UAV is made at the beginning of each time slot. The UAV provides services for the ground users in a TDMA manner, and only serves the closest ground user that has not been served in each time slot. When the data transmission needs of all ground users are met, the task is declared to be completed.

[0023] Step S13, constructing a UAV flight model for UAV path planning: at time slot t+1, the position coordinates of the UAV are According to the formula , , update, wherein, is the position coordinates of the UAV at time slot t, and are time-varying horizontal coordinates, is a time-varying vertical coordinate, and the UAV updates its position time slot by time slot according to the three-coordinate update rule, is the flight speed of the UAV at time slot t, is the vertical flight angle of the UAV at time slot t, is the horizontal flight angle of the UAV at time slot t; Step S14, constructing a signal and communication model between the UAV, the ground user and the base station: the signal and communication model includes a base station-UAV link and a UAV-ground user link. Since the UAV operates as a relay node, its effective communication rate depends on the transmission rate of the two links.

[0024] The UAV provides services for the ground users in a TDMA manner. In order to support the relay communication task of the UAV, each time slot is further divided into three sub-slots, namely, a base station-UAV sub-slot, a UAV-ground user sub-slot and a UAV-base station sub-slot. The time lengths of these sub-slots are , and , and satisfy In this embodiment, The functions of the three sub-slots are summarized as follows: Base station-UAV sub-slot: the base station transmits control instructions and data signals to the UAV, and the UAV receives these signals; UAV-ground user sub-slot: the UAV forwards the received data signals to the ground user using the ISAC signal, and collects environmental perception echoes; UAV-base station sub-slot: the UAV uploads the perception data to the digital twin server through the base station, updates the virtual environment, and generates a new flight strategy for the next time slot.

[0025] To reflect the actual propagation environment, the channel model contains both large-scale path loss and small-scale fading. Considering the base station height and the aerial position of the UAV, the base station-UAV link is assumed to be line-of-sight (LoS) transmission condition. The large-scale fading of this link at time slot t is modeled as where is the shadowing fading in the LoS link, is the path loss, is the Euclidean distance between the base station and the UAV at time slot t, is the carrier frequency, and c is the speed of light.

[0026] Therefore, the channel gain is expressed as where is the small-scale fading of the base station-UAV link at time slot t, which captures the fast fluctuations of the channel, and B denotes the available spectral bandwidth, is the constant transmit power of the base station, is the power of the zero-mean additive white Gaussian noise.

[0027] For the UAV-ground user link, the air-to-ground (A2G) channel is classified as a line-of-sight (LoS) or non-line-of-sight (NLoS) channel according to the real-world topology constructed by the digital twin server, the reported ground user locations, and the position of the UAV. The communication range is defined as If the distance between the UAV and the ground user k exceeds , the link does not exist, and if it does not exceed, the digital twin server will evaluate whether the link between the UAV and the ground user k is blocked by buildings and assign for the LoS link and for the NLoS link.

[0028] The large-scale fading of the channel corresponding to the UAV-ground user link is expressed as , and the path loss is Therefore, the transmission rate of the UAV-ground user link is where is used to describe the shadowing effect in the NloS link, is the Euclidean distance between the UAV and the kth ground user at time slot t, denotes the transmit power of the UAV, is the channel gain, is the small-scale fading of the UAV-ground user link at time slot t, which describes the fast changes of the channel.

[0029] Therefore, the effective communication rate of the UAV to the ground user k at time slot t is​​​ The amount of data transmitted to ground user k at time slot t is To meet the throughput requirement of ground user k The minimum number of time slots required for the UAV to serve this user is Must satisfy Once the UAV has transmitted the required amount of data to all ground users, the mission is considered complete. Better channel quality would increase allowing more data to be transmitted at each time slot, reducing the total number of time slots required, and thus shortening the mission time.

[0030] In the single-UAV scenario, with no inter-UAV interference and ground user scheduling based on TDMA, it is reasonable to set to a constant maximum value.

[0031] In this example, the shadow fading of the LoS link is 0.1 dB, while the shadow fading of the NLoS link is 21 dB, the throughput requirement of each ground user is 10 Mbits, and the power of the additive white Gaussian noise (AWGN) is -75 dBm.

[0032] The goal of the present invention is to optimize the trajectory of the UAV to minimize the mission completion time while ensuring that obstacles are avoided, then the joint optimization function is expressed as where, is the set of speeds of the UAV throughout the duration of the mission, , is the maximum speed of the UAV, is the set of vertical flight angles of the UAV throughout the duration of the mission, is the set of horizontal flight angles of the UAV throughout the duration of the mission, is the minimum number of time slots required for the UAV to serve ground user k, , is the amount of data transmitted by the UAV to ground user k at time slot t, is the throughput requirement of ground user k, is the set of ground users. The constraints limit the speed, vertical angle, and horizontal angle of the UAV, respectively. and ensure that the UAV remains within the designated mission area and avoids colliding with buildings. Guarantee that the cumulative amount of data transmitted to each ground user meets the minimum requirement.

[0033] Step S2, based on the system model and the joint optimization problem, a digital twin driven trajectory framework is constructed, which includes a physical entity layer, a digital twin layer and a connection between the physical entity layer and the digital twin layer, the digital twin layer runs on a digital twin server, during task execution, the digital twin server uses the constructed virtual environment to continuously monitor the actions of the unmanned aerial vehicle and evaluate potential safety risks; Specifically, in the digital twin driven trajectory framework, the physical entity layer represents the real-world task scene, the digital twin layer is the virtual counterpart of the physical entity layer, and the connection is used to realize the bidirectional interaction between the two. The physical entity layer specifically includes a physical unmanned aerial vehicle, a task execution area and a ground user waiting for service. Due to its limited computing and storage capacity, the unmanned aerial vehicle cannot train the path optimization model by itself, nor can it store a large amount of perception data generated during the task. Therefore, the unmanned aerial vehicle only serves as an agent, responsible for observing the environment, executing the control actions received from the base station, providing communication services for the ground user, and uploading perception data to the digital twin server; The digital twin layer runs on a digital twin server with rich computing and storage resources, responsible for building and maintaining a high-fidelity virtual copy of the physical entity. This virtual environment captures the mobility of the unmanned aerial vehicle, the distribution of the ground user, and the environmental characteristics of the building, etc. Based on this virtual environment, the path optimization model is trained and executed in the digital twin layer, so as to generate effective and safe control actions to guide the unmanned aerial vehicle; The unmanned aerial vehicle continuously uploads perception data, which enables the digital twin layer to gradually improve the virtual environment. As feedback, the digital twin layer generates actions through the trained path optimization model and transmits them to the unmanned aerial vehicle through the base station for execution. In order to ensure the reliable and timely consistency between the unmanned aerial vehicle and its virtual counterpart, sufficient communication resources must be allocated to maintain a stable link.

[0034] Since the topology of the task area is unknown in advance, directly training the model using real scene data has certain challenges, and unknown buildings may threaten the safety of the unmanned aerial vehicle. To solve this problem, we use digital twin technology, which supports offline model training before task execution and provides continuous safety assurance during flight. In the training phase, the digital twin server generates multiple virtual environments, allowing virtual unmanned aerial vehicles to interact with diverse simulated task scenes. These interactions generate experiences of Markov decision processes, which are stored in a replay buffer for training the path optimization model. With the rich computing resources of the digital twin server, the training process is significantly accelerated, so that a fully optimized model can be obtained, which is then deployed on the virtual unmanned aerial vehicle to enhance its decision-making ability.

[0035] The virtual environment built by the digital twin server captures the mobility of the UAV, the distribution of the ground users, and the environmental features including buildings, based on which a three-dimensional trajectory planning process is performed in the digital twin layer; the physical entity layer and the digital twin layer are connected through a bidirectional link to ensure synchronization and coordination; during task execution, before each action is transmitted to the UAV, the twin server checks whether the action will collide with obstacles or have a boundary crossing behavior, and if so, the action is canceled and the UAV is instructed to remain stationary in that time slot. This closed-loop verification ensures that the UAV can operate safely under the guidance of the trained path optimization model.

[0036] Step S3, based on the system model and the joint optimization problem, a user scheduling based on simulated annealing is designed, and the joint optimization problem is modeled as a Markov decision process, with the UAV as an agent, aiming to maximize the cumulative reward related to task completion time and collision avoidance; The user scheduling based on simulated annealing, as shown in Figure 2 , specifically includes the following steps: Step S31, initialization based on the nearest greedy strategy: starting from the initial position of the UAV, the nearest unserved ground user is iteratively selected from the remaining ground user set to construct a greedy scheduling vector . This initialization provides a high-quality starting point and accelerates convergence; Step S32, scheduling optimization: taking the scheduling vector as the initial scheduling, the total service distance is set to the initial value and the minimum value at the same time, in each iteration, a candidate solution is generated by applying a neighborhood operator to the current scheduling, the objective function of iteration is defined as the total flight distance of the UAV serving all ground users, if the candidate solution can shorten the total distance, the candidate solution is accepted, otherwise, according to the Metropolis criterion, it is accepted with a certain probability, which decreases gradually with the cooling of the temperature, and the known best scheduling vector is updated whenever a shorter distance is found, wherein the neighborhood operator set is , and the Metropolis criterion is a prior art; According to the temperature-dependent probability, the operator o∈O is selected to achieve a balance between exploration and exploitation: at high temperature, it tends to large perturbation (such as subsequence reversal operation), at low temperature, it tends to small perturbation (such as insertion movement operation), the selection probability is defined as , and represent the preference coefficient and the sensitivity coefficient of the operator , respectively.

[0037] Modeling the joint optimization problem as a Markov decision process specifically includes the following steps: Step S33, constructing state space: state space including each time slot all possible states , each state defined by the locations of the UAV and the ground users, the detected buildings, and the user scheduling policy derived from step S31 and step S32, while aiming to simultaneously solve the objectives of providing communication services for all ground users and avoiding obstacle collision. defined as the combination of two parts , matrix and matrix are both matrices, m = 100 and n = 100 correspond to the grid cell indices for discretizing the task area.

[0038] The matrix facilitates the UAV to provide communication services for the ground users according to the user scheduling vector. The present application employs the Artificial Potential Field (APF) technique to integrate the user scheduling and the ground user location information. The APF method generates a virtual force field, which is a mathematical construct that simulates attractive forces to guide the UAV motion. In this force field, each ground user generates an attractive potential that pulls the UAV towards its location along the negative gradient direction. To prioritize the ground users at the front of the scheduling, the potential energy is linearly increasing with distance within a certain distance range, centered at each ground user location, ensuring that the UAV serves the ground users with higher priority first. Specifically, the potential energy at an arbitrary location is defined as , is a positive constant representing the attractive strength of the ground user k, =|| - || represents the distance between the location and , is the potential energy distance threshold. To prioritize the users at the front of the scheduling sequence, we arrange in descending order according to the user scheduling vector, ensuring that the UAV serves the ground users with higher priority first. In the matrix, the UAV is represented as a circular region with its location as the center and as the radius, corresponding to the value .

[0039] The matrix realizes the obstacle avoidance of the UAV by integrating its location and the detected building information. In the matrix, the UAV is represented as a circular region with its location as the center and as the radius, corresponding to the value The detected buildings are discretized in a matrix according to their perceived position by the drone, the discretized values corresponding to their perceived height; Step S34, building the action space: the action space contains all feasible actions for each time slot t, each action representing a flight decision for the drone, the flight decision including the speed, the vertical flight angle and the horizontal flight angle of the drone, denoted as ; Step S35, building the reward space: the reward function generates a reward at each time slot according to the state and the action , the reward being defined as , balancing the minimization of the mission duration, the collision avoidance and the exploration, where is a shaping reward function, .

[0040] is a collision penalty, if an action is predicted to cause a collision of the drone with a building or to fly out of the mission area, the action is cancelled and a collision penalty is applied, ; is a speed penalty, a speed penalty is applied when the speed of the drone is below a certain threshold, , ; is a time penalty, in order to minimize the mission duration, the drone is subject to a penalty that increases with time at each time slot, ; is a service reward, a reward is assigned for serving the ground users in order to maximize the data transmission rate and thus shorten the mission time, , ; is a distance reward, a large reward is given when the distance between the drone and the target ground user is significantly reduced, a small reward is given for a slight reduction or small increase in distance, while a penalty is applied for a large increase in distance, in order to handle obstacle avoidance, is a constant coefficient, is a distance reward threshold, is the amount of change in the distance between the drone and the target ground user at time slot t, in the present embodiment, , ; To explore the reward, to prevent the UAV from continuously hovering in the safe area due to obstacle avoidance, and to reduce the long flight time caused by repeatedly flying over the explored path, an exploration reward mechanism is introduced, , The change value of the number of grid cells newly covered by the UAV communication range in time slot t; To complete the reward, once the UAV completes the service to all users, a significant reward will be allocated, The moment of task completion, ; Step S36, constructing state transition probability: the state transition probability represents the possibility of the agent transitioning to the next state when performing action , , is expressed as .

[0041] Step S4, based on the system model, user scheduling and Markov decision process, the path optimization model is trained using the TD3 algorithm, and after the training is completed, the trained path optimization model is used for three-dimensional trajectory optimization of the UAV; The UAV follows the initial scheduling, while continuously adjusting its trajectory through the TD3 model according to real-time observations. When the UAV explores the environment, it will detect previously unknown buildings and temporarily adjust its path through TD3 to avoid collision. If this adjustment disrupts the original service order, the UAV will prioritize serving the nearest unserved ground user within its communication range. Once the obstacle is bypassed, the UAV will resume the initial scheduling and prioritize serving the unserved ground users at the top of the list. By utilizing TD3 for dynamic decision-making, the UAV has achieved a balance between safe navigation and service efficiency, thereby minimizing the task completion time in a complex, previously unknown environment.

[0042] As shown in Figure 3 , the training specifically includes the following steps: Step S41, initializing the parameters of the TD3 algorithm related to the online Actor network with parameter , two online Critic networks , , and two target Critic networks and , the online Actor network approximates the agent's policy and generates actions, the target Actor network generates a target policy, and with parameters and They estimate the action value function and select the smaller one as the action value. and The parameters are and They generate the target Q value by taking the smaller of the two calculated Q values.

[0043] In this embodiment, the Actor network maps the current state to an action. The input state matrix is ​​first processed through a series of convolutional layers with dimensions of 2x32, 32x64, 64x128, and 128x256. After flattening, the features are passed to a set of fully connected layers with dimensions of 36864x512, 512x256, and 256x128. The final branches generate different components of the action, ultimately producing the action. The Critic network structure evaluates the action taken by the Actor network. The state input is processed through convolutional layers (2x32, 32x64, 64x128, 128x256) similar to those in the Actor network.

[0044] Step S42: Based on initialization, the drone agent generates actions. The digital twin server generates a state based on the created virtual environment. and pass it on to the online Actor network. The network deterministically maps this state to an action. To balance exploration and exploitation, random noise is added to the action, which is then represented as follows: ,in, The mean is zero and the variance is Additive white Gaussian noise, For the range of action constraints, and The lower and upper limits of the action are defined respectively. This is a pruning function used to limit the sum of the Actor network and noise within the action constraints. Inside; Step S43: The drone agent interacts with the environment, receives a reward, and enters the next state. Specifically, the drone performs an action. Receive rewards Observe the next state , experience tuple Stored in the playback buffer; Step S44: Determine whether the number of empirical tuples stored in the buffer has reached a set threshold. If yes, proceed to step S45; otherwise, proceed to step S41. Step S45: After collecting a sufficient number of empirical tuples, uniformly sample a buffer of size [size missing]. The mini-batch data is used to update the network, for each experience tuple, the target policy smoothing mechanism is applied, and a smoothed target action is calculated by the target Actor network , denoted as , is an additive white Gaussian noise with mean zero and variance , and is the maximum exploration noise. Step S46, the target Q value is calculated by the clipped double Q-learning technique. Specifically, two target Critic networks evaluate the smoothed target action of the next state , respectively generating two Q values and , and the smaller one of the two values is used to calculate the target value, denoted as , is used to balance the immediate reward and the future reward. Step S47, the two online Critic networks are updated by minimizing the temporal difference error between the predicted Q value and the target Q value, obtaining the updated online Critic networks and , wherein the loss function used for updating is , and the gradient descent method is used to optimize the parameters and to minimize the respective losses, denotes the number of experience tuples uniformly sampled from the buffer. Step S48, it is judged whether the number of updates of the two online Critic networks reaches a set threshold , if yes, step S49 is entered, otherwise, step S50 is entered. Step S49, the online Actor network is updated. In order to realize delayed policy updating, the Actor network and the target Actor network are updated only after every Critic network updates.

[0045] The update target of the online Actor network is to maximize the expected Q value of its action. The deterministic policy gradient is calculated based on the Q value of the first online Critic network , denoted as , wherein represents the expected return under the policy .

[0046] Step S50, the target network parameters are updated in a soft update manner, and the target network includes the target Actor network and the target Critic network , , a soft update mechanism is adopted to ensure smooth convergence, expressed as a soft update manner , is an update rate; Step S511, judge whether the time slot index reaches T, if yes, enter step S512, otherwise, enter step S51; Step S512, end this round of training, reset the environment, and observe the initial state , judge whether the training round reaches the set threshold , if yes, complete the training, otherwise, enter step S51.

[0047] Based on the above steps, the trained path optimization model is obtained, which provides communication services for ground users while avoiding obstacles and reducing the time required to provide communication services for all users. The ISAC signal enables the UAV to actively perceive unknown obstacles, and dynamic path planning is performed in combination with user location information, solving the problem of efficient and safe communication services in the case of missing environmental information. By fusing greedy initialization and adaptive neighborhood search, a lightweight and efficient solution is provided for user scheduling in TDF scenarios. The complex path planning problem is modeled as MDP, and the UAV is abstracted as an agent. The trial-and-error mechanism of reinforcement learning enables the UAV to autonomously learn and optimize its long-term behavior strategy to maximize cumulative rewards. The advanced TD3 algorithm effectively solves the problem of overestimation of Q values in traditional DRL algorithms, and improves the stability of training and the performance of the final strategy through policy smoothing and delayed update mechanisms, so that the UAV can generate a better flight trajectory, thereby significantly shortening the total time required to complete all communication tasks while ensuring safety.

[0048] Figure 4 In the figure, the horizontal coordinate is the iteration number, and the value range is [0, 3500], and the vertical coordinate is the reward value obtained by training. Figure 4 As shown in the figure, the method of the present application and each comparison method can converge, but the convergence speed, convergence reward value and stability of the present application are better. Specifically, the introduction of the twin technology can converge faster to a higher reward value, and the curve is smoother and the variance band is narrower, showing stronger stability; in contrast, the method without introducing digital twin has larger fluctuations in the training process and slower learning progress. The fundamental reason is that digital twin can reduce environmental uncertainty through a continuously updated virtual environment, providing more reliable state representation and feedback for model training, so that the UAV can focus on learning on potential high-quality trajectories, and avoid wasting training resources in inefficient exploration.

[0049] Figure 5In the graph, the horizontal axis represents the number of ground users (GUs), with values ​​of 5, 10, 15, 20, 25, and 30 respectively, and the vertical axis represents the task completion time in seconds. Figure 5 As can be seen, the task completion time of the proposed method and all comparative methods increases with the number of users, reflecting the inherent challenges of expanding communication services in multi-user scenarios. However, the introduction of digital twins in this invention significantly improves the effectiveness and stability of the algorithm. Specifically, this invention achieves the shortest task completion time in all scenarios; although the digital twin-enhanced DDPG algorithm originally had instability and learning limitations, its performance is comparable to the TD3 algorithm without digital twin enhancement. This is mainly due to the fact that digital twins can continuously generate rich and accurate state information, effectively compensating for the algorithm's shortcomings in instability and learning ability. In addition, this invention performs better in terms of scalability; its task completion time growth curve with the number of users is relatively flat, fully demonstrating that its dynamic decision-making mechanism assisted by digital twins can efficiently support a larger scale of user access.

[0050] In this invention, the terms "first," "second," and "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. The use of terms such as "upper," "lower," "left," "right," "front," and "rear" to indicate orientation or positional relationships is based on the orientation or positional relationships shown in the accompanying drawings and is only for the convenience of describing the invention, not to indicate or imply that the device referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the scope of protection of this invention. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0051] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0052] The above are merely specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial modifications made to the present invention using this concept shall be considered as infringing upon the protection scope of the present invention.

Claims

1. A digital twin-enhanced method for optimizing the 3D trajectory of a UAV in an unknown environment, characterized by: Includes the following steps: Step S1: Construct a system model and a joint optimization problem. The system model includes a scenario of UAV path planning in an unknown environment, a time model of UAV path planning based on the scenario, a UAV flight model of UAV path planning based on the scenario and time model, and a signal and communication model between UAV, ground users and base stations based on the scenario. The joint optimization problem is constructed based on the system model with the goal of minimizing the task completion time. Step S2: Based on the system model and joint optimization problem, construct a digital twin-driven trajectory framework. This framework includes a physical entity layer, a digital twin layer, and the connection between the physical entity layer and the digital twin layer. The digital twin layer runs on a digital twin server. During mission execution, the digital twin server uses the constructed virtual environment to continuously monitor the drone's actions and assess potential safety risks. Step S3: Based on the system model and the joint optimization problem, design a user scheduling based on simulated annealing, and model the joint optimization problem as a Markov decision process. The UAV, as an intelligent agent, aims to maximize the cumulative reward related to task completion time and collision avoidance. Step S4: Based on the system model, user scheduling, and Markov decision process, train the path optimization model using the TD3 algorithm.

2. The digital twin-enhanced method for optimizing the 3D trajectory of a UAV in an unknown environment according to claim 1, characterized in that: In step S1, constructing the system model specifically includes the following steps: Step S11: Construct a scenario for UAV path planning in an unknown environment: in the mission area The ground area contains buildings and K randomly distributed ground users, with the buildings occupying an area of ​​[missing information]. The base station transmits communication data to the drone, and the drone uses ISAC signals to relay the data to ground users. The drone's communication coverage is greater than its sensing coverage. Step S12: Construct a time model for UAV path planning: The entire mission duration is divided into T time slots of equal length. The UAV makes motion decisions at the beginning of each time slot. The UAV provides services to ground users using TDMA. In each time slot, only the nearest ground user who has not yet completed the service is served. Step S13: Construct a UAV flight model for UAV path planning: In time slot t+1, the UAV's position coordinates According to the formula , , Update, in which, Let be the position coordinates of the UAV in time slot t. Let be the flight speed of the UAV in time slot t. Let be the vertical flight angle of the UAV in time slot t. Let be the horizontal flight angle of the UAV in time slot t. The duration of each time slot; Step S14: Construct a signal and communication model between the UAV, ground users, and the base station: The signal and communication model includes a base station-UAV link and a UAV-ground user link. The transmission rate of the base station-UAV link is... The transmission rate of the UAV-ground user link is The effective communication rate of the UAV transmitting data to ground user k in time slot t is: Where B represents the available spectrum bandwidth, This indicates the constant transmit power of the base station. This represents the power of zero-mean additive white Gaussian noise. , This indicates large-scale fading of the base station-drone link in time slot t. This indicates the small-scale fading of the base station-drone link in time slot t. Indicates the drone's transmission power. , This indicates large-scale fading of the UAV-ground user link in time slot t. This indicates the small-scale fading of the UAV-Ground User link in time slot t.

3. The digital twin-enhanced method for optimizing the 3D trajectory of a UAV in an unknown environment according to claim 2, characterized in that: In step S1, the joint optimization function is expressed as follows: ,in, The set of speeds of the drone throughout the entire mission duration. , The maximum speed of the drone, This refers to the set of vertical flight angles of the drone throughout the entire mission duration. This refers to the set of horizontal flight angles of the drone throughout the entire mission duration. The minimum number of time slots required to serve ground user k for a drone. , The amount of data transmitted by the UAV to ground user k in time slot t. For the throughput requirements of ground user k, For ground users.

4. The digital twin-enhanced method for optimizing the three-dimensional trajectory of a UAV in an unknown environment according to claim 3, characterized in that: In step S2, the physical entity layer of the digital twin driven trajectory framework includes a physical drone, a mission execution area, and ground users waiting for service. The drone, as an intelligent agent, is responsible for observing the environment, executing control actions received from the base station, providing communication services to ground users, and uploading the perception data to the digital twin server. The virtual environment constructed by the digital twin server captures the mobility of the drone, the distribution of ground users, and environmental features including buildings. Based on the virtual environment, the three-dimensional trajectory planning process is carried out in the digital twin layer. The physical entity layer and the digital twin layer are connected through a bidirectional link to ensure synchronization and coordination. During task execution, before each action is transmitted to the drone, the twin server checks whether the action will collide with obstacles or cross the boundary. If so, the action is canceled and the drone is instructed to remain stationary in that time slot.

5. The digital twin-enhanced method for optimizing the three-dimensional trajectory of a UAV in an unknown environment according to claim 4, characterized in that: In step S3, designing user scheduling based on simulated fallback specifically includes the following steps: Step S31: Initialization based on the nearest greedy strategy: Starting from the initial position of the drone, iteratively initialize from the remaining set of ground users. Select the nearest unserved ground user to construct a greedy scheduling vector. ; Step S32, Scheduling Optimization: Adjust the scheduling vector As the initial scheduler, in each iteration, a candidate solution is generated by perturbing the current scheduler with a neighborhood operator. The objective function of the iteration is defined as the total flight distance of the UAV serving all ground users. If the candidate solution can shorten the total distance, it is accepted; otherwise, it is accepted with a certain probability according to the Metropolis criterion, which decreases as the temperature gradually cools. Whenever a shorter distance is found, the known optimal scheduler vector is updated. The set of neighborhood operators is as follows: .

6. The digital twin-enhanced method for optimizing the three-dimensional trajectory of a UAV in an unknown environment according to claim 5, characterized in that: In step S3, modeling the joint optimization problem as a Markov decision process specifically includes the following steps: Step S33: Construct the state space: State space Including each time slot All possible states , will state Defined as a combination of two parts , Matrix and Matrix The matrices, m and n, correspond to the grid cell indices used to discretize the task region. The matrix facilitates the provision of communication services from UAVs to ground users based on user scheduling vectors. In the matrix, the drone is represented as a circle centered at its position. A circular region with radius , corresponding to the value . Based on the positive constant of the attraction intensity of the user scheduling vector to ground user k Sort in descending order. The matrix integrates the drone's location and detected building information to enable obstacle avoidance for the drone. In the matrix, the drone is represented by a circle with its position center, A circular region with radius , corresponding to the value . The detected buildings are placed in the matrix based on their location as perceived by the drone. The height is discretized, and the discretized value corresponds to its perceived height. Step S34: Construct the action space: Action space Includes all feasible actions for each time slot t. Each action represents a flight decision for the drone, which includes the drone's speed, vertical flight angle, and horizontal flight angle, represented as... ; Step S35: Construct the reward space: The reward is defined as... ,in, , , As a penalty for collision, , Penalty for speed , , As a time penalty, , As a service reward, , , As a distance reward, , The distance reward threshold, Let be the change in distance between the UAV and the target ground user at time slot t. To explore rewards, , The change in the number of newly covered grid cells within the UAV communication range during time slot t. To complete the reward, , The moment the task is completed; Step S36: Construct state transition probabilities: State transition probabilities represent the probability of an agent transitioning from state S36 to state S36. Execute action When, transition to the next state. The possibility is expressed as .

7. The digital twin-enhanced three-dimensional trajectory optimization method for unmanned aerial vehicles in unknown environments according to claim 6, characterized in that: Step S4 specifically includes the following steps: Step S41: Initialize the online Actor network for the TD3 algorithm. Target Actor Network Two online Critic networks , and two target Critic networks and ; Step S42: Based on initialization, the digital twin server generates a state according to the created virtual environment. and pass it on to the online Actor network. The network deterministically maps this state to an action. To balance exploration and exploitation, random noise is added to the action, which is then represented as follows: ,in, The mean is zero and the variance is Additive white Gaussian noise, For the range of action constraints, , , This is the clipping function; Step S43: The drone agent performs actions. Receive rewards Observe the next state , experience tuple Stored in the playback buffer; Step S44: Determine whether the number of experience tuples stored in the buffer has reached the set threshold. If yes, proceed to step S45; otherwise, proceed to step S41. Step S45: For each empirical tuple, compute a smooth target action using the target Actor network. , represented as , It is additive white Gaussian noise with zero mean. To maximize noise exploration; Step S46: Calculate the target Q value , ; Step S47: The two online Critic networks are updated by minimizing the time difference error between the predicted Q-value and the target Q-value, resulting in the updated online Critic network. and The loss function used for the update is... , This represents the number of empirical tuples uniformly sampled from the buffer; Step S48: Determine whether the number of updates of the two online Critic networks has reached the set threshold. If yes, proceed to step S49; otherwise, proceed to step S50. Step S49: Update the online Actor network with the goal of maximizing the expected Q value of its action; Step S50: Soft update method Update the target Actor network and the target Critic network. For update rate; Step S511: Determine whether the time slot index has reached T. If yes, proceed to step S512; otherwise, proceed to step S51. Step S512: End this round of training, reset the environment, and observe the initial state. Determine whether the training rounds have reached the set threshold. If so, complete the training; otherwise, proceed to step S51.

Citation Information

Cited By

  • Passive network anomaly monitoring and unmanned aerial vehicle automatic inspection method and system

    CN122090327A