A decision-making method for multi-UAV networks based on imperfect digital twins

By using imperfect digital twin systems and hybrid learning methods, the flight trajectory of drone swarms is optimized, solving the problems of training cost and energy consumption in multi-drone network communication systems, and achieving efficient and high-quality drone communication.

CN116828507BActive Publication Date: 2026-04-17XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2023-07-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies in multi-drone network communication systems have failed to effectively balance the high efficiency and quality of drone communication with training costs and flight energy consumption costs. Furthermore, digital twin systems assume a perfect replication of the physical environment, ignoring imperfections and construction costs.

Method used

By employing an imperfect digital twin system, and through pre-trained environment decision networks and action decision networks, combined with unlabeled unsupervised learning and reinforcement learning, the flight trajectory of the drone swarm is optimized, training costs are reduced, and noise bias of virtual drones is taken into account, thereby optimizing the number and flight direction of real drones.

Benefits of technology

It achieves both high-efficiency and high-quality UAV communication, while reducing the training cost and flight energy consumption cost of UAV swarms and optimizing the average transmission rate of UAV networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116828507B_ABST
    Figure CN116828507B_ABST
Patent Text Reader

Abstract

This invention discloses a decision-making method for multi-UAV networks based on imperfect digital twins, comprising: obtaining the first position of each UAV in a UAV swarm with specific environmental parameters at the current moment, and the second position of each user device; the specific environmental parameters characterize the number of real UAVs in the UAV swarm, and the noise deviation between real UAVs and virtual UAVs; the specific environmental parameters are determined by a pre-trained environmental decision network according to preset parameters; the pre-trained environmental decision network is trained on the initial environmental decision network through unlabeled unsupervised learning based on multiple sets of sample data and a first pre-trained action decision network; the first and second positions are input into a second pre-trained action decision network to obtain the flight direction of each UAV at the current moment; the second pre-trained action decision network is trained on the initial action decision network through reinforcement learning using samples obtained under specific environmental parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, specifically relating to a decision-making method for multi-UAV networks based on imperfect digital twins. Background Technology

[0002] Drones, as low-cost and highly flexible mobile network access nodes, are widely used to improve network performance. Since drones are mobile, and communication rates in mobile networks are negatively correlated with the distance between network users and access nodes, optimizing drone flight trajectories to reduce the distance between all users and drones is crucial for optimizing average network performance over time. Traditionally, this type of temporal optimization can be solved using dynamic programming. However, as the number of drones in the network increases, the dimensionality of the state space in dynamic programming also grows dramatically. Therefore, when multiple drones exist in the network, searching the entire state space using dynamic programming to find the optimal trajectory is very time-consuming. Fortunately, reinforcement learning methods based on deep neural networks can effectively optimize drone trajectories with only a few matrix multiplications. Therefore, much work now focuses on using deep reinforcement learning to solve the optimization problem of drones in wireless networks. Meanwhile, the rapid development of digital twin technology provides a reliable digital simulation platform for the rapid deployment of various algorithms, enabling the digital reproduction of physical environments. However, research on optimizing the training costs of neural networks using deep reinforcement learning, especially the hardware costs of purchasing drones and the energy consumption costs incurred by drone flight during training, remains lacking. Meanwhile, current research largely assumes that digital twin technology can perfectly reproduce all characteristics of the physical environment. However, current research has not yet addressed issues such as how imperfect digital twin systems can assist in system optimization and the construction cost of digital twins.

[0003] In existing multi-UAV network communication systems, the following are the main methods for optimizing the movement trajectory of UAV swarms.

[0004] The first approach uses a dynamic programming-based method. This method employs a recursive search, comparing the communication performance gains of all possible flight paths to determine which flight path the drone should use. However, the recursive search latency of this method increases exponentially with the number of drones and their range of motion, making it difficult to apply to real-time path planning in multi-drone networks.

[0005] The second approach utilizes deep reinforcement learning to optimize the flight methods and trajectories of drone swarms. This method achieves real-time trajectory optimization through lightweight neural network design. Furthermore, due to the powerful representational capabilities of neural networks, it can analyze and extract effective features from the environment, thereby improving the communication performance of the drone network. However, deep reinforcement learning algorithms require extensive neural network training before use. To ensure the performance of the neural network, a large amount of data must be used for training. However, in drone communication systems, collecting training data relies on the drone swarm extensively trying various flight paths, leading to extremely high training costs.

[0006] The third approach utilizes digital twin systems to simulate and analyze the impact of different flight trajectories on network performance in digital space, thereby selecting the optimal flight path for deployment in the physical system. However, these algorithms typically do not consider the construction cost of the digital twin system and assume that the digital space in the digital twin can perfectly replicate all characteristics of the physical space, which is difficult to achieve both economically and technically.

[0007] As mentioned above, current methods for optimizing UAV communication networks using reinforcement learning rarely address the optimization of the training costs of reinforcement learning itself, particularly the purchase cost of the UAV, the energy consumption costs during UAV training, and the cost of transmitting data between the UAV and the environment to the controller. In other words, there is currently no method that can simultaneously achieve efficient and high-quality UAV communication while also minimizing UAV training costs. Summary of the Invention

[0008] To address the aforementioned problems in related technologies, this invention provides a decision-making method for multi-UAV networks based on imperfect digital twins. The technical problem to be solved by this invention is achieved through the following technical solution:

[0009] This invention provides a decision-making method for multi-UAV networks based on imperfect digital twins, comprising:

[0010] The system obtains the first position of each drone in a drone swarm with specific environmental parameters at the current moment, and the second position of each user device among the multiple user devices served by the drone swarm; the specific environmental parameters represent the number K of real drones in the drone swarm, and the noise deviation δ between real drones and virtual drones; the specific environmental parameters are determined by a pre-trained environmental decision network based on preset environmental index parameters; the pre-trained environmental decision network is obtained by training an initial environmental decision network through unlabeled unsupervised learning based on multiple sets of sample data and a first pre-trained action decision network;

[0011] The first position and the second position are processed by a second pre-trained action decision network to obtain the flight direction of each UAV at the current time. The second pre-trained action decision network is obtained by training the initial action decision network with reinforcement learning using sample data obtained under the specific environmental parameters.

[0012] In some embodiments, the drone swarm includes multiple real drones and multiple virtual drones; the preset environmental indicator parameters include: parameter η for characterizing the importance of performance indicators, parameter α for characterizing the cost of virtual drones in the drone swarm, parameter ζ for characterizing the cost of real drones in the drone swarm, and parameter β for characterizing the construction difficulty of virtual drones in the drone swarm.

[0013] In some embodiments, the gradient update function of the environment decision network, the gradient update function of the action decision network, and the reward function of the action decision network are all determined according to an objective function; the objective function is used to reduce training costs while maximizing the transmission rate of all user devices in the average time.

[0014] In some embodiments, the expression of the objective function is as follows:

[0015]

[0016] in, Let i be the position of user equipment i at time t. Let j be the position of the drone at time t. Let g be the position of UAV j at time t-1, g0 be the path loss constant, and L0 be the reference distance. Let p be the distance between user device i and drone j at time t. i Let σ be the transmission power of user equipment i. 2 Let M be the noise power, N be the total number of drones in the drone swarm, η be the total number of user devices, α be the cost of virtual drones in the drone swarm, ζ be the cost of real drones in the drone swarm, β be the construction difficulty of virtual drones in the drone swarm, K be the number of real drones in the drone swarm, δ be the noise deviation between real and virtual drones in the drone swarm, log2(.) be the log function, and v be the noise power. t Let loc be the velocity of the drone at time t. user For the location of the user equipment, loc uav,t-1 Let |t| be the position of the UAV at time t-1, and ||.|| be the absolute value function. The action decision network is described above, where θ is the path loss factor; constraint (4a) represents the distance between user equipment i and UAV j; constraint (4b) represents that the flight speed of each UAV in the UAV swarm is a constant; in constraint (4c), the action decision network... This is used to determine the flight direction of the UAV at time t, where the noise bias δ ​​affects the loss function during neural network training. The parameters affect The performance of the drone; constraint (4d) constrains the update of the drone's position; constraint (4e) constrains the number of real drones in the drone swarm to be at most M and at least 0; constraint (4f) constrains δ to be greater than or equal to 0.

[0017] In some embodiments, the gradient update function of the environmental decision network is expressed as follows: in, Let the objective function be... For the environmental decision network Network parameters, For the environmental decision network The learning rate For about Partial derivatives of the parameters, Let be the partial derivatives with respect to δ and K. This refers to the action decision network.

[0018] In some embodiments, the gradient update function of the action decision network is expressed as follows: in, For the action decision network Network parameters, Let r be the learning rate of the action decision network Q. t Let v be the reward value of the drone at time t. t+1 Let s be the speed of the drone at time t+1. t+1 Let v be the position of the drone at time t+1. t Let s be the velocity of the drone at time t. t Let t be the position of the drone at time t. For about Partial derivatives of the parameters, for The target network, In order to make v when taking the maximum value t+1 The value of .

[0019] In some embodiments, before obtaining the first position of each drone in a drone swarm with specific environmental parameters at the current moment, and the second position of each user device among the multiple user devices served by the drone swarm, the method further includes:

[0020] Obtain multiple sets of sample data corresponding to multiple time points; each set of sample data includes: the position of the drone at time t, the flight direction of the drone at time t, and the reward value of the drone at time t; the reward value of the drone at time t represents the reward value obtained by the drone at the position at time t and executing the flight direction at time t; t is an integer greater than 0;

[0021] The output of the initial action decision network is used as the input of the initial environment decision network, and the output of the initial environment decision network is also used as the input of the initial action decision network to obtain the initial serial network;

[0022] Using the multiple sets of sample data, the initial action decision network in the initial concatenated network is iteratively trained through reinforcement learning to obtain a first pre-trained concatenated network that includes the first pre-trained action decision network and the initial environment decision network.

[0023] Using the multiple sets of sample data, the initial environment decision network in the first pre-trained concatenated network is iteratively trained through unlabeled unsupervised learning to obtain the pre-trained environment decision network.

[0024] In some embodiments, the expression for the drone's reward value at time t is as follows:

[0025]

[0026] Where M is the total number of drones in the drone swarm, N is the total number of user devices, g0 is the path loss constant, and L0 is the reference distance. Let p be the distance between user device i and drone j at time t. i Let σ be the transmission power of user equipment i. 2 Let θ be the noise power and θ be the path loss factor. It is a Gaussian random variable with a mean of 0 and a variance of Kδ.

[0027] In some embodiments, before processing the first position and the second position using a second pre-trained action decision network to obtain the flight direction of each UAV at the current moment, the method further includes:

[0028] Under the specific environmental parameters, multiple sets of sample data corresponding to multiple times are obtained; each set of sample data under the specific environmental parameters includes: the position of the UAV at time c, the flight direction of the UAV at time c, and the reward value of the UAV at time c; the reward value of the UAV at time c represents the reward value obtained by the UAV at the position at time c and the flight direction at time c; c is an integer greater than 0.

[0029] Using the multiple sets of sample data under the specific environmental parameters, the initial action decision network is iteratively trained through reinforcement learning to obtain the second pre-trained action decision network.

[0030] In some embodiments, the network structures of the initial environment decision network and the action decision network are the same as those of the MLP network.

[0031] The present invention has the following beneficial technical effects:

[0032] This invention addresses the common problem of simultaneously reducing training costs and improving the average transmission rate of the drone network over time. By establishing a noise deviation δ between the physical and digital spaces in a digital twin system (i.e., the noise deviation δ between real and virtual drones in a drone swarm), and training two neural networks using low-cost, label-free, unsupervised learning and reinforcement learning, this invention performs hybrid optimization on scalar variables K (i.e., the number of real drones in the drone swarm) and δ, as well as time-series variable v (i.e., the flight direction of the drone at each moment). This achieves the goal of simultaneously ensuring high network transmission rates for drones while reducing the deep reinforcement learning training costs and flight energy costs caused by the presence of real drones in the drone swarm.

[0033] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0034] Figure 1 A flowchart illustrating a decision-making method for multi-UAV networks based on imperfect digital twins, provided in an embodiment of the present invention.

[0035] Figure 2 This is an exemplary schematic diagram illustrating the process of synthesizing acquired images provided in an embodiment of the present invention;

[0036] Figure 3 A schematic diagram of an exemplary upsampling process provided for an embodiment of the present invention;

[0037] Figure 4 A schematic diagram illustrating the principle of improving a digital twin model using real-time data from the physical space, as provided in an embodiment of the present invention.

[0038] Figure 5 A schematic diagram illustrating the principle of an exemplary training agent provided in an embodiment of the present invention;

[0039] Figure 6 A schematic diagram comparing the convergence performance of an exemplary physical drone and a different number of digital twin systems provided in this embodiment of the invention;

[0040] Figure 7 A schematic diagram comparing the convergence performance of an exemplary physical drone and digital twin system under different noise levels, provided for embodiments of the present invention;

[0041] Figure 8 A comparative diagram illustrating the network utility of an exemplary physical drone and different digital twin systems provided for embodiments of the present invention;

[0042] Figure 9 An exemplary diagram illustrating the relationship between network benefits and the construction cost alpha of a digital twin system is provided for embodiments of the present invention.

[0043] Figure 10 This is a schematic diagram illustrating the relationship between exemplary network benefits and the construction cost beta of a digital twin system, provided for embodiments of the present invention. Detailed Implementation

[0044] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0045] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0046] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0047] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0048] The applicant has discovered that current methods for optimizing UAV networks with digital twin assistance all assume that the digital twin can perfectly replicate all features of the physical space, ignoring the errors between the simulated digital space and the real physical space caused by the limitations of real-world technology, and failing to consider the construction costs inherent in the digital twin system itself. Based on this discovery, the applicant proposes a decision-making method for multi-UAV networks based on imperfect digital twins. This method can solve the problem of low-cost optimization of UAV swarm flight paths in UAV network communication systems, reduce the training cost of deep reinforcement learning optimization methods, and comprehensively consider economic and technical characteristics to address how to leverage imperfect digital twin systems to achieve high-efficiency, high-quality performance improvements in UAV communication systems, while also comprehensively optimizing the construction costs of digital twin systems.

[0049] Figure 1 This is a flowchart of a decision-making method for multi-UAV networks based on imperfect digital twins provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes the following steps:

[0050] S101. Obtain the first position of each drone in the drone swarm with specific environmental parameters at the current moment, and the second position of each user device among the multiple user devices served by the drone swarm; the specific environmental parameters represent the number K of real drones in the drone swarm, and the noise deviation δ between real drones and virtual drones; the specific environmental parameters are determined by a pre-trained environmental decision network based on preset environmental index parameters; the pre-trained environmental decision network is obtained by training the initial environmental decision network through unlabeled unsupervised learning based on multiple sets of sample data and the first pre-trained action decision network.

[0051] Here, the drone swarm consists of multiple real drones and multiple virtual drones. Preset environmental parameters include: η, which characterizes the importance of performance indicators; α, which characterizes the cost of virtual drones in the swarm; ζ, which characterizes the cost of real drones in the swarm; and β, which characterizes the construction difficulty of virtual drones in the swarm. The values ​​of η, ζ, α, and β can be preset according to actual needs. Different settings for η, ζ, α, and β will result in different generated environmental parameters (i.e., K and δ).

[0052] S102. The second pre-trained action decision network is used to process the first position and the second position to obtain the flight direction of each UAV at the current time. The second pre-trained action decision network is obtained by training the initial action decision network with sample data obtained under specific environmental parameters through reinforcement learning.

[0053] Here, the network structures of the initial environment decision network and the action decision network are the same as those of the MLP network.

[0054] For example, in a drone swarm, all drones have the same speed. By making decisions about the flight direction of each drone at a given time, the flight trajectory of each drone in the swarm can be determined, thereby providing better network services to user equipment.

[0055] Here, the environmental decision-making network gradient update function Action Decision Network gradient update function The reward function r of both the action decision network and the action decision network are determined based on the objective function, which is used to reduce training costs while maximizing the transmission rate of all user devices in the average time.

[0056] For example, the expression for the objective function is as follows:

[0057]

[0058] in, Let i be the position of user equipment i at time t. Let j be the position of the drone at time t. Let g be the position of UAV j at time t-1, g0 be the path loss constant, and L0 be the reference distance. Let p be the distance between user device i and drone j at time t. i Let σ be the transmission power of user equipment i. 2Let M be the noise power, N be the total number of drones in the drone swarm, η be the total number of user devices, α be the cost of virtual drones in the drone swarm, ζ be the cost of real drones in the drone swarm, β be the construction difficulty of virtual drones in the drone swarm, K be the number of real drones in the drone swarm, δ be the noise deviation between real and virtual drones in the drone swarm, log2(.) be the log function, and v be the noise power. t Let loc be the velocity of the drone at time t. user For the location of the user equipment, loc uav,t-1 Let |t| be the position of the UAV at time t-1, and ||.|| be the absolute value function. It is an action decision network, where θ is the path loss factor; constraint (4a) represents the distance between user equipment i and UAV j; constraint (4b) represents that the flight speed of each UAV in the UAV swarm is a constant; in constraint (4c), the action decision network... This is used to determine the flight direction of the UAV at time t, where the noise bias δ ​​affects the loss function during neural network training. The parameters thus affect The performance of the drone; constraint (4d) constrains the update of the drone's position; constraint (4e) constrains the number of real drones in the drone swarm to be at most M and at least 0; constraint (4f) constrains δ to be greater than or equal to 0.

[0059] For example, The expression is as follows: in, Let be the objective function. For environmental decision-making networks Network parameters, For environmental decision-making networks The learning rate For about Partial derivatives of the parameters, Let be the partial derivatives with respect to δ and K. For action decision networks.

[0060] For example, θ Q The expression is as follows: in, For action decision networks Network parameters, For action decision networks The learning rate, r y Let v be the reward value of the drone at time t. t+1 Let s be the speed of the drone at time t+1.t+1 Let v be the position of the drone at time t+1. t Let s be the velocity of the drone at time t. t Let t be the position of the drone at time t. For about Partial derivatives of the parameters, for The target network, In order to make v when taking the maximum value t+1 The value of .

[0061] For example, the expression for the drone's reward value at time t, calculated based on the reward function r, is as follows: Where M is the total number of drones in the drone swarm, N is the total number of user devices, g0 is the path loss constant, and L0 is the reference distance. Let p be the distance between user device i and drone j at time t. i Let σ be the transmission power of user equipment i. 2 Let θ be the noise power and θ be the path loss factor. It is a Gaussian random variable with a mean of 0 and a variance of Kδ.

[0062] In some embodiments, before S101, steps S201 to S204 are included, through which a pre-trained environmental decision network can be obtained.

[0063] S201. Obtain multiple sets of sample data corresponding to multiple time points; one set of sample data includes: the position of the UAV at time t. t The flight direction of the UAV at time t is a t The reward value r of the drone at time t t The reward value of the drone at time t represents the reward value obtained by the drone at its position at time t and executing the flight direction at time t; t is an integer greater than 0.

[0064] Here, multiple sets of sample data correspond to multiple time points, and each set of sample data can be denoted as... t a t r t For example, the flight direction a of each drone at time t. t It can be one of the directions {00, 01, 10, 11}, where "00" represents a direction of 0°, "01" represents a direction of 90°, "10" represents a direction of 180°, and "11" represents a direction of 270°. This reduces the action space and thus reduces complexity.

[0065] ​In some embodiments, the reward value also has a discount factor γ, which represents the degree of importance attached to future rewards.

[0066] S202, Initial Action Decision Network The output is used as the initial environment decision network. The input, and simultaneously, the initial environmental decision network. The output is used as the initial action decision network. The input yields the initial cascaded network (i.e., the cascaded network). With the network ).

[0067] S203. Using multiple sets of sample data, reinforcement learning is applied to the initial action decision network in the initial concatenated network. Perform iterative training to obtain the action decision network containing the first pre-trained network. and initial environment decision network The first pre-trained concatenated network.

[0068] S204. Using multiple sets of sample data, the initial environment decision network in the first pre-trained concatenated network is trained through unlabeled unsupervised learning. Iterative training is performed to obtain a pre-trained environment decision network.

[0069] Here, since K and δ are both scalar variables and are constant over time, modeling the optimization of K and δ as a Markov decision process (MDP) is quite challenging. Furthermore, obtaining the optimal values ​​of K and δ through supervised learning as training... Labeling is very time-consuming, but since the objective function mentioned above is differentiable, it can be used in neural networks. Update guide parameters The gradient can be obtained using the chain rule. Specifically, we can first calculate the derivative of the objective function with respect to K and δ. Therefore, by expressing the objective function as... Update using gradient descent and through The above formula is used to perform neural network operations. The network parameters are updated to achieve [the desired effect]. Training.

[0070] Here, the neural network is controlled by gradients. and Coupled together, The output v (flight direction) affects The parameters are updated to reflect the gradient. The higher the trajectory optimization performance, The more accurate the parameter gradient, the better. This will reduce training costs and To achieve a better trade-off between training performance and [other factors]. Meanwhile, The trade-off in performance will, in turn, affect The gradient calculation accuracy. This training method enables... The upper layer learns the joint optimization of v, K, and δ, and makes the optimization of v by... The decision reduced the overall difficulty of training.

[0071] In some embodiments, during training After the training, K and δ can be obtained, and K and δ can be used as hyperparameters. Sample data can be constructed using some UAV flight data obtained under hyperparameters K and δ, and the network can be trained using deep reinforcement learning. Specifically, training steps S301 to S302.

[0072] S301. Under specific environmental parameters, acquire multiple sets of sample data corresponding to multiple times. A set of sample data under specific environmental parameters includes: the position of the UAV at time c, the flight direction of the UAV at time c, and the reward value of the UAV at time c. The reward value of the UAV at time c represents the reward value obtained by the UAV at the position at time c and the flight direction at time c. c is an integer greater than 0.

[0073] Here, a drone swarm is constructed based on hyperparameters K and δ, and multiple sets of sample data corresponding one-to-one with multiple time points are generated based on the flight data of the drones in the swarm. Similarly, each set of sample data can be denoted as... c a c r c For example, the flight direction a of each drone at time c. c It can be one of the directions {00, 01, 10, 11}, where "00" represents a direction of 0°, "01" represents a direction of 90°, "10" represents a direction of 180°, and "11" represents a direction of 270°. This reduces the action space and thus reduces complexity.

[0074] S302. Using multiple sets of sample data under these specific environmental parameters, reinforcement learning is used to refine the initial action decision network. Iterative training is performed to obtain the second pre-trained action decision network.

[0075] ​Here, since the trajectory control of the UAV is a continuous decision-making process, the optimal direction in each time slot is determined by the current network state. Therefore, this process can be modeled as a Markov decision process. In this invention, the agent controls the flight direction of all UAVs in the network (UAV swarm). At the beginning of each time slot, the agent provides the flight direction of each UAV based on the positions of the user equipment and the UAV. Then, the agent improves its own flight direction decision strategy based on the reward value obtained after the UAV executes the flight direction. In this way, the system achieves... Training.

[0076] Here, reinforcement learning methods are used for training. At that time, the target network can be introduced. To aid training, and in updating When using the above network parameters The calculation formula is updated with parameters.

[0077] In some embodiments, the objective function is constructed through the following process:

[0078] 1. Theoretical basis for model establishment:

[0079] The establishment of digital twin models is based on the real world; ideally, a digital twin model is a complete replica of the real world. However, considering economic costs and the limitations of computing power, this invention uses imperfect digital twin models for training reinforcement learning models. The main research task in imperfect digital twin models is the optimization of drone flight trajectories.

[0080] 1) Objects contained in the model

[0081] V = {v1, v2, v3, ..., v} M Let V be the set of all unmanned aerial vehicles (UAVs) in the digital space, where v i (1≤i≤M) represents the i-th drone. U={u1,u2,u3,…,u…} N}, where set U is the total set of user equipment in the digital space, and u j (1≤j≤N) represents the j-th user.

[0082] 2) The internal logic of the model

[0083] This invention uses OFDM to calculate the UAV v in digital space i and user equipment u j The transmission rate R between ij Considering that the channel noise is a Gaussian random variable, the expression for the transmission rate can be obtained as follows: This formula is an important basis for establishing imperfect digital twin models.

[0084] 3) Optimization objective of the model

[0085] Let the mapping relationship f represent the optimization objective of the imperfect digital twin model, namely, the fastest transmission rate: In this formula, for any user distribution V0, there is a definite UAV array distribution U0 that corresponds to it, maximizing the transmission rate. Specifically, there are three optimization variables: the UAV flight direction a... i The number of drones M deployed in the digital space, and the noise variance δ in the digital space; the matrix A = [a1, a2, a3, ..., a...] for the drone flight directions. M It can be used to control the flight trajectory of drones.

[0086] 2. Digitalization of physical space

[0087] To make the digital twin model as close to the real world as possible, its creation must be supported by real data. Since this invention uses an imperfect digital twin model to study the optimization of UAV flight trajectories, only a subset of key data needs to be collected. Specifically, digitization mainly includes three aspects: data perception, integration of multi-source heterogeneous data, and data transmission.

[0088] 1) Data perception

[0089] The creation of a digital twin model requires a large amount of accurate data, and data perception determines the final effect of the digital twin model. In this invention, we not only collected common types of data from drones, such as altitude, speed, and orientation, but also collected high-altitude image data taken by drones and used machine learning algorithms to synthesize images to obtain a complete view of a certain area. This helps to make imperfect digital twin models more intuitive.

[0090] In this invention, we define a data matrix D. For the i-th UAV, the data collected during a particular flight can be represented by the data matrix D. i express: This formula gives the data matrix D. i The basic structure and sample data. Data matrix D i The first line t n This is a time series; lines two through four (x n y n h n The first row represents the spatial coordinates of the UAV, corresponding to the time series; the fifth row represents the UAV speed, also corresponding to the time series. It should be noted that an outlier value of -1.0 appeared for the UAV speed in the sample data, which is due to the low sampling rate of the speed sensor used in this invention. Furthermore, in this invention, we also use machine learning algorithms to synthesize images acquired by the UAV to construct a comprehensive view of a specific area. For example, as shown... Figure 2 As shown in Figure 2 Among them, the left N images ∑P′ i are all original images collected by drones. After image synthesis, the right image P′0 is obtained. The reason for image synthesis is that the flight altitude of a single drone is limited, so it is impossible to directly photograph the全貌 of the area.

[0091] 2) Multi-source heterogeneous data integration

[0092] Data matrix D i The basic structure and sample data are as follows: Since the sampling rate of the speed sensor used in this invention is low, therefore, the method of "compensating values" is used to simulate the time series. Therefore, this invention uses a multi-source heterogeneous data integration method for data processing. Specifically, this invention upsamples the drone speed sequence values, and the upsampling process is as Figure 3 shown, where y(nT1) is the original digital signal, the sampling period is T1, x(nT2) is the upsampled signal, and the sampling period becomes T2 (T2 < T1), a(nT2) is the output signal obtained after the upsampled signal y(nT2) passes through a low-pass filter, and the sampling period is still T2. In this process, the cut-off frequency ω of the low-pass filter c = 1 / 2KT2, and its function is to prevent the frequency domain mirror generated by direct interpolation. The sample data D i ′ after upsampling is: At this point, the data processing is basically completed. Next, a model is established based on the data matrix D′ and the regional image P′0 matrix.

[0093] 3) Data transmission

[0094] Since this invention uses drones in the physical space and the digital space as the training set together, the data transmission ability determines the upper limit of the model's performance. The fifth-generation mobile communication network technology 5G has the characteristics of high bandwidth, high security, and low latency, and can be used for data transmission between the physical space and the digital space. Based on the above requirements, this invention uses the Y10-5G drone communication module, with a downlink rate of 2Gbps, an uplink rate of 230Mbps, and a maximum latency of 400ms, which can efficiently complete the goal of hybrid deployment of drones in the physical space and the digital space.

[0095] 3. Modeling method

[0096] The construction of the digital twin core model is the key link to ensure the stable operation of the entire digital twin model. For the construction of a non-perfect digital twin model, it mainly includes multi-model fusion and optimization of the non-perfect digital twin model.

[0097] 1) Multi-model fusion

[0098] This invention employs a multi-model fusion strategy when establishing an imperfect digital twin model. It not only performs data modeling to meet the vast majority of application scenarios but also performs knowledge modeling and mechanism modeling to handle potential extreme cases. Data modeling involves building a digital twin model based on data from the physical space. Its data foundation consists of an upsampled data matrix D′ and a region image P′0 obtained through image synthesis. The optimization objective is to achieve the fastest transmission rate. Knowledge modeling and mechanism modeling take into account extreme scenarios that may occur in practical applications. This invention considers one of the most likely extreme scenarios: the user distribution in a certain region is unknown. For this scenario, this invention assumes that the user distribution in this region follows a two-dimensional uniform distribution: V′~U D (D); The optimization goal remains the fastest transmission rate:

[0099] 2) Optimization of imperfect digital twin models

[0100] In practical applications, due to economic costs and limitations in computing power, digital twin models cannot perfectly replicate the real world. To address this, this invention employs a feedback optimization method, using real-time data from the physical space to improve the digital twin model. Its principle is as follows: Figure 4 ,refer to Figure 4 To optimize the imperfect digital twin model, we simultaneously run the physical and digital spaces at time t, obtaining their outputs h1(t) and h2(t) respectively. Then, we feed these outputs into a comparator and use the comparator's feedback to adjust the parameters of the digital space, thus completing the feedback optimization. In this invention, the similarity between the digital space and the physical space, adjusted using the feedback optimization method, exceeds 90%. At this point, the imperfect digital twin model is complete.

[0101] 4. Establish a communication transmission model

[0102] This patent considers a model with K users located in remote villages or areas where infrastructure is absent or damaged. Due to the effectiveness and flexibility of drones, M drones are used as mobile access points to provide communication services to the users, maximizing the average communication rate for any user. Inspired by the successful application of Orthogonal Frequency Division Multiplexing (OFDM) in 4G / 5G networks, this paper employs OFDM to eliminate interference between users and drones. Therefore, the transmission rate between user i and drone j can be calculated as: Where B is the bandwidth, g0 is the path loss constant, L0 is the reference distance, and L ij Let i be the distance between user i and drone j. Position of user i Let be the position of UAV j, p be the transmission power, and σ be the position of UAV j.2 Let N(0, δ) be the noise power, and let N(0, δ) be a Gaussian random variable with mean 0 and variance δ. Assume all K users are randomly distributed on a finite plane, their positions remain constant, and all drones initially reside in the same location, their release point considered as the drone hangar. Therefore, to optimize drone positions to improve user transmission rates, flight trajectories need to be optimized to increase the average time transmission rate. Drone positions are optimized by changing flight directions, while the flight speed of each drone remains constant. In this patent, the consideration is that the distance between any user and a drone is not very far, thus allowing one user to establish communication with multiple drones. Due to the use of OFDM and the neglect of interference, the communication rate for a single user is the maximum rate between that user and all drones.

[0103] 5. Establish a digital twin model

[0104] To effectively optimize flight trajectories, this invention employs an imperfect digital twin system, which has three main functions: physical data reproduction, simulation generation, and decision control. Considering economic costs, the more powerful the digital twin's functions, the higher its construction cost. First, the imperfect digital twin system can collect data from physical space and reproduce the characteristics of the physical entity in digital space using certain technical means. Since the characteristics of the digital copy in the digital twin system must always remain consistent with the physical entity, it is necessary to frequently transmit entity information from physical space to the digital twin to ensure that all characteristic changes in physical space are reflected in digital space in real time. Due to frequent information synchronization, transmission costs cannot be simply ignored. Assuming that all physical entities transmit information at the same frequency, the transmission cost is determined by the number of physical entities; more entities lead to higher transmission costs.

[0105] Beyond reproducing information in physical space, digital twins can analyze the interactions between different entities and their environment in a digitally simulated space, even if these actions and entities do not actually occur in the physical space. Therefore, if only K drones are deployed in physical space (real space), a digital twin can generate an additional MK drones to analyze the different trajectory effects when there are M drones. However, due to technological limitations and the complexity of predicting the future, the simulation results cannot be completely identical to the actual physical situation. In this paper, we consider the scenario where a drone is generated in digital space (virtual space) without a corresponding physical entity. In this case, a noise bias δ ​​will appear when calculating the drone's transmission rate to the user, where the magnitude of the noise bias δ ​​is related to the construction cost of the digital twin. Generally, digital twin systems are equipped with abundant computing resources, so the drone trajectory controller is deployed in digital space to reduce decision latency. Furthermore, since digital space contains all environmental information, including information collected from physical space and additional drone data generated in digital space, deploying the controller in digital space eliminates the cost of transmitting data to the controller.

[0106] 6. Hybrid Deployment Strategy of Physical and Virtual Drones

[0107] Traditionally, to train a deep reinforcement learning agent to optimize the trajectory of a network with M drones, an equal number of drones need to be deployed in physical space. Different flight trajectories are then attempted, and the agent continuously interacts with the environment to evaluate the decision-making performance. Clearly, this training method results in high hardware costs for purchasing drones, and the deep reinforcement learning agent needs to try many different trajectories to converge, leading to high drone flight energy costs. However, even without physically deployed drones, a digital twin system can evaluate the effectiveness of drone flight trajectories in digital space. However, since the digital twin system is considered imperfect in this invention, its analysis results in digital space may differ from the actual physical situation. In this invention, the case considered is that the transmission rate simulated by adding noise differs from the actual transmission rate. Therefore, the transmission rate between the virtually generated drone j and user i is: in, It is a Gaussian random variable with mean 0 and variance δ. The magnitude of δ is related to the construction cost of the digital twin system; the smaller δ is, the greater the construction cost of the digital twin system. To evaluate the construction cost ∈ of the digital twin system, in this invention, the construction cost of the digital twin system is considered to increase exponentially with the decrease of noise δ. Specifically, the construction cost of the digital twin system is: ∈ = αe βδTo address the noise bias and hardware cost issues in training within digital twin systems, this invention proposes a method for the hybrid deployment of physical and virtual drones. Specifically, to train an agent (controller) that controls M drones, only K drones need to be deployed in physical space, while the remaining MK drones can be virtually generated in digital space by the digital twin system. Drones virtually generated by the digital twin system do not require the purchase of physical equipment or the expenditure of energy to attempt different flight trajectories, thus reducing hardware costs during training. Simultaneously, the deployment of physical drones improves the training performance of deep reinforcement learning by reducing simulation errors in the digital twin system. However, more physical drones also increase the data transmission cost between physical and digital spaces. Assuming the data synchronization period between physical and digital spaces is a constant value, and the number of training iterations for the deep reinforcement learning agent is the same in all scenarios, then the flight distance and data transmission volume of each physical drone during training are identical. Therefore, the cost of deploying drones in physical space is linearly related to the number of drones.

[0108] 7. Establish a mathematical model for the problem

[0109] Therefore, using an imperfect digital twin system to assist in training deep reinforcement learning agents in a drone network presents two main challenges: optimizing flight trajectories to improve network performance and reducing training costs. By employing a hybrid deployment strategy of physical and virtual drones, this invention fully leverages the advantages of both physical and digital spaces while also considering the construction cost of the digital twin system. Therefore, this invention primarily involves three optimization variables: the flight direction v of each drone, the number K of drones deployed in the physical space, and the noise bias δ ​​within the digital twin system. Thus, the mathematical model for the problem to be solved is the aforementioned objective function.

[0110] For example, Figure 5 This is a schematic diagram illustrating the principle of training the intelligent agent according to the present invention. Figure 5 As shown, drones are deployed in both physical and digital spaces, and these drones can communicate with corresponding user devices via networks. Figure 5Solid lines with arrows represent the actual communication transmission rate of the drones, while dashed lines with arrows represent the communication transmission rate estimated by the digital space (DT). The DRL agent is also deployed in the digital space. The number of drones deployed in the physical space and the noise deviation δ of the DT are determined by the output of another agent, which in turn affects the training of that agent. During training, the current communication rate, state (position), and reward of the drones in the physical space, as well as the current communication rate, state, and reward of the drones in the digital space, are transmitted to the DRL agent in the digital space. The DRL agent uses this data to determine the next action (flight direction) of each drone. Each drone flies according to its flight direction and speed, generating a corresponding flight trajectory and obtaining the next moment's communication rate, state (position), and reward. This data is then transmitted back to the DRL agent in the digital space to determine the actions of each drone in the moment after that. After a period of time, the flight data from this period is used as sample data to train the DRL agent.

[0111] The technical effects of the embodiments of the present invention will be further illustrated below using simulation experimental data.

[0112] In this invention, N = 100, M = 4, p = 100mW, σ 2 = -174dBm / Hz to evaluate the proposed method. In addition, the altitude of all UAVs is 5 meters and the speed is 8 m / s. The simulation results of the proposed algorithm are presented in two aspects. First, we evaluate the convergence performance of different schemes. Second, we compare the network utility of the proposed scheme with other benchmark schemes by considering the summation rate and cost. In order to evaluate the performance of the proposed method, the following algorithms are used for comparison: (1) training deep reinforcement learning agents using UAVs fully deployed in physical space; (2) training deep reinforcement learning agents using a fixed number of UAVs deployed in physical space, while the remaining UAVs are generated in digital space; (3) jointly optimizing the number of UAVs deployed in physical space, the noise bias of the digital twin system and the UAV trajectory design (i.e., the method of the present invention).

[0113] To demonstrate performance in real-world deployment scenarios, we tested the convergence rate in physical space. Figure 6 The convergence performance of physical drones (real drones) and different numbers of digital twin systems is shown (when the total number of drones M is fixed, the number of digital twin systems (virtual drones) varies with the number of real drones K). Figure 6 The five colored lines in the graph represent the moving average of the past 200 events. Figure 6The noise level of all digital twin systems is 0.8. It can be seen that the convergence performance of training using digital twin systems is inferior to that of physical drones. This is because the noise of the digital twin systems interferes with the agent's learning, making it difficult for the agent to find the optimal behavioral strategy. Furthermore, it can be observed that the convergence performance increases with the number of digital twin systems (i.e., ...). Figure 6 The decrease is due to the reduction of K in the equation. With the noise level being the same for each digital twin system, increasing the number of digital twin systems will make the system noisier, thus degrading its performance.

[0114] Figure 7 The convergence performance of physical drones and digital twin systems under different noise levels is shown, with four digital twin systems used in all cases except for the physical drone. It can also be seen that the convergence performance of the digital twin system is inferior to that of the physical drone, and the system performance deteriorates with increasing noise.

[0115] Figure 8 This paper compares the network utility of physical drones and different digital twin systems. It shows that training with a digital twin system yields better results than training with a physical drone. While physical drones are unaffected by noise, their high cost reduces network utility. Furthermore, it demonstrates that the noise level of the digital twin system significantly impacts network utility. This is because while building a more accurate digital twin system can reduce noise interference, it also increases the system's cost. Therefore, it is necessary to find the optimal digital twin system based on network characteristics.

[0116] Figure 9 The graph illustrates the relationship between network utility and cost weights, represented by alpha. The fixed digital twin system in the figure uses four digital twin systems with a noise level of 0.9. It can be seen that the network utility of each method decreases as the cost weight increases. The physical drone's network utility decreases the fastest because it has the highest cost. With low cost weights, the fixed digital twin system's network utility is lower than that of the physical drone due to its weaker performance. With high cost weights, the fixed digital twin system's network utility is worse than that of the physical drone because of its lower cost. The reinforcement learning-based digital twin system approach effectively balances system performance and cost, achieving optimal network utility.

[0117] Figure 10The relationship between network utility and cost weights, denoted as beta, is shown. Since cost and noise are negatively correlated, the value of beta is always negative. It can be seen that the advantage of the digital twin system approach becomes more pronounced as the absolute value of beta increases. This is because a larger absolute value of beta means that the cost decreases faster as noise increases, while the cost of the physical drone remains unchanged, thus highlighting the cost advantage of the digital twin system approach. Similarly, deep reinforcement learning's choice of digital twin systems enables it to achieve better network utility than other methods.

[0118] The simulation results show that the method proposed in this invention significantly reduces the training costs of deep reinforcement learning in actual training, such as energy and hardware procurement, and ensures high network transmission rates by optimizing UAV trajectories. By applying the proposed method in the network, even in imperfect digital twin systems, the training cost based on deep reinforcement learning optimization can be reduced. This means that effective decision-making can be made in multi-UAV networks under imperfect digital twin systems.

[0119] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A decision-making method for multi-UAV networks based on imperfect digital twins, characterized in that, include: Obtain the first position of each drone in a drone swarm with specific environmental parameters at the current moment, and the second position of each user device among the multiple user devices served by the drone swarm; The specific environmental parameter characterizes the number of actual drones in the drone swarm. And the noise deviation between real drones and virtual drones. The specific environmental parameters are determined by the pre-trained environmental decision network based on preset environmental index parameters; the pre-trained environmental decision network is obtained by training the initial environmental decision network through unlabeled unsupervised learning based on multiple sets of sample data and the first pre-trained action decision network. The first position and the second position are processed by a second pre-trained action decision network to obtain the flight direction of each UAV at the current time; the second pre-trained action decision network is obtained by training the initial action decision network with reinforcement learning using sample data obtained under the specific environmental parameters. The gradient update function of the environment decision network, the gradient update function of the action decision network, and the reward function of the action decision network are all determined based on an objective function. This objective function aims to maximize the transmission rate of all user devices within the average time while reducing training costs. The expression for the objective function is as follows: ; in, For user equipment exist Location at any given moment For drones exist Location at any given moment For drones exist Location at any given moment This is the path loss constant. For reference distance, for Time User Equipment With drones The distance between them For user equipment Transmission power, For noise power, The total number of drones in the drone swarm. The total number of user devices. These are parameters used to characterize the importance of performance indicators. This is a parameter used to characterize the cost of virtual drones in a drone swarm. These are parameters used to characterize the cost of real drones in a drone swarm. These are parameters used to characterize the difficulty of constructing virtual drones in a drone swarm. The number of actual drones in the drone swarm. This refers to the noise deviation between real and virtual drones in a drone swarm. For the log function, For drones The speed of time For the location of the user equipment, For drones Location at any given moment It is an absolute value function. It is the action decision network, The road loss factor; constraint (4a) characterizes the user equipment. and drones The distance between them; constraint (4b) characterizes the flight speed of each UAV in the UAV swarm as a constant; in constraint (4c), the action decision network Used to determine the drone's location Flight direction at any given time, including noise deviation Influence by affecting the loss function during neural network training. The parameters affect Performance; Constraint (4d) constrains the update of the drone's position; Constraint (4e) constrains the number of real drones in the drone swarm to a maximum of M and a minimum of 0; Constraint (4f) constrains Greater than or equal to 0.

2. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 1, characterized in that, The drone swarm includes multiple real drones and multiple virtual drones; The preset environmental indicator parameters include: parameters used to characterize the importance of performance indicators. Parameters used to characterize the cost of virtual drones in the drone swarm Parameters used to characterize the cost of the actual drones in the drone swarm And parameters used to characterize the construction difficulty of virtual drones in the drone swarm. .

3. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 1, characterized in that, The expression for the gradient update function of the environmental decision network is as follows: ,in, Let the objective function be... For the environmental decision network Network parameters, For the environmental decision network learning rate, For about Partial derivatives of the parameters, For about The partial derivative with K, This refers to the action decision network.

4. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 1, characterized in that, The gradient update function of the action decision network is expressed as follows: ,in, For the action decision network Network parameters, For the action decision network learning rate, For drones Reward value at any moment For drones The speed of time For drones Location at any given moment For drones The speed of time For drones Location at any given moment , for The target network, In order to make When taking the maximum value The value of .

5. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 1, characterized in that, Before obtaining the first position of each drone in the drone swarm with specific environmental parameters at the current moment, and the second position of each user device among the multiple user devices served by the drone swarm, the method further includes: Obtain multiple sets of sample data corresponding to multiple time points; each set of sample data includes: the position of the drone at time t, the flight direction of the drone at time t, and the reward value of the drone at time t; the reward value of the drone at time t represents the reward value obtained by the drone at the position at time t and executing the flight direction at time t; t is an integer greater than 0; The output of the initial action decision network is used as the input of the initial environment decision network, and the output of the initial environment decision network is also used as the input of the initial action decision network to obtain the initial serial network. Using the multiple sets of sample data, the initial action decision network in the initial concatenated network is iteratively trained through reinforcement learning to obtain a first pre-trained concatenated network that includes the first pre-trained action decision network and the initial environment decision network. Using the multiple sets of sample data, the initial environment decision network in the first pre-trained concatenated network is iteratively trained through unlabeled unsupervised learning to obtain the pre-trained environment decision network.

6. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 5, characterized in that, The expression for the drone's reward value at time t is as follows: ; in, The total number of drones in the drone swarm. The total number of user devices. This is the path loss constant. For reference distance, for Time User Equipment With drones The distance between them For user equipment Transmission power, For noise power, , The mean is 0 and the variance is Gaussian random variables.

7. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 1, characterized in that, Before processing the first position and the second position using the second pre-trained action decision network to obtain the flight direction of each UAV at the current moment, the method further includes: Under the specific environmental parameters, multiple sets of sample data corresponding to multiple times are obtained; each set of sample data under the specific environmental parameters includes: the position of the UAV at time c, the flight direction of the UAV at time c, and the reward value of the UAV at time c; the reward value of the UAV at time c represents the reward value obtained by the UAV at the position at time c and the flight direction at time c; c is an integer greater than 0. Using the multiple sets of sample data under the specific environmental parameters, the initial action decision network is iteratively trained through reinforcement learning to obtain the second pre-trained action decision network.

8. The decision-making method for multi-UAV networks based on imperfect digital twins according to claim 1, characterized in that, The network structures of the initial environment decision network and the action decision network are the same as those of the MLP network.