A multi-unmanned aerial vehicle communication topology network optimization method based on improved Q-learning
Patent Information
- Application Number
- CN202510273788.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-03-07
AI Technical Summary
[0002]大规模无人机编队具备灵活性高、协同性强和成本低廉的特点,但其在协同工作时由于编队队形动态变化会导致局部通信代价高和通信时延大等问题
[0014]本发明技术方案提出一种基于改进Q-Learning的多无人机通信拓扑网络优化方法,首先分析无人机编队通信网络影响因素,构建多无人机通信拓扑网络路由评价模型;其次,在Q-Learning算法的基础上,设计贪婪因子自适应调节机制,提升Q-Learning算法学习探索能力;最后,依据编队控制所需最小通信的要求,考虑多无人机编队控制对通信网络需求,优化设计编队无人机间的通信连通情况。本发明方法能够克服大规模编队约束下基于规则的路由设计方法使得算法的计算复杂度上升的问题,最佳平衡编队控制基础上最小路由需求,解决了攻防博弈对抗与信息流最小需求的无人机编队通信网络优化问题。
Smart Images

Figure CN120342522B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-UAV communication topology design technology, specifically to a multi-UAV communication topology network optimization method based on improved Q-Learning, which is particularly suitable for the design of large-scale UAV swarm communication topology networks. Background Technology
[0002] Large-scale UAV swarms offer advantages such as high flexibility, strong coordination, and low cost. However, dynamic changes in formation during collaborative operations can lead to issues like high local communication costs and significant communication latency. Considering that UAV swarm flight formation control is primarily supported by communication links, optimizing communication routing costs through formation can effectively assess the load balancing, resilience, and robustness of the swarm communication topology network in dynamic environments.
[0003] Q-Learning technology has been widely researched and applied in fields such as artificial intelligence, machine learning, and optimization. It is one of the core technologies driving system intelligence, enabling nodes to learn the optimal action strategy for the current node through interaction with the environment without prior information, thus rapidly improving formation communication capabilities within a limited time. Therefore, how to use Q-Learning algorithms to achieve fast communication topology optimization for large-scale formations is a crucial problem that urgently needs to be solved. Summary of the Invention
[0004] The objective of this invention is to provide a method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on an improved Q-Learning algorithm.
[0005] To achieve the above-mentioned objectives, the present invention provides a method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning, comprising the following steps:
[0006] Based on the process of UAV formation information interaction network, the influencing factors of UAV formation communication routing model are analyzed to obtain the corresponding evaluation model of each influencing factor, thereby constructing a formation communication network performance evaluation model.
[0007] Treating the members of a drone swarm as nodes in a communication network, and considering that the communication routing process requires traversing each node, the size N of the swarm members is defined as the state space S. UAVs And define the leader drone node, and set the state space S UAVs The current drone node's neighboring node S i Defined as action A i Forming Action Space A UAVs Action A will be executed. i The harvest as a reward R i+1The policy process of state nodes selecting actions according to fixed rules is defined as policy π(A|S), thus obtaining the Markov decision model;
[0008] The cumulative reward function is defined based on the optimal strategy of maximizing rewards; considering the randomness of the algorithm's learning process, the mathematical expectation of the cumulative reward function is defined as the state value function of the formation node; combined with the aforementioned action A... i The state value function is transformed into a state action value function. Under the constructed Markov decision model, the optimal state action value function is finally determined by training with the Q-Learning algorithm.
[0009] By acquiring information such as the structure, communication equipment, location, and data link distance limitations of the UAV formation, and based on the formation communication network performance evaluation model, the reward matrix R(S) is designed to maximize the reward of the formation communication network performance evaluation model. i A i By combining the optimal state action value function, the reward matrix of the formation drones is obtained;
[0010] Considering the contradiction between node exploration and development when constructing a large-scale UAV swarm communication topology network, a greedy factor function that adaptively changes based on the iteration number of the Q-Learning algorithm is constructed to make the UAV swarm reward matrix converge. The node with the farthest physical distance from the leader UAV node is selected as the initial node. The optimal action for constructing the routing link is based on the converged UAV swarm reward matrix, and finally the main communication link of the UAV swarm is obtained.
[0011] Preferably, the evaluation models corresponding to each influencing factor include a communication strength evaluation model, a communication cost evaluation model, and a probability evaluation model of being detected by the other party.
[0012] Preferably, the communication strength assessment model is established based on the physical distance between UAVs, the maximum reachable distance of the UAV communication chain, and the path dissipation index between UAVs; the communication cost assessment model is established based on the optimal effective distance of the UAV formation communication chain, the physical distance between UAVs, the maximum reachable distance of the UAV communication chain, and the path dissipation index between UAVs; and the probability assessment model of being detected by the other party is established based on the bandwidth of the UAV formation communication network, the power consumption of the terminal device, the physical distance between UAVs, and the optimal effective distance of the UAV formation communication chain.
[0013] Preferably, if an initial node exists outside the main communication link, the node with the shortest physical distance from the slave is selected and connected to the main communication link to optimize the main communication link of the UAV formation.
[0014] This invention proposes a multi-UAV communication topology network optimization method based on improved Q-Learning. First, it analyzes the influencing factors of UAV formation communication networks and constructs a routing evaluation model for multi-UAV communication topology networks. Second, based on the Q-Learning algorithm, a greedy factor adaptive adjustment mechanism is designed to enhance the learning and exploration capabilities of the Q-Learning algorithm. Finally, considering the minimum communication requirements for formation control and the communication network demands of multi-UAV formation control, the communication connectivity between UAVs in the formation is optimized. This method overcomes the problem of increased computational complexity caused by rule-based routing design methods under large-scale formation constraints, achieving an optimal balance between minimum routing requirements for formation control and solving the UAV formation communication network optimization problem involving both offensive and defensive game theory and minimum information flow requirements. Attached Figure Description
[0015] Figure 1 This is a block diagram of a comprehensive evaluation architecture for the performance of a drone communication network provided in an embodiment of the present invention;
[0016] Figure 2 This is a schematic diagram illustrating the initial communication link connectivity of a drone formation provided in an embodiment of the present invention;
[0017] Figure 3 This is a network structure diagram of an improved Q-Learning algorithm provided in an embodiment of the present invention;
[0018] Figure 4 This is a graph showing the change in reward function values of UAVs 7, 8, and 10 during the training process of an improved Q-Learning algorithm provided in this embodiment of the invention.
[0019] Figure 5 This is a schematic diagram of a drone formation communication topology main chain provided in an embodiment of the present invention;
[0020] Figure 6 This is a schematic diagram of the entire chain of UAV formation communication topology provided in an embodiment of the present invention. Detailed Implementation
[0021] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0022] This invention provides a method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning, comprising the following steps:
[0023] Step 1: Analyze the influencing factors of the UAV formation communication routing model and construct a formation communication network performance evaluation model. Specifically, Step 1 includes the following steps:
[0024] Analysis of the UAV swarm information exchange network reveals that the main factors influencing the final communication routing model of the swarm include communication strength between UAVs, communication cost, and the probability of being detected by the other side. Further analysis of the specific causes of each influencing factor leads to the establishment of a communication strength assessment model using mathematical functions. Communication cost assessment model and the probability assessment model of being detected by the other party
[0025] Consider the physical distance D between drones m and n mn The maximum reachable distance of the UAV communication link is D l,max A communication strength assessment model is established based on the path dissipation index υ between drones and other communication systems. as follows:
[0026]
[0027] Considering the optimal effective range D of the UAV formation communication link l The physical distance D between drones m and n mn The maximum reachable distance of the UAV communication link is D l,max Establish a communication cost assessment model based on the path dissipation index υ between drones and other communication systems. as follows:
[0028]
[0029] Considering the bandwidth B of the drone swarm communication network, the power consumption P of the terminal unit, and the physical distance D between drones m and n. mn Optimal range of action D for communication link with UAV formation l Establish a probability assessment model for being detected by the other party as follows:
[0030]
[0031] Among them, f w B represents the communication frequency band coefficient. c and B sta P represents the communication bandwidth of the formation network in its current and standard states, respectively. max and P sta These represent the maximum power and nominal power consumption of the network terminal, respectively. min This indicates the closest distance to the other drone.
[0032] Taking into account the factors mentioned above that affect the performance of UAV swarm communication networks, a comprehensive evaluation model for the performance of UAV swarm communication networks is established. mn for:
[0033]
[0034] Step 2: Construct a Markov decision model for UAV formation using the Q-Learning algorithm, and train the state-action value function using the Q-Learning algorithm under the constructed Markov decision model until the optimal state-action value function is determined. Specifically, Step 2 includes the following steps:
[0035] A Markov decision model for UAV formation is constructed based on the components of the Q-Learning algorithm in reinforcement learning. Reinforcement learning consists of a state space S. UAVs Action Space A UAVs Reward R, policy π, and state transition probability Composed of five key elements, its interaction with the environment can be modeled as a Markov decision process. In this embodiment of the invention, the members of the UAV swarm are regarded as nodes in a communication network. Considering that the communication routing process needs to traverse each node, the swarm size N is defined as the state space S. UAVs :
[0036] S UAVs ={S1,S2,…,S N+1}
[0037] Among them, S N+1 This represents a virtual leadership drone node.
[0038] Set the current drone node S i The surrounding adjacent nodes are defined as action space A. UAVs The elements in the set are used as the action A. i The drone formation will be placed at node S i Execute action A i The reward R(S) obtained i A i S′) is defined as R i+1 The strategy process of selecting actions for state nodes according to fixed rules is defined as π(A|S)=P(A UAVs =A|S UAVs =S), then the Markov property in the interaction process with the environment can be expressed as:
[0039] P{S′,R i+1 |S i A i ,S i-1 A i-1,…,S1,A1}=P{S′,R i+1 |S i A i}
[0040] Furthermore, to maximize the reward for the optimal strategy, the cumulative reward function is defined as follows:
[0041]
[0042] In the formula, γ is a constant between 0 and 1, called the discount factor, which is used to balance the real-time reward and the expected reward. Historical cumulative returns R are used at each Markov decision step. i a This allows us to evaluate the effectiveness of the current strategy and then select the optimal strategy.
[0043] Considering the randomness of the strategy during the algorithm learning process, the cumulative reward function R... i a The mathematical expectation is defined as the value function of the formation node state Vπ(S) i )for:
[0044]
[0045] Consideration node execution action A i Then the state value function V π (S i This can be transformed into a state-action value function Q. π (S i i,A i ), i.e. Q π (S i A i ) = E π (R i a |S i UAVs =S i A i UAVs =A i Then, from the formula, we can obtain:
[0046] Q π (S i A i ) = E π (R i+1 +γQ π (S′i,A′)|S i UAVs =S i A i UAVs =A i )
[0047] Furthermore, the state-action value function can be derived as follows:
[0048]
[0049] Then, through training and learning, the algorithm's optimal policy π * The corresponding optimal state action value function Q * (S i A i )for:
[0050]
[0051] The Q-Learning algorithm, through interaction with the environment, aims to find the optimal state-action value function Q. π (S i i,A i The process can be divided into two parts: designing the reward matrix of the Q-Learning algorithm and training and updating the reward matrix to obtain the optimal state-action value function.
[0052] Step 3: Based on the spacing D between drone formation members mn Maximum reachable distance D of the information communication link l,max Establish the maximum available communication topology for the formation, with the initial reward matrix Q0 as follows:
[0053]
[0054] The reward matrix design takes into account the UAV swarm communication performance evaluation model established in step 1. mn Define the maximum reward for establishing a communication routing link between the leader and slave drones in a drone swarm as Θ. max Then, the reward matrix R(S) will be designed during the training and learning process of the formation drones. i A i )for:
[0055]
[0056] Among them, (S) N+1 ,S n ) represents the virtual leader drone node S N+1 To node S n Establish a communication connection, (S m ,S n ) represents the drone node S m To node S n Establish a connection.
[0057] As can be seen from step 2, the state of the formation drone i is S. i When, for example, if you choose to perform action A iAfter interacting with the environment, the state becomes S′, and a reward R(S′) is generated. i A i Based on the state-action value function, the update formula for the reward matrix Q of the formation UAVs during the learning process is defined as follows:
[0058] Q(S′,A′)=(1-α)Q(S i A i )+α[R(S i A i )+γmaxQ'(S′,A′)]
[0059] In the formula, α is a constant between 0 and 1, called the learning rate, which represents the degree of correlation between the current result and the previous training result. When updating matrix Q through the formula, if the value of Q(S′,A′) decreases, it indicates that the selected action is not the optimal action, and the selection of this action will be reduced in the next training.
[0060] Step 4: Design a greedy factor adaptive adjustment mechanism to improve the traditional Q-Learning algorithm's exploration ability in the early stages of environmental interaction and its learning ability in the later stages. Specifically, Step 4 includes the following steps:
[0061] Considering that the general ε-greedy strategy uses a fixed probability ε to randomly select the next drone to connect with the current node, it is insufficient to quickly balance the contradiction between the exploration and development capabilities of nodes when constructing a large-scale drone swarm communication topology network. Therefore, a greedy factor function ζ(Iter) is designed based on the adaptive variation of the number of iterations of the Q-Learning algorithm. i ):
[0062]
[0063] In the formula, Iter i Indicates the current iteration number of the algorithm. and This is the overshoot parameter. The adjustment function is ζ(Iter i The value of ζ (Iter) is relatively large in the initial iterations of the algorithm, allowing the drone nodes to fully explore connectable neighbor nodes in the early stages, increasing the use of information in the environment; in the later stages of the algorithm iterations, considering that the routing nodes have learned the experience of optimal node connectivity during training, the adjustment function ζ (Iter) is adjusted. i The value of ) decreases compared to the initial iteration, which reduces the probability of the UAV routing node randomly selecting a connected node, effectively increasing the reward of the overall formation routing network construction process. Furthermore, the adaptive greedy factor avoids the problem of UAV nodes failing to learn from experience during the exploration phase due to a constant greedy factor, ultimately choosing a suboptimal communication route data link.
[0064] Repeating the above steps will yield a large amount of data regarding the drone node connectivity actions and rewards, making the reward matrix Q... g It eventually converges.
[0065] Step 5: Optimize the formation communication network topology based on the improved Q-Learning algorithm, achieving dynamic optimization and updating of the formation communication network. Specifically, Step 5 includes the following steps:
[0066] The specific process of optimizing the UAV formation communication topology network using the improved Q-Learning algorithm designed in this invention is as follows:
[0067] S1. Obtain information such as the structure, communication equipment, location, and data link distance limits of the UAV formation, and calculate the reward matrix Q of the current state according to the formula in step (3). mn ;
[0068] S2. Model the formation members as routing nodes and abstract them into the structural form of Q network input to construct a UAV formation Markov decision model;
[0069] S3. Based on the drone formation, a routing action space can be established. Neighboring nodes are regarded as optional actions. The optional action with the largest reward Q is selected. The instant reward value of the action is calculated according to the formula and the reward Q matrix is updated according to the formula. Then, the process is iterated repeatedly until the reward Q matrix converges.
[0070] S4. Select the slave device that is physically farthest from the leader as the initial node, and construct the optimal action for routing links based on the converged reward Q matrix to complete the construction of the main communication link of the UAV formation; if there is a slave device that is not in the main communication link, select the node that is physically farthest from the slave device and connect it to the main communication link to complete the optimization of the UAV formation communication topology network.
[0071] This invention provides a method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning. First, it analyzes the influencing factors of UAV formation communication networks, constructs a UAV communication routing evaluation model, designs a greedy factor adaptive adjustment mechanism and a reward matrix based on the routing evaluation model, and applies the Q-Learning algorithm to UAV formation communication topology optimization, thus forming a method for solving the communication topology optimization problem of UAV formations.
[0072] Example 1
[0073] Step 1: Analysis of the drone swarm information exchange network process reveals that the factors influencing the final communication routing model of the swarm mainly include the communication strength between drones, communication cost, and the probability of being detected by the other side, such as... Figure 1As shown, a communication strength assessment model is established using mathematical functions. Communication cost assessment model and the probability assessment model of being detected by the other party as follows:
[0074]
[0075] Among them, D mn Let m be the physical distance between drones m and n, and υ = 1 be the path dissipation exponent between drones in an unobstructed environment. l,max =30m is the maximum reachable distance of the UAV communication link, D l =20m is the optimal operating distance for the UAV formation communication link, B and P are the bandwidth of the UAV formation communication network and the power consumption of the terminal device, f w =1 indicates the communication frequency band coefficient, B c and B sta =40MHz represents the communication bandwidth of the formation network in its current and standard states, respectively. P max =4W and P sta =2W represents the maximum power and nominal power consumption of the network terminal, respectively.
[0076] Taking into account the factors mentioned above that affect the performance of UAV swarm communication networks, a comprehensive evaluation model for the performance of UAV swarm communication networks is established. mn for:
[0077]
[0078] In the formula, take γ c =[γ 1,c ,γ 2,c ,γ 3,c = [0.5, 0.2, 0.3].
[0079] Step 2: Set the drone formation size to 15, and determine whether initial communication routes can be established between drone formations. Figure 2 As shown, the state space S is defined based on the size of the drone swarm members. UAVs for:
[0080] S UAVs ={S1,S2,…,S 15 ,S 16}
[0081] Where S 16The virtual leader drone node has the following location information for drones 1 to 15: (0,60), (15,75), (30,90), (52.5,97.5), (76.5,87), (102,78), (31.5,55.5), (46.5,75), (97.5,70.5), (7.5,52.5), (22.5,33), (49.5,27), (63,51), (85.5,37.5), (100.5,49.5), (63,60). The location of the opposing drone is (135, -30), (157.5, -15), (150, -27), (150,27).
[0082] Define the nodes adjacent to the current drone node S7 as the action space A7. UAVs Element A in the set i Then A7 UAVs ={A2,A8,A 10 A 11}, where A2 represents the action of selecting the next drone node as S2. The reward R(S) obtained by executing action Ai at node Si with the drone formation is... i A i S′) is defined as R i+1 The strategy process of a state node selecting an action according to fixed rules is defined as π(A|S)==P(A UAVs =A|S UAVs =S), then the Markov property in the interaction process with the environment can be expressed as:
[0083] P{S′,R i+1 |S i A i ,S i-1 A i-1 ,…,S1,A1}=P{S′,R i+1 |S i A i}
[0084] To maximize the reward for the optimal strategy, the cumulative reward function is defined as follows:
[0085]
[0086] Because the strategy is random during the learning process, R i a The mathematical expectation is defined as the node state value function V. π (S i )for:
[0087]
[0088] Furthermore, consider the action parameter A of the node. i Then the state value function V π (S i This can be transformed into a state-action value function Q. π (S i i,A i ), i.e. Q π (S i A i ) = E π (R i a |S i UAVs =S i A i UAVs =A i Then, from the formula, we can obtain:
[0089]
[0090] Then, through training and learning, the algorithm's optimal policy π * The corresponding optimal state action value function Q * (S i A i )for:
[0091]
[0092] Step 3: Based on Figure 2 The initial reward matrix Q0 is established based on the connectivity of the drone swarm member nodes shown:
[0093]
[0094] Design an instant reward function R(S) i A i )for:
[0095]
[0096] Where, Θ max =1000 represents the maximum reward when the leader and slave drones in a drone swarm establish a communication routing link, (S 16 ,S8) or (S 16 ,S 13 ) represents the virtual leader node S 16 To drone node S8 or S 13 Establish a communication link, (S m ,S n ), n = 1, 2, ..., 15, m ≠ 0 represents node S m To node S n Once a communication link is established, the immediate feedback function R is:
[0097]
[0098] As can be seen from step 2, the state of the formation drone i is S. i When selecting motion parameter A i After interacting with the environment, the state becomes S′, and a reward R(S′) is generated. i A i The interactive training process is as follows: (S′), Figure 3 As shown. Based on the state-action value function, the update formula for the reward matrix Q of the formation UAVs during the learning process is defined as follows:
[0099] Q(S′,A′)=(1-α)Q(S i A i )+α[R(S i A i )+γmaxQ'(S′,A′)]
[0100] In the formula, α = 0.1 is the learning rate, which is used to represent the degree of correlation between the current result and the previous training result, and γ = 0.9 is the discount factor.
[0101] Step 4: Considering that the general ε-greedy strategy uses a fixed probability ε to randomly select the next drone to connect with the current node, it is insufficient to quickly balance the contradiction between the exploration and development capabilities of nodes when constructing a large-scale drone swarm communication topology network. Therefore, a greedy factor function ζ(Iter) is designed based on the adaptive variation of the Q-Learning algorithm iteration count. i ):
[0102]
[0103] In the formula, Iter i Indicates the current iteration number of the algorithm. and This is the overshoot parameter.
[0104] Step 5: Optimize the formation communication network topology based on the improved Q-Learning algorithm, realizing dynamic optimization and updating of the formation communication network, including the following steps:
[0105] S1. Obtain information such as the structure, communication equipment, location, and data link distance limits of the UAV formation, and calculate the reward matrix Q of the current state according to the formula in step (3). mn ;
[0106] S2. As shown in the equation, the formation members are modeled as routing nodes and abstracted into the structural form of Q network input to construct the UAV formation Markov decision model.
[0107] S3. Based on drone formations, a routing action space can be established, treating nearby nodes as optional actions and selecting the optional action with the highest reward Q. For example... Figure 4 As shown, Figure 4 (a) in the diagram is a diagram of the UAV7 training iteration process. Figure 4 (b) in the diagram is a diagram of the UAV8 training iteration process. Figure 4 (c) in the diagram shows the UAV10 training iteration process. The instantaneous reward value of the action is calculated according to the formula, and the reward Q matrix is updated based on the formula. This process is repeated until the reward Q matrix converges. The converged Q matrix is shown below:
[0108]
[0109] S4. For example Figure 5 As shown, the slave device with the greatest physical distance from the leader is selected as the initial node. Based on the converged reward Q-matrix, the optimal action for routing links is constructed, thus completing the construction of the main communication link for the UAV formation. Figure 6 As shown, if a slave device is not in the main communication link, the node with the shortest physical distance to the slave device is selected to connect it to the main communication link, thus optimizing the UAV formation communication topology network.
[0110] The foregoing details the mathematical principles and specific steps of this invention. While embodiments are provided to illustrate the specific processes and mathematical principles of this invention, it should be noted that this invention is not limited to the above embodiments. Under conditions consistent with the design concept of this invention, various variations and improvements to the embodiments can yield the same final results. Therefore, it should be understood that the embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this invention.
Claims
1. A method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning, characterized in that, Includes the following steps: Based on the process of UAV formation information interaction network, the influencing factors of UAV formation communication routing model are analyzed to obtain the corresponding evaluation model of each influencing factor, thereby constructing a formation communication network performance evaluation model. Treating the members of a drone swarm as nodes in a communication network, and considering that the communication routing process requires traversing each node, the size N of the swarm members is defined as the state space S. UAVs And define the leader drone node, and set the state space S UAVs The current drone node's neighboring node S i Defined as action A i Forming Action Space A UAVs Action A will be executed. i The harvest as a reward R i+1 The policy process of state nodes selecting actions according to fixed rules is defined as policy π(A|S), thus obtaining the Markov decision model; The cumulative reward function is defined based on the optimal strategy of maximizing rewards; considering the randomness of the algorithm's learning process, the mathematical expectation of the cumulative reward function is defined as the state value function of the formation node; combined with the aforementioned action A... i The state value function is transformed into a state action value function. Under the constructed Markov decision model, the optimal state action value function is finally determined by training with the Q-Learning algorithm. By acquiring information such as the structure, communication equipment, location, and data link distance limitations of the UAV formation, and based on the formation communication network performance evaluation model, the reward matrix R(S) is designed to maximize the reward of the formation communication network performance evaluation model. i A i By combining the optimal state action value function, the reward matrix of the formation drones is obtained; Considering the contradiction between node exploration and development when constructing a large-scale UAV swarm communication topology network, a greedy factor function that adaptively changes based on the iteration number of the Q-Learning algorithm is constructed to make the UAV swarm reward matrix converge. The node with the farthest physical distance from the leader UAV node is selected as the initial node. The optimal action for constructing the routing link is based on the converged UAV swarm reward matrix, and finally the main communication link of the UAV swarm is obtained.
2. The method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning as described in claim 1, characterized in that, The evaluation models corresponding to each influencing factor include a communication strength evaluation model, a communication cost evaluation model, and a probability of being detected by the other party evaluation model.
3. The method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning as described in claim 2, characterized in that, The communication strength assessment model is established based on the physical distance between UAVs, the maximum reachable distance of the UAV communication chain, and the path dissipation index between UAVs; the communication cost assessment model is established based on the optimal effective distance of the UAV formation communication chain, the physical distance between UAVs, the maximum reachable distance of the UAV communication chain, and the path dissipation index between UAVs; and the probability assessment model of being detected by the other side is established based on the bandwidth of the UAV formation communication network, the power consumption of the terminal device, the physical distance between UAVs, and the optimal effective distance of the UAV formation communication chain.
4. The method for optimizing the communication topology of multiple unmanned aerial vehicles (UAVs) based on improved Q-Learning as described in claim 1, characterized in that, If an initial node exists outside the main communication link, the node with the shortest physical distance from the slave is selected and connected to the main communication link to optimize the main communication link of the UAV formation.
Citation Information
Patent Citations
Distributed formation method of unmanned aerial vehicle cluster based on reinforcement learning
CN110007688A
Unmanned aerial vehicle ad hoc network adaptive routing method based on Q-Learning
CN114449608A