Multi-unmanned aerial vehicle communication topology network optimization method based on improved Q-Learning

By improving the Q-Learning algorithm, the drone communication topology network routing evaluation model is built, the greedy factor adaptive adjustment mechanism is designed, and the drone formation communication topology is optimized, which solves the high cost and delay problems of large-scale drone formation communication networks, and achieves load balancing and robustness.

CN120342522AActive Publication Date: 2025-07-18EAST CHINA INST OF COMPUTING TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510273788.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

When large-scale drone formations work together, due to the dynamic changes in formation, local communication costs and communication delays are high, and the prior art is difficult to quickly optimize communication topology networks.

Method used

The improved Q-Learning algorithm is used to build a multi-UAV communication topology network routing evaluation model, design a greedy factor adaptive adjustment mechanism, optimize the communication and connectivity between drone formations, and train the optimal state action value function through the Q-Learning algorithm to optimize the communication topology of drone formations.

Benefits of technology

It effectively solves the problem of high computational complexity of routing design under large-scale formation constraints, realizes load balancing and robustness of formation communication networks, and optimizes the communication topology network under formation control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342522A_ABST
    Figure CN120342522A_ABST
Patent Text Reader

Abstract

The invention provides a multi-unmanned aerial vehicle communication topology network optimization method based on improved Q-Learning, and the method comprises the steps: firstly analyzing an unmanned aerial vehicle formation communication network influence factor, and constructing a multi-unmanned aerial vehicle communication topology network routing evaluation model; secondly, on the basis of the Q-Learning algorithm, designing a greedy factor adaptive adjustment mechanism, and improving the learning and exploration capability of the Q-Learning algorithm; and finally, according to the requirement of minimum communication required by formation control, considering the requirement of multi-unmanned aerial vehicle formation control for a communication network, and optimally designing the communication connection condition between the formation unmanned aerial vehicles. According to the method, the problem that the calculation complexity of an algorithm is increased due to a rule-based routing design method under large-scale formation constraint can be solved, the minimum routing requirement on the basis of optimal balance formation control is met, and the unmanned aerial vehicle formation communication network optimization problem of the attack and defense game confrontation and the minimum information flow requirement is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-UAV communication topology design, and in particular to a multi-UAV communication topology network optimization method based on improved Q-Learning, which is particularly suitable for the design of large-scale UAV swarm communication topology network. Background Art

[0002] Large-scale UAV formations are characterized by high flexibility, strong coordination and low cost, but when working together, the dynamic changes in the formation will lead to high local communication costs and long communication delays. Considering that the formation control of UAV formation flight is mainly supported by communication links, optimizing the communication routing cost through formation formation can effectively evaluate the load balancing, anti-destruction and robustness of the formation communication topology network in a dynamic environment.

[0003] Q-Learning technology has been widely studied and applied in the fields of artificial intelligence, machine learning, and optimization. It is one of the core technologies that promote system intelligence. It can learn the optimal action strategy of the current node by interacting with the environment without prior information, so that the communication capability of the formation can be quickly improved within a limited time. Therefore, how to use the Q-Learning algorithm to achieve fast communication topology network optimization for large-scale formations is an important issue that needs to be solved urgently. Summary of the invention

[0004] The purpose of the technical solution of the present invention is to provide a multi-UAV communication topology network optimization method based on an improved Q-Learning algorithm.

[0005] In order to achieve the above-mentioned invention object, the technical solution of the present invention provides a multi-UAV communication topology network optimization method based on improved Q-Learning, comprising the following steps:

[0006] According to the process of UAV formation information interaction network, the influencing factors of UAV formation communication routing model are analyzed to obtain the evaluation model corresponding to each influencing factor, so as to construct the formation communication network effectiveness evaluation model;

[0007] The members of the UAV formation are regarded as nodes in the communication network. Considering that the communication routing process needs to traverse each node, the size of the formation member N is defined as the state space S UAVs And define the leader drone node, the state space S UAVs The nodes S adjacent to the current drone node in i Defined as performing action A i Forming action space A UAVs , will execute action A i The harvest as a reward R i+1, the process of the policy for the state node to select actions according to fixed rules is defined as policy π(A|S), and a Markov decision model is obtained;

[0008] Define the cumulative reward function according to the optimal policy of maximizing the reward; according to the randomness of the algorithm learning process, define the mathematical expectation of the cumulative reward function as the state value function of the formation nodes; combine the executed action A i Convert the state value function into a state-action value function, and under the constructed Markov decision model, train it through the Q-Learning algorithm, and finally converge to determine the optimal state-action value function;

[0009] Obtain information such as the structure, communication equipment, position, and data link distance limit of the UAV formation. According to the formation communication network effectiveness evaluation model, obtain the reward matrix design R(S i ,A i ) when the maximum reward of the formation communication network effectiveness evaluation model is obtained, and combine the optimal state-action value function to obtain the UAV formation reward matrix;

[0010] Considering the contradiction between node exploration and development in the construction of a large-scale UAV formation communication topology network, construct a greedy factor function that adaptively changes according to the number of iterations of the Q-Learning algorithm, so that the UAV formation reward matrix converges, select the node with the farthest physical distance from the leading UAV node as the initial node, and construct the optimal action of the routing link according to the converged UAV formation reward matrix, and finally obtain the main communication link of the UAV formation.

[0011] Preferably, the evaluation models corresponding to the respective influencing factors include a communication intensity evaluation model, a communication cost evaluation model, and a probability of being detected by the other party evaluation model.

[0012] Preferably, establish the communication intensity evaluation model according to the physical distance between UAVs, the maximum reachable distance of the UAV communication link, and the path dissipation index between UAVs; establish the communication cost evaluation model according to the optimal operating distance of the UAV formation communication link, the physical distance between UAVs, the maximum reachable distance of the UAV communication link, and the path dissipation index between UAVs; establish the probability of being detected by the other party evaluation model according to the communication network bandwidth of the UAV formation, the terminal power consumption, the physical distance between UAVs, and the optimal operating distance of the UAV formation communication link.

[0013] Preferably, if there is an initial node outside the main communication link, select the node with the shortest physical distance from the slave and connect it to the main communication link to optimize the main communication link of the UAV formation.

[0014] The technical solution of the present invention proposes a method for optimizing a multi-UAV communication topology network based on improved Q-Learning. First, analyze the influencing factors of the UAV formation communication network and construct a routing evaluation model for the multi-UAV communication topology network. Secondly, based on the Q-Learning algorithm, design a greedy factor adaptive adjustment mechanism to improve the learning and exploration ability of the Q-Learning algorithm. Finally, according to the requirement of the minimum communication required for formation control, consider the communication network requirements of multi-UAV formation control, and optimize the communication connectivity between formation UAVs. The method of the present invention can overcome the problem that the computational complexity of the algorithm increases due to the rule-based routing design method under large-scale formation constraints, best balance the minimum routing requirements based on formation control, and solve the optimization problem of the UAV formation communication network for offensive and defensive game confrontation and minimum information flow requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 FIG. is a block diagram of a comprehensive evaluation architecture for the effectiveness of a UAV communication network provided by an embodiment of the present invention;

[0016] Figure 2 FIG. is a schematic diagram of the initial communication link connectivity of a UAV formation provided by an embodiment of the present invention;

[0017] Figure 3 FIG. is a network structure diagram of an improved Q-Learning algorithm provided by an embodiment of the present invention;

[0018] Figure 4 FIG. is a graph showing the change of the reward function values of UAVs 7, 8, and 10 during the training process of an improved Q-Learning algorithm provided by an embodiment of the present invention;

[0019] Figure 5 FIG. is a schematic diagram of the main chain of the UAV formation communication topology provided by an embodiment of the present invention;

[0020] Figure 6 FIG. is a schematic diagram of the full chain of the UAV formation communication topology provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0022] An embodiment of the present invention provides a method for optimizing a multi-UAV communication topology network based on improved Q-Learning, including the following steps:

[0023] Step 1: Analyze the influencing factors of the UAV formation communication routing model and construct an evaluation model for the formation communication network effectiveness. Specifically, Step 1 includes the following steps:

[0024] By analyzing the process of the UAV formation information interaction network, it can be seen that the factors affecting the final communication routing model of the formation mainly include the communication intensity, communication cost, and the probability of being detected by the other party between UAVs. Further analyze the specific causes of each influencing factor, and establish a communication intensity evaluation model in the form of a mathematical function Communication cost evaluation model And the probability of being detected by the other party evaluation model

[0025] Consider the physical distance D between UAVs m and n mn 、the maximum reachable distance D of the UAV communication link l,max And the path dissipation index υ between UAVs to establish a communication intensity evaluation model As follows:

[0026]

[0027] Consider the optimal operating distance D of the UAV formation communication link l 、the physical distance D between UAVs m and n mn 、the maximum reachable distance D of the UAV communication link l,max And the path dissipation index υ between UAVs to establish a communication cost evaluation model As follows:

[0028]

[0029] Consider the communication network bandwidth B of the UAV formation, the terminal power consumption P, the physical distance D between UAVs m and n mn And the optimal operating distance D of the UAV formation communication link l Establish a probability of being detected by the other party evaluation model As follows:

[0030]

[0031] Among them, f w Represents the communication frequency band coefficient, B c And B sta Respectively represent the current and standard state communication bandwidths of the formation network, P max And P sta Respectively represent the maximum power of the network terminal and the nominal terminal power consumption, r min Represents the closest distance to the other UAV.

[0032] Considering the above factors affecting the communication network efficiency of UAV formations comprehensively, an integrated evaluation model Θ of the communication network efficiency of UAV formations is established. mn It is:

[0033]

[0034] Step 2: Construct a Markov decision model for the Q-Learning algorithm of UAV formations, and train the state-action value function through the Q-Learning algorithm under the constructed Markov decision model, and finally converge to determine the optimal state-action value function. Specifically, Step 2 includes the following steps:

[0035] Construct a Markov decision model for UAV formations based on the components of the Q-Learning algorithm in reinforcement learning. Reinforcement learning consists of a state space S UAVs , an action space A UAVs , a reward R, a policy π, and a state transition probability . These five key elements form a Markov decision process when interacting with the environment. In the embodiments of the present invention, the members of the UAV formation are regarded as nodes in the communication network. Considering that the communication routing process needs to traverse each node, the formation member scale N is defined as the state space S UAVs :

[0036] S UAVs = {S1, S2, …, S N+1}

[0037] Among them, S N+1 represents the virtual leader UAV node.

[0038] Define the nodes adjacent to the current UAV node S i as the elements in the action space A UAVs . Take the elements in the set as the executed action A i . Define the reward R(S i , A i ) obtained by the UAV formation when executing the action A i at the node S i as R i+1 . Define the policy process of the state node selecting an action according to a fixed rule as π(A|S) = P(A UAVs = A|S UAVs = S). Then the Markov property in the process of interacting with the environment can be expressed as:

[0039] P{S′, R i+1 |S i , A i , S i-1 , A i-1,…,S1,A1}=P{S′,R i+1 |S i ,A i}

[0040] Furthermore, to give the optimal policy the maximum reward, the cumulative reward function is defined as:

[0041]

[0042] where γ is a constant between 0 and 1, called the discount factor, which is used to balance the immediate reward and the expected reward. When making each step of the Markov decision, the historical cumulative reward R i a is used to evaluate the quality of the current policy, and then the optimal policy is selected.

[0043] Considering the randomness of the policy during the algorithm learning process, the mathematical expectation of the cumulative reward function R i a is defined as the formation node state value function Vπ(S i ) as:

[0044]

[0045] Considering the execution action A i of the node, the state value function V π (S i ) can be transformed into the state-action value function Q π (S i i,A i ), that is, Q π (S i ,A i )=E π (R i a |S i UAVs =S i ,A i UAVs =A i ), then from the formula we can get:

[0046] Q π (S i ,A i )=E π (R i+1 +γQ π (S′i,A′)|S i UAVs =S i ,A i UAVs =A i )

[0047] Furthermore, the state-action value function can be derived as follows:

[0048]

[0049] Then, through training and learning, the optimal policy π of the algorithm * The corresponding optimal state-action value function Q * (S i , A i ) is:

[0050]

[0051] Solving through the interaction between the Q-Learning algorithm and the environment is to find the optimal state-action value function Q π (S i i, A i ). This process can be divided into two parts: designing the reward matrix of the Q-Learning algorithm and training and updating the reward matrix to obtain the optimal state-action value function.

[0052] Step 3: Based on the distance D between UAV formation members mn and the maximum reachable distance D of the information communication link l,max , establish the maximum available communication topology of the formation. The initial reward matrix Q0 is:

[0053]

[0054] When designing the reward matrix, consider the UAV formation communication effectiveness evaluation model Θ mn established in Step 1, and define the maximum reward when the leading UAV and the following UAV in the UAV formation establish a communication routing link as Θ max . Then, design the reward matrix R(S i , A i ) during the training and learning process of the formation UAVs as:

[0055]

[0056] Among them, (S N+1 , S n ) represents that the virtual leader UAV node S N+1 establishes a communication connection with the node S n , and (S m , S n ) represents that the UAV node S m establishes a connection with the node S n .

[0057] As can be seen from Step 2, when the formation UAV i is in the state S i , if it chooses to execute the action A i, then the state after interacting with the environment is S′, and at the same time, a reward R(S i ,A i ,S′) is generated. According to the state-action value function, the update formula for the reward matrix Q of the formation UAVs during the learning process is defined as follows:

[0058] Q(S′,A′) = (1 - α)Q(S i ,A i ) + α[R(S i ,A i ) + γmaxQ'(S′,A′)]

[0059] In the formula, α is a constant between 0 and 1, called the learning rate, which is used to represent the correlation degree between the current and the previous training results. When updating the matrix Q through the formula, if the value of Q(S′,A′) becomes smaller, it means that the action selected this time is not the optimal action, and the selection of this action will be reduced in the next training.

[0060] Step 4: Design a greedy factor adaptive adjustment mechanism to improve the exploration ability in the early stage and the learning ability in the later stage of the traditional Q-Learning algorithm when interacting with the environment. Specifically, Step 4 includes the following steps:

[0061] Considering that the general ε-greedy strategy uses a fixed probability ε to randomly select the next UAV connected to the current node UAV, which is not sufficient to quickly balance the contradiction between the exploration ability and the exploitation ability of the nodes when constructing the communication topology network of large-scale UAV formations, design a greedy factor function ζ(Iter i ) that changes adaptively according to the number of iterations of the Q-Learning algorithm:

[0062]

[0063] In the formula, Iter i represents the current iteration generation of the algorithm, and are overshoot parameters. The adjustment function ζ(Iter i ) takes a relatively large value at the initial iteration of the algorithm, so that in the initial stage, the UAV nodes can fully explore the connectable neighbor nodes and increase the use of information in the environment; in the later stage of the algorithm iteration, considering that the routing nodes have learned the experience of the optimal node connection, the value of the adjustment function ζ(Iter i ) is reduced compared to the initial iteration, that is, the probability that the UAV routing node randomly selects a connected node is reduced, which can effectively increase the reward in the process of constructing the overall formation routing network. In addition, the adaptive conditional greedy factor can avoid the problem that the UAV nodes cannot learn experience due to the constant greedy factor and finally select a sub-optimal communication routing data link.

[0064] Repeating the above steps will obtain a large amount of data on the connectivity actions and rewards of UAV nodes, enabling the reward matrix Q g to finally converge.

[0065] Step 5: Optimize the formation communication network topology based on the improved Q-Learning algorithm designed, and achieve the dynamic optimization and update of the formation through the network. Specifically, Step 5 includes the following steps:

[0066] The specific process of optimizing the UAV formation communication topology network using the improved Q-Learning algorithm designed by the present invention is as follows:

[0067] S1. Obtain information such as the structure, communication equipment, positions, and data link distance limitations of the UAV formation, and calculate the reward matrix Q of the current state according to the formula in step (3) mn ;

[0068] S2. Model the formation members as routing nodes and abstract them into the structural form of the input of the Q network to construct a Markov decision model for the UAV formation;

[0069] S3. Based on the action space where routing can be established for the UAV formation, regard neighboring nodes as optional actions, select the optional action with the largest reward Q, calculate the immediate reward value of the action according to the formula, and complete the update of the reward Q matrix based on the formula, and then iterate repeatedly until the reward Q matrix converges;

[0070] S4. Select the slave UAV with the farthest physical distance from the leader UAV as the initial node, and infer and construct the optimal action of the routing link according to the converged reward Q matrix to complete the construction of the main communication link of the UAV formation; if there is a slave UAV not in the main communication link, select the node with the shortest physical distance from the slave UAV and connect it to the main communication link to complete the optimization of the UAV formation communication topology network.

[0071] A method for optimizing the communication topology network of multiple UAVs based on improved Q-Learning provided by the embodiment of the present invention first analyzes the influencing factors of the UAV formation communication network, constructs a UAV communication routing evaluation model, designs a greedy factor adaptive adjustment mechanism and a reward matrix based on the routing evaluation model, and applies the Q-Learning algorithm to the optimization of the UAV formation communication topology, forming a method for solving the communication topology optimization of the UAV formation.

[0072] Embodiment 1

[0073] Step 1: Analyzing the process of the UAV formation information interaction network shows that the factors affecting the final communication routing model of the formation mainly include the communication intensity, communication cost, and the probability of being detected by the other party between UAVs, such as Figure 1As shown, a communication intensity evaluation model is established by means of a mathematical function. Communication cost evaluation model and the probability of being detected by the other party evaluation model are as follows:

[0074]

[0075] Among them, D mn is the physical distance between UAVs m and n, υ = 1 is the path dissipation index between UAVs in a barrier-free environment, D l,max = 30m is the maximum reachable distance of the UAV communication link, D l = 20m is the optimal operating distance of the UAV formation communication link, B and P are the communication network bandwidth and terminal power consumption of the UAV formation, f w = 1 represents the communication frequency band coefficient, B c and B sta = 40MHz represent the current and standard state communication bandwidths of the formation network respectively, P max = 4W and P sta = 2W represent the maximum power of the network terminal and the nominal terminal power consumption respectively.

[0076] Taking into account the above factors affecting the performance of the UAV formation communication network, a comprehensive evaluation model Θ of the UAV formation communication network is established mn as:

[0077]

[0078] In the formula, take γ c = [γ 1,c , γ 2,c , γ 3,c = [0.5, 0.2, 0.3].

[0079] Step 2: Set the UAV formation size to 15. Whether a communication route can be initially established between UAV formations is as Figure 2 shown. Based on the state space S UAVs defined by the UAV formation member size, it is:

[0080] S UAVs = {S1, S2, …, S 15 , S 16}

[0081] Among them, S 16It is a virtual leading UAV node. The position information of UAVs 1 to 15 is (0, 60), (15, 75), (30, 90), (52.5, 97.5), (76.5, 87), (102, 78), (31.5, 55.5), (46.5, 75), (97.5, 70.5), (7.5, 52.5), (22.5, 33), (49.5, 27), (63, 51), (85.5, 37.5), (100.5, 49.5), (63, 60), and the positions of the opposing UAVs are (135, -30), (157.5, -15), (150, -27), (150, 27).

[0082] Define the adjacent nodes around the current UAV node S7 as the action space A7 UAVs Element A in the set i , then A7 UAVs ={A2, A8, A 10 , A 11}, where A2 represents the action of selecting the next UAV node as S2. Define the reward R(S i , A i , S′) obtained by the UAV formation when performing the action Ai at the node Si as R i+1 . Define the policy process of selecting actions by the state node according to fixed rules as π(A|S) == P(A UAVs == A|S UAVs == S), then the Markov property in the process of interacting with the environment can be expressed as:

[0083] P{S′, R i+1 |S i , A i , S i-1 , A i-1 , …, S1, A1} == P{S′, R i+1 |S i , A i}

[0084] To give the optimal policy the maximum reward, define the cumulative return function as:

[0085]

[0086] Due to the randomness of the policy during the learning process, define the mathematical expectation of R i a as the node state value function V π (S i ) as:

[0087]

[0088] Furthermore, consider the action parameter A of the formula node i , then the state value function V π (S i ) can be transformed into the state-action value function Q π (S i i,A i ), that is, Q π (S i ,A i ) = E π (R i a |S i UAVs = S i ,A i UAVs = A i ). Then, from the formula, we can get:

[0089]

[0090] Then, through training and learning, the optimal policy π * of the algorithm corresponds to the optimal state-action value function Q * (S i ,A i ) as:

[0091]

[0092] Step 3: Based on Figure 2 the connectivity of the UAV formation member nodes shown, establish the initial reward matrix Q0 as:

[0093]

[0094] Design the immediate reward function R(S i ,A i ) as:

[0095]

[0096] where, Θ max = 1000 represents the maximum reward when the leader UAV and the follower UAV in the UAV formation establish a communication routing link. (S 16 , S8) or (S 16 , S 13 ) represents that the virtual leader node S 16 establishes a communication link with the UAV node S8 or S 13 . (S m , S n ), n = 1, 2, … 15, m ≠ 0 represents that the node S m establishes a communication link with the node S n . Then the immediate reward function R is:

[0097]

[0098] As can be seen from step 2, the state of formation UAV i is S i When, if action parameter A is selected i , then the state after interacting with the environment is S′, and at the same time, a reward R(S i , A i , S′) is generated. The interaction training process is as Figure 3 shown. According to the state-action value function, the update formula of the reward matrix Q of the formation UAVs during the learning process is defined as follows:

[0099] Q(S′, A′) = (1 - α)Q(S i , A i ) + α[R(S i , A i ) + γmaxQ'(S′, A′)]

[0100] In the formula, α = 0.1 is the learning rate, which is used to represent the correlation degree between the current and the previous training results, and γ = 0.9 is the discount factor.

[0101] Step 4: Considering that the general ε-greedy strategy uses a fixed probability ε to randomly select the next UAV connected to the current node UAV, which is not sufficient to quickly balance the contradiction between the exploration ability and exploitation ability of nodes when constructing the communication topology network of large-scale UAV formations, design a greedy factor function ζ(Iter i ) that adapts to the number of iterations of the Q-Learning algorithm:

[0102]

[0103] In the formula, Iter i represents the current iteration generation of the algorithm, and are overshoot parameters.

[0104] Step 5: Optimize the formation communication network topology based on the designed improved Q-Learning algorithm to achieve the dynamic optimization and update of the formation through the network, including the following steps:

[0105] S1. Obtain information such as the structure, communication equipment, position, and data link distance limit of the UAV formation, and calculate the reward matrix Q of the current state according to the formula in step (3) mn ;

[0106] S2. As shown in the formula ~, model the formation members as routing nodes and abstract them into the structural form of the input of the Q network to construct a Markov decision model for the UAV formation;

[0107] S3. Based on the action space where routes can be established by the UAV formation, consider adjacent nodes as optional actions, and select the optional action with the maximum return Q. As Figure 4 shown, Figure 4 in (a) is the training iteration process diagram of UAV7, Figure 4 in (b) is the training iteration process diagram of UAV8, Figure 4 in (c) is the training iteration process diagram of UAV10. Calculate the immediate return reward value of the action according to the formula and complete the update of the return Q matrix based on the formula, and then iterate repeatedly until the return Q matrix converges. The converged Q matrix is as follows:

[0108]

[0109] S4. As Figure 5 shown, select the slave UAV with the farthest physical distance from the leader UAV as the initial node, and infer and construct the optimal action of the routing link according to the converged return Q matrix to complete the construction of the main communication link of the UAV formation; as Figure 6 shown, if there is a slave UAV not in the main communication link, select the node with the shortest physical distance from the slave UAV and connect it to the main communication link to complete the optimization of the communication topology network of the UAV formation.

[0110] The above content details the mathematical principles and specific steps of the present invention. Although the present invention gives embodiments to illustrate the specific process and mathematical principles of the present invention, it should be noted that the present invention is not limited to the above embodiments. Under the condition of conforming to the design concept of the present invention, the present invention can obtain the final results for various changed and improved embodiments. Therefore, it should be understood that the embodiments here are only used to illustrate the present invention and are not limited to the present invention.

Claims

1. A method for optimizing a multi-UAV communication topology network based on improved Q-Learning, characterized in that, It includes the following steps: According to the process of the UAV formation information interaction network, analyze the influencing factors of the UAV formation communication routing model to obtain the corresponding evaluation models for each influencing factor, thereby constructing an evaluation model for the formation communication network effectiveness; Regarding the members of the UAV formation as nodes in a communication network, considering that the communication routing process needs to traverse each node, the formation member scale N is defined as the state space S UAVs And define the leader UAV node, and the state space S UAVs The nodes adjacent to the current UAV node in i Are defined as performing action A i To form the action space A UAVs , the harvest of performing action A i Is used as the reward R i+1 , the process of the policy for the state node to select actions according to fixed rules is defined as the policy π(A|S), and the Markov decision model is obtained; Define the cumulative reward function according to the optimal strategy of maximizing the reward; according to the randomness of the algorithm learning process, define the mathematical expectation of the cumulative reward function as the state value function of the formation nodes; combine the executed action A i Convert the state value function into a state-action value function, and under the constructed Markov decision model, train it through the Q-Learning algorithm, and finally converge to determine the optimal state-action value function; Obtain information such as the structure, communication equipment, position, and data link distance limit of the UAV formation. According to the formation communication network effectiveness evaluation model, obtain the reward matrix design R(S i ,A i ) when the maximum return of the formation communication network effectiveness evaluation model is obtained. Combine the optimal state-action value function to obtain the UAV reward matrix of the formation; Considering the contradiction between node exploration and development in the construction of the large-scale UAV formation communication topology network, construct a greedy factor function that adaptively changes according to the number of iterations of the Q-Learning algorithm, so that the formation UAV reward matrix converges. Select the node with the farthest physical distance from the leading UAV node as the initial node, and construct the optimal action of the routing link according to the converged formation UAV reward matrix, and finally obtain the main communication link of the UAV formation.

2. The multi-UAV communication topology network optimization method based on improved Q-Learning according to claim 1, wherein The evaluation models corresponding to the respective influencing factors include a communication intensity evaluation model, a communication cost evaluation model, and a probability of being detected by the other party evaluation model.

3. The multi-UAV communication topology network optimization method based on improved Q-Learning according to claim 2, characterized in that, Establish the communication intensity evaluation model based on the physical distance between UAVs, the maximum reachable distance of the UAV communication link, and the path dissipation index between UAVs; establish the communication cost evaluation model based on the optimal operating distance of the UAV formation communication link, the physical distance between UAVs, the maximum reachable distance of the UAV communication link, and the path dissipation index between UAVs; establish the probability of being detected by the other party evaluation model based on the bandwidth of the UAV formation communication network, the terminal power consumption, the physical distance between UAVs, and the optimal operating distance of the UAV formation communication link.

4. The multi-UAV communication topology network optimization method based on improved Q-Learning according to claim 1, characterized in that, If there is an initial node outside the main communication link, select the node with the shortest physical distance from the slave and connect it to the main communication link to optimize the main communication link of the UAV formation.

Citation Information

Patent Citations

  • Distributed formation method of unmanned aerial vehicle cluster based on reinforcement learning

    CN110007688A

  • Unmanned aerial vehicle ad hoc network adaptive routing method based on Q-Learning

    CN114449608A

  • Decentralized policy gradient descent and ascent for safe multi-agent reinforcement learning

    US20230113168A1