Multi-unmanned aerial vehicle cooperation strategy learning method and device based on time sequence hierarchical decision-making
By adopting a strategy learning method based on timing hierarchical decision-making in multi-UAV collaborative tasks, a macro action space is constructed and a hierarchical decision-making mechanism is introduced, the problem of inefficiency exploration under sparse reward conditions is solved, and efficient collaboration strategy optimization and performance improvement are achieved.
Patent Information
- Application Number
- CN202510626050.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In multi-UAV collaborative tasks, under the conditions of sparse rewards, the exploration efficiency of the agent in the huge collaborative decision-making space is inefficient, and the strategy convergence is difficult, resulting in the core bottleneck of technology implementation.
The multi-UAV collaborative strategy learning method based on timing hierarchical decision-making is adopted to construct the macro action space of the drone through timing abstraction, introduce a hierarchical decision-making mechanism of macro strategies and micro strategies, and gradually replace the hierarchical strategy through parallel strategy optimization mechanisms.
The optimization efficiency of multi-drone collaboration strategy has been improved, the performance of drone collaboration strategy has been effectively improved, and the problems of high complexity of sequential decision-making and large decision-making space have been overcome.
Smart Images

Figure CN120146145A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multi-UAV collaborative strategy learning methods, and particularly to a multi-UAV collaborative strategy learning method and device based on temporal hierarchical decision-making. Background Art
[0002] In recent years, multi-UAV collaborative tasks (such as area search, formation control, disaster rescue, etc.) have become a research hotspot in the field of intelligent unmanned systems due to their flexibility and efficiency. To achieve autonomous collaboration in complex environments, multi-agent reinforcement learning has been widely applied to UAV collaborative strategy learning. However, in actual scenarios, task objectives often correspond to extremely sparse external reward signals (such as only giving a one-time reward when the task is completed), resulting in low exploration efficiency of agents in the huge collaborative decision-making space and difficult policy convergence, which has become the core bottleneck restricting the implementation of the technology.
[0003] Under sparse reward conditions, intrinsic rewards drive agents to search for positive (external) rewards in the state space. Essentially, it drives agents to search for effective decision sequences that can reach positive reward states in the collaborative sequential decision-making space. Intrinsic rewards provide an effective decision sequence search strategy, but do not optimize the sequential decision-making space itself that needs to be searched. The size of the collaborative sequential decision-making space grows exponentially with the increase in the number of agents and the length of the decision sequence. In long-term collaborative tasks, agents need to complete tasks through continuously coordinated action sequences, which poses a huge challenge to the search and utilization of sparse positive rewards.
[0004] In addition, multi-UAV collaboration needs to solve practical constraints such as distributed decision-making under limited communication, heterogeneous ability coordination, and dynamic environment disturbance, which further exacerbates the difficulty of policy learning under sparse reward conditions. Existing methods often have problems such as divergent exploration directions and disordered collaborative behaviors in long-term tasks, resulting in the task success rate and algorithm robustness being difficult to meet actual requirements. Summary of the Invention
[0005] Based on this, it is necessary to provide a multi-UAV collaborative strategy learning method and device based on temporal hierarchical decision-making for the above technical problems.
[0006] A multi-UAV collaborative strategy learning method based on temporal hierarchical decision-making, the method includes: Using a temporal hierarchical decision-making mechanism to perform temporal abstraction on the original sequential decision sequence, and constructing a UAV macro action space capable of reaching a positive reward state; each macro action consists of original actions.
[0007] Train the pre-constructed hierarchical policy according to the current observation, macro-action, and primitive action; the hierarchical policy includes a macro-policy and a micro-policy; the macro-policy is used to select a macro-action according to the current observation, where each drone selects a macro-action every steps. The micro-policy is used to select a primitive action at each step according to the current observation and the most recently selected macro-action.
[0008] The drone performs hierarchical decision-making search in the macro-action space through the hierarchical policy and stores the experience samples generated by the drone's hierarchical decision-making in the experience cache.
[0009] After the drone searches for a positive extrinsic reward using the hierarchical policy, it offline trains the pre-constructed parallel policy using the experience samples containing primitive actions sampled from the experience cache, and then gradually replaces the hierarchical policy with the parallel policy to obtain a multi-drone cooperation policy; the parallel policy is used to output a primitive action according to the current observation, and only the extrinsic reward is used during training.
[0010] In one embodiment, a temporal hierarchical decision-making mechanism is used to perform temporal abstraction on the original sequential decision-making sequence to construct a drone macro-action space capable of reaching the positive reward state, including: Based on the analysis and induction of known successful cooperation behaviors, extract the primitive action sequences that coexist in multiple cooperation modes as macro-actions.
[0011] Decompose the multi-drone cooperation action sequence of the successful task into a macro-action sequence, and abstract a reasonable macro-action space from multiple spatio-temporal coordination modes.
[0012] A multi-drone cooperation policy learning device based on temporal hierarchical decision-making, the device includes: A drone macro-action space construction module, which is used to perform temporal abstraction on the original sequential decision-making sequence using a temporal hierarchical decision-making mechanism to construct a drone macro-action space capable of reaching the positive reward state; each macro-action consists of primitive actions.
[0013] A hierarchical policy training module, which is used to train the pre-constructed hierarchical policy according to the current observation, macro-action, and primitive action; the hierarchical policy includes a macro-policy and a micro-policy; the macro-policy is used to select a macro-action according to the current observation, where each drone selects a macro-action every steps. The micro-policy is used to select a primitive action at each step according to the current observation and the most recently selected macro-action.
[0014] A hierarchical decision-making module, which is used for the drone to perform hierarchical decision-making search in the macro-action space through the hierarchical policy and store the experience samples generated by the drone's hierarchical decision-making in the experience cache.
[0015] The multi-UAV cooperation strategy determination module is used to, after the UAVs search for positive external rewards using the hierarchical strategy, offline train a pre-constructed parallel strategy with experience samples containing original actions sampled from the experience cache, and then gradually replace the hierarchical strategy with the parallel strategy to obtain the multi-UAV cooperation strategy; the parallel strategy is used to output the original actions according to the current observation, and only the external rewards are used during training.
[0016] The above multi-UAV cooperation strategy learning method and device based on temporal hierarchical decision-making overcome the problems of high complexity and large decision space in multi-UAV sequential decision-making by performing temporal abstraction on the original sequential decision sequence, thereby improving the efficiency of multi-UAV cooperation strategy optimization. Temporal abstraction is performed on the original sequential decision sequence to construct a UAV macro-action space capable of reaching the positive reward state, and a temporal hierarchical decision-making method including a macro strategy and a micro strategy is introduced. The macro strategy makes top-level decisions based on macro actions and is updated using the intrinsic reward of equivalent state novelty value combined with the external reward. The micro strategy makes bottom-level decisions based on the original actions and is generated based on rule definition or pre-training. To overcome the weakening of the cooperation strategy performance caused by hierarchical decision-making, a strategy parallel optimization mechanism based on experience reuse is designed. This method improves the strategy optimization efficiency and effectively enhances the performance of the UAV cooperation strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of a multi-UAV cooperation strategy learning method based on temporal hierarchical decision-making in an embodiment; Figure 2 It is a schematic diagram of a multi-UAV task scenario in another embodiment, where Figure 2 (a) is a schematic diagram of scenario 2UAV-1ADS-2, Figure 2 (b) is a schematic diagram of scenario 3UAV-1ADS, Figure 2 (c) is a schematic diagram of scenario 3UAV-2ADS, Figure 2 (d) is a schematic diagram of scenario 5UAV-2ADS; Figure 3 It is a curve graph of cumulative external rewards and task success rates in scenario 2UAV-1ADS-2 in another embodiment, where Figure 3 (a) is a curve graph of cumulative external rewards, Figure 3 (b) is a curve graph of task success rates; Figure 4 It is a curve graph of cumulative external rewards and task success rates in scenario 3UAV-1ADS in another embodiment, where Figure 4 (a) is a curve graph of cumulative external rewards, Figure 4 (b) is a curve graph of task success rates; Figure 5 For the cumulative extrinsic reward and mission success rate curve graphs of Scenario 3 UAV-2 ADS in another embodiment, where Figure 5 (a) is the cumulative extrinsic reward curve graph, Figure 5 (b) is the mission success rate curve graph; Figure 6 For the cumulative extrinsic reward and mission success rate curve graphs of Scenario 5 UAV-2 ADS in another embodiment, where Figure 6 (a) is the cumulative extrinsic reward curve graph, Figure 6 (b) is the mission success rate curve graph; Figure 7 For the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in another embodiment, where Figure 7 (a) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 2 UAV-1 ADS-2, Figure 7 (b) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 3 UAV-1 ADS, Figure 7 (c) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 3 UAV-2 ADS, Figure 7 (d) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 5 UAV-2 ADS. Detailed implementation manners
[0018] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] In one embodiment, as Figure 1 shown, a multi-UAV cooperative strategy learning method based on hierarchical decision-making over time is provided, and the method includes the following steps: Step 100: Use a hierarchical decision-making mechanism over time to perform temporal abstraction on the original sequential decision-making sequence, and construct a macro action space of UAVs capable of reaching the positive reward state; each macro action consists of original actions.
[0020] Specifically, model the multi-UAV system as a distributed partially observable Markov decision process (Dec-POMDP) model. The size of the multi-UAV cooperative sequential decision-making space is , where is the size of the individual action space of the UAV, is the number of UAVs, is the length of the decision-making sequence, that is, the maximum number of decision-making steps per round.
[0021] Multi - UAV collaborative policy learning aims to search and optimize an effective action sequence that can reach the positive - reward state as soon as possible from the collaborative sequential decision - making space where is the joint action of multiple UAVs, is the number of state transitions from the initial state to the positive - reward state
[0022] For long - period collaborative tasks, the state containing the positive reward is far from the initial state, that is, is large, and the maximum decision - making step is usually set to a large value, which leads to two problems affecting policy optimization: One is that the size of the multi - UAV collaborative sequential decision - making space increases exponentially with the increase of the decision - sequence length . Driven only by the novelty value of equivalent states, the search efficiency for effective decision sequences still needs to be improved; The other is that the feedback delay of the positive - reward signal is long. When combined with the time - discount factor for policy update, the guiding effect decreases as increases.
[0023] Introduce a temporal hierarchical decision - making mechanism to reduce the collaborative sequential decision - making space, lower the decision - making complexity, and improve the efficiency of multi - UAV collaborative policy optimization. The temporal hierarchical decision - making mechanism performs temporal abstraction (Temporal Abstract) on the original decision sequence and constructs a macro - action space of UAVs Each macro - action is composed of original actions, that is, , is an integer greater than zero.
[0024] Step 102: Train the pre - constructed hierarchical policy according to the current observation, macro - action, and original action; the hierarchical policy includes a macro - policy and a micro - policy; the macro - policy is used to select a macro - action according to the current observation, where each UAV selects a macro - action every steps; the micro - policy is used to select an original action at each step according to the current observation and the most recently selected macro - action.
[0025] Specifically, based on the macroscopic action space, this application proposes a multi-UAV cooperative policy learning method based on temporal hierarchical decision-making (Temporal-Hierarchical decision making based Multi-UAV cooperative Policy Learning method, TH-MUPL), trains the UAV hierarchical cooperation policy, and further optimizes it through parallel policy learning based on offline training. The macroscopic policy of the UAV and the microscopic policy are the basis for hierarchical decision-making. The input of the macroscopic policy is the current moment observation , and the output is the macroscopic action . The macroscopic action selection is re-performed every steps. The microscopic policy receives the most recently selected macroscopic action and the current moment observation as inputs, and the output is the primitive action .
[0026] In the multi-UAV cooperative policy learning method based on temporal hierarchy, the macroscopic policy is responsible for determining the conversion timing of each macroscopic action according to the changing situation, and the microscopic policy is responsible for completing the currently specified macroscopic action by selecting an appropriate primitive action , thereby decoupling the collaborative decision-making behavior in terms of time sequence and reducing the complexity of decision-making. In addition to the macroscopic policy and the microscopic policy, the UAV also includes a parallel policy . The input of the parallel policy is the current moment observation , and the output is the primitive action .
[0027] To improve the stability and efficiency of training, the macroscopic policy and the microscopic policy of the UAV in this method are optimized independently.
[0028] Step 104: The UAV performs hierarchical decision-making search in the macroscopic action space through the hierarchical policy, and stores the experience samples generated by the UAV hierarchical decision-making in the experience cache.
[0029] Step 106: After the UAV searches for a positive external reward using the hierarchical policy, it offline trains the pre-constructed parallel policy using the experience samples containing primitive actions sampled from the experience cache, and then gradually replaces the hierarchical policy with the parallel policy to obtain the multi-UAV cooperative policy; the parallel policy is used to output the primitive action according to the current observation, and only the external reward is used during training.
[0030] Specifically, since the defined macro-action is a primitive action sequence of a fixed length, under the time-series hierarchical decision-making mechanism, the macro-actions switched at fixed intervals result in a decrease in the flexibility of state transition, and the performance of the multi-UAV hierarchical cooperation strategy generated by training is accordingly reduced. Therefore, a parallel strategy offline training mechanism is proposed. After obtaining a positive external reward, an experience sample containing primitive actions is used to offline train the parallel strategy , and gradually replace the hierarchical strategy, thereby further improving the optimization of the multi-UAV cooperation strategy.
[0031] In the above multi-UAV cooperation strategy learning method based on time-series hierarchical decision-making, the method overcomes the problems of high complexity and large decision space in multi-UAV sequential decision-making by performing time-series abstraction on the original sequential decision sequence, thereby improving the optimization efficiency of the multi-UAV cooperation strategy. Time-series abstraction is performed on the original sequential decision sequence to construct a UAV macro-action space capable of reaching the positive reward state. A time-series hierarchical decision-making method including a macro-strategy and a micro-strategy is introduced. The macro-strategy makes top-level decisions based on macro-actions and is updated using the intrinsic reward of equivalent state novelty combined with the external reward. The micro-strategy makes bottom-level decisions based on primitive actions and is generated by rule definition or pre-training; to overcome the weakening of the cooperation strategy performance caused by hierarchical decision-making, a strategy parallel optimization mechanism based on experience reuse is designed. This method improves the strategy optimization efficiency and effectively enhances the performance of the UAV cooperation strategy.
[0032] In one embodiment, step 100 includes: analyzing and summarizing based on known successful cooperation behaviors, and extracting the primitive action sequences common to multiple cooperation modes as macro-actions; decomposing the multi-UAV cooperation action sequence of the successful task into macro-action sequences, and abstracting a reasonable macro-action space from multiple spatio-temporal cooperation modes; where the macro-action space is: ;
[0033] where is the macro-action space.
[0034] Specifically, under the time-series hierarchical decision-making mechanism, the decision-making process of the UAV changes from: to hierarchical decision-making: ;
[0035] Each UAV selects a macro-action every steps; at each step, according to the current observation and the most recently selected macro-action , a primitive action is selected. For multiple UAVs The size of the macro decision space for sequential decision-making is , where is the size of the individual macro action space. Therefore, if the underlying decision is a predefined deterministic mapping and the macro action satisfies , the compression of the sequential decision space can be achieved, that is, . The construction of the macro action space needs to meet the reachability requirement from the initial state to the positive reward state, that is, for any initial state , there exists a macro decision sequence , , , corresponding to the micro decision sequence , satisfying , , , such that the state can be transferred from to the positive reward state , that is, there exists a state transition sequence , satisfying: ;
[0036] where, represent the transition probabilities of the two states respectively.
[0037] To construct a macro action space that meets the above requirements, the method adopted is to analyze and summarize based on known successful cooperation behaviors, and extract the necessary original action sequences coexisting in multiple cooperation modes as macro actions. Decompose the multi-UAV cooperation action sequence of successful penetration and assault into a macro action sequence, and abstract a reasonable macro action space from multiple spatio-temporal cooperation modes.
[0038] Abstract the individual behaviors during the process of multi-UAV cooperation in performing tasks as: approaching the target, moving away from the target, flying clockwise around the target, flying counterclockwise around the target, and hovering in place. The above individual behaviors of the UAVs need to be executed in a specific time sequence under certain spatio-temporal constraints to exert the cooperation effect. Therefore, define a macro action space based on temporal abstraction .
[0039] In one embodiment, step 102 includes: the micro policy is defined by means of pre-training or rule-based definition, and the macro policy is pre-trained by using the reinforcement learning method; input the current observation and the macro action into the micro policy to output the original action; use the ACER method to update the macro policy of the UAV, and the macro policy includes the UAV macro value network and the macro policy network.
[0040] In one of the embodiments, the current observation and the macro action are input into the micro policy, and the original action is output, including: when the macro action selects to approach the target, the micro policy expression of the UAV is: ; ; wherein, is the micro policy of the UAV when the macro action is to approach the target, represents the current observation of UAV i, represents approaching the target, represents the UAV the unit vector of the heading angle after selecting each original action, represents from the UAV the unit vector pointing to the protected target, represents the k th original action; represents the original action space, respectively represent the position coordinates of the protected target.
[0041] When the macro action selects to approach the threat k , the micro policy expression of the UAV is: ; ; wherein, represents the micro policy of the UAV when the macro action is to approach threat k, represents the UAV the unit vector pointing to the threat area , represents approaching threat k, respectively represent the threat k 's position coordinates, respectively represent the UAV i at t time's position coordinates.
[0042] When the macro action selects to move away from the threat area , the micro policy expression of the UAV is: ; wherein, represents the micro policy of the UAV when the macro action is to move away from the threat area k , represents moving away from the threat area .
[0043] When the macro action selects to fly clockwise around the closest threat area, the micro policy expression of the UAV is: ; ;
[0044] Among them, represents the UAV micro strategy when the macro action is to fly around the nearest threat area clockwise, represents flying around the nearest threat area clockwise, represents the unit vector orthogonal to the vector pointing from the UAV to the nearest threat area, and this vector points in the clockwise direction, represents the position of the nearest threat area to the UAV .
[0045] When the macro action selects to fly around the nearest threat area counterclockwise, the UAV micro strategy expression is: ; Among them, represents the UAV micro strategy when the macro action is to fly around the nearest threat area counterclockwise; represents flying around the nearest threat area counterclockwise; When the macro action selects to hover in place, the UAV micro strategy expression is: ; Among them, represents the UAV micro strategy when the macro action is to hover in place; is the hovering center, and the hovering center is based on the moment when the macro action is selected. At that time, the position point is 0.5 km extended outward from the connection line between the nearest threat area and the UAV at that time, represents the reference moment of the macro action.
[0046] Hovering center is: ; Among them, is the predicted position of the UAV at the next moment after executing the original action , represents the angle between the flight of the UAV i at this moment and the horizontal plane, represents the time interval, represents the macro value network parameter.
[0047] In one embodiment, the update method of the UAV macro value network is: ; ; ;
[0048] Among them, represents the parameters of the UAV's macroscopic value network, represents the hyperparameters for updating the parameters of the UAV's macroscopic value network, represents the gradient of the mean squared error of the macroscopic value network, represents the value estimation of the macroscopic policy based on the Retrace method, represents the UAV i current observation, represents the UAV i macroscopic action, represents the reward, represents the truncated importance sampling coefficient value at time , represents the value network, represents the total number of threat areas, represents the hyperparameter, represents the value function of this observation, respectively represent the UAV at this moment i observation and macroscopic action.
[0049] In one embodiment, the update method of the UAV's macroscopic policy network is as follows: ; Among them, represents the parameters of the UAV's macroscopic policy network, represents the hyperparameters for updating the parameters of the UAV's macroscopic policy network, represents the ACER policy gradient, represents the total number of threat areas, represents the value function of this observation, respectively represent t the two different truncated importance sampling coefficient values at time represents the UAV's macroscopic value network, represents the UAV's macroscopic policy network, represents the i th macroscopic action, represents the UAV i current observation, represents the value estimation of the macroscopic policy based on the Retrace method.
[0050] For the sake of simplicity in the text, part is and are respectively used to replace the UAV's macroscopic policy network and the UAV's macroscopic value network 。
[0051] In one embodiment, a concurrent experience replay mechanism is adopted to store and extract UAV experience samples; the concurrent experience replay mechanism includes: the experience cache capacities of each UAV are equal, and the storage of experience samples consistently follows the first-in-first-out principle to ensure that each experience cache is located at the same index position The experience tuples are from the same state transition process; when sampling the experience cache, the indexes of the batch samples sampled by all UAVs are the same.
[0052] Specifically, the microscopic policy is pre-trained or defined based on rules, and then the macroscopic policy is learned based on the stable and effective microscopic policy, so as to continuously optimize the switching timing of each macroscopic action. Since the set macroscopic actions of the UAVs are relatively simple and there is no complex decision-making judgment, the microscopic policy is defined in a rule-based manner 。When the macroscopic actions are set to be more complex and difficult to be completed by a rule-based method, the reinforcement learning method can be used for pre-training. The original individual action space of the UAV: ; wherein, represents that the angular acceleration of the UAV heading is successively adjusted to , and the speed is successively adjusted to .
[0053] The microscopic policy is a deterministic policy, that is, according to the current observation and the macroscopic action outputs a specific action , rather than the probability distribution of the action. Therefore, the microscopic policy is expressed as .
[0054] When the macroscopic action selects "approach the target", that is, , the microscopic policy of the UAV is as shown in the expression of the microscopic policy of the UAV when the macroscopic action selects to approach the target.
[0055] When the macroscopic action selects to approach threat k, that is, , the microscopic policy of the UAV is as shown in the expression of the microscopic policy of the UAV when the macroscopic action selects to approach the threat.
[0056] When the macroscopic action selects to move away from threat area k, that is, , the microscopic policy of the UAV is as shown in the expression of the microscopic policy of the UAV when the macroscopic action selects to move away from the threat area.
[0057] When the macroscopic action selects "fly clockwise around the closest threat area", that is, When the macro-action selects to fly clockwise around the nearest threat area as described above, the micro-strategy of the UAV is as shown in the micro-strategy expression of the UAV.
[0058] When the macro-action selects "fly counterclockwise around the nearest threat area", that is When, the micro-strategy of the UAV is as shown in the micro-strategy expression of the UAV when the macro-action selects to fly counterclockwise around the nearest threat area.
[0059] When the macro-action selects "hover in place", that is When, the micro-strategy of the UAV is as shown in the micro-strategy expression of the UAV when the macro-action selects to hover in place.
[0060] The micro-strategy of the UAV selects the original action that makes the predicted position of the UAV the closest to the hovering center by the hovering radius of 0.5 km . The macro-strategy of the UAV is updated using the ACER method, which involves the macro-value network of the UAV and the macro-strategy network . .
[0061] The parameters of the macro-value network are updated by minimizing the mean square error as follows: ; where represents the value estimation of the macro-strategy based on the Retrace method, , the reward used to update the macro-value and the policy function is the sum of the rewards generated by the original action sequence corresponding to the macro-action of the UAV, , the reward includes the extrinsic reward and the intrinsic reward of the equivalent state novelty value , respectively represent t the values of two different truncation importance sampling coefficient values at time , is the historical macro-strategy of the UAV when it selects the macro-action according to the observation .
[0062] The parameters of the UAV macro-strategy network are updated based on the ACER policy gradient as follows: ; During the multi-UAV collaborative policy learning process, after the execution of the macroscopic action, the generated experience tuples are stored in the macroscopic experience cache . After each decision-making round, randomly sample past experience samples from the experience cache to train the current policy. Given the UAV sampled from the macroscopic experience cache i historical decision-making trajectory is as follows: ; Use the historical decision-making trajectory to calculate the gradient of the macroscopic value network parameters; the value network parameters and the policy network parameters are both updated based on the gradient.
[0063] In one embodiment, after the UAV searches for a positive external reward using the hierarchical policy, use the experience samples containing the original actions in the cache to offline train the pre-constructed parallel policy, including: after the UAV stably searches for a positive external reward through the hierarchical policy, start using the offline experience samples to train the parallel policy, and optimize the parallel policy through D3QN; the parallel policy includes an action value network and a target value network; the action value network and the target value network have the same structure but different parameters; the update method of the action value network parameters is: ; ; ; ; where represents the action value network parameters, represents the action value network, represents the loss function, is the external reward, represents the value, represents the learning rate, represents t the value of the first truncation importance sampling coefficient at time represents the next observation, , respectively represent the UAV i current observation and the original action, represents the advantage function network parameters, represents the advantage function network, represents the state value function network parameters, Represents the state value function network.
[0064] The update method of the target value network parameters is as follows: ; Among them, represents the target value network parameters, represents the learning rate.
[0065] In one embodiment, a parallel strategy is gradually used to replace the hierarchical strategy, including: adjusting the strategy replacement process through the hyperparameter , , is a positive integer; after the positive reward is searched, the parallel strategy starts to be trained; when , start making decisions using the parallel strategy instead of the hierarchical strategy, and record the decision round number as ; among them represents the success rate of the drone completing the task through the hierarchical cooperation strategy; after transition rounds, the parallel strategy completely replaces the hierarchical strategy; the parallel strategy of the drone does not show in an independent parameterized form, but is defined according to the action value function; all drones are based on method for action selection: ; Among them , the drone selects the action corresponding to the maximum of the action value function with a probability of , and randomly selects all candidate actions with a probability of , usually takes a smaller value between 0 and 1, represents the original action space, represents the action value network, k represents the action sequence number, represents the action value network parameters, , respectively represent the current observation and the original action of the drone i , represents the k th original action, represents the parallel strategy of the drone i .
[0066] Specifically, introducing the time-series hierarchical decision-making mechanism effectively reduces the collaborative sequential decision space. By setting a reasonable macro-action space, redundant decision branches in long-cycle collaborative tasks are avoided, and the learning efficiency of collaborative strategies under sparse reward conditions is effectively improved. However, the hierarchical strategy based on the macro-action space has the following two limitations: First, the macro-action is an original action sequence of fixed length, which switches only at certain intervals, and all drones switch simultaneously. Second, the macro-action space is artificially defined. Although it is generated based on the analysis of the successful collaborative behaviors spontaneously emerged by drones in previous work, it is difficult to accurately split the optimal collaborative mode into original action sequences of fixed length. Therefore, although the setting of macro-actions reduces the decision complexity, the setting of fixed and limited action sequences also reduces the flexibility of state transition, resulting in redundant actions when the trained hierarchical collaborative strategy completes tasks, and the strategy performance needs to be improved.
[0067] To solve this problem, a parallel policy optimization mechanism based on offline training is proposed. That is, when the drone searches for positive external rewards through the hierarchical strategy, the generated experience samples are used to offline train the parallel policy. . Different from the macro policy, the parallel policy outputs the original action according to the current observation and only uses the external reward during training. Considering that the drone selects actions through and hierarchical selection, and the micro policy is defined based on rules, the probability distribution of each action in the historical action strategy is difficult to directly obtain and record. Therefore, the offline learning method D3QN (Dueling Double DeepQ-Network) that does not require the use of is selected to optimize the parallel policy. Each drone maintains its own value network . Based on the Dueling structure, the action value function consists of two parts: the state value and the advantage function: ; When the drone stably searches for positive external rewards through the hierarchical strategy, the parallel policy (i.e., the D3QN network ) starts to be trained using offline experience samples. The experience samples generated by the drone's hierarchical decision-making are stored in the experience cache . To alleviate the instability caused by the simultaneous update of multi-drone strategies, the concurrent experience replay mechanism is referred to store and extract drone experiences: (1) The experience cache With equal capacity, the storage of empirical samples consistently follows the "first in, first out" principle to ensure that each empirical cache is located at the same index position. of the empirical tuple originates from the same state transition process; (2)When sampling from the empirical cache, the indices of the batch samples obtained by all drones are the same, , that is . After each episode ends, samples are obtained by sampling from the empirical cache to , and the parameters of the action-value network are optimized by minimizing the loss function : . .
[0068] To avoid the "overestimation" phenomenon, when estimating the value using the D3QN method based on the temporal difference method , a target value network with the same structure as the value network but with a lag in parameter update is used : .
[0069] The update methods of the action-value network parameters and the target value network parameters are as follows: ; ; .
[0070] As the parallel policy is continuously optimized, it gradually replaces the macro policy and the micro policy. Specifically, the hyperparameter is used to adjust the policy replacement process, , where is a positive integer,
[0071] After the positive reward is found, the parallel policy starts to be trained; when , the parallel policy starts to be used instead of the hierarchical policy for decision-making, and the number of decision-making episodes is recorded as ; after transition episodes, the parallel policy completely replaces the hierarchical policy. The parallel policy of the drone is not presented in an independent parameterized form but is defined according to the value function . All drones make action selections based on method.
[0072] It should be understood that although Figure 1The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0073] In a verification embodiment, with the background that a UAV task unit arrives at the periphery of the task area and identifies protected targets in multiple threat areas, the following four scenarios are set for algorithm verification: 2UAV-1ADS-2, 3UAV-1ADS, 3UAV-2ADS, and 5UAV-2ADS. The schematic diagram of the multi-UAV task scenario is as Figure 2 shown, where Figure 2 (a) is the schematic diagram of scenario 2UAV-1ADS-2, Figure 2 (b) is the schematic diagram of scenario 3UAV-1ADS, Figure 2 (c) is the schematic diagram of scenario 3UAV-2ADS, Figure 2 (d) is the schematic diagram of scenario 5UAV-2ADS. The scenario name includes the number of UAVs and the number of threats in the task area. The parameter settings of the multi-UAV task scenario are shown in Table 1.
[0074] Table 1 Parameter settings of multi-UAV task scenario
[0075] Three algorithms are tested and compared. The three algorithms are respectively: A multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making, TH-MUPL; A multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making without a parallel strategy optimization mechanism, TH-MUPL-np; A multi-UAV cooperative strategy learning method that only utilizes extrinsic rewards, No-Intrinsic-Reward.
[0076] The UAV macro policy network, from the input layer to the output layer, consists of a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, a fully connected layer (with 6 layers), and a softmax function in sequence. The UAV macro value network, from the input layer to the output layer, consists of a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 6 layers) in sequence. The UAV parallel value network, from the input layer to the output layer, consists of a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 7 layers).
[0077] The cumulative extrinsic reward and task success rate curve of Scenario 2 UAV-1 ADS-2 is as follows Figure 3 shown, where Figure 3 (a) is the cumulative extrinsic reward curve Figure 3 (b) is the task success rate curve. In Scenario 2 UAV-1 ADS-2, the cumulative extrinsic reward and task completion rate of the TH-MUPL algorithm increase rapidly with the number of training rounds. The task completion rate of the TH-MUPL algorithm reaches 60%, 80%, and 90% after approximately 4.6×10 3 , 5.5×10 3 , 7.5×10 3 rounds of training, respectively, and the training cost is significantly reduced compared to the No-Intrinsic-Reward method. In terms of the cumulative extrinsic reward metric, the curve of the TH-MUPL algorithm grows very fast at the beginning, but it takes about 15,000 rounds of training to converge to 0.4. This phenomenon is because after the TH-MUPL algorithm's hierarchical decision-making searches for positive rewards, it maximizes the cumulative extrinsic reward by optimizing the macro policy. When the macro policy reaches a certain task success rate, it then transitions to the parallel policy based on the original action space. There are some redundant decision-making steps in the macro cooperation policy generated by the TH-MUPL during the early training, so the cumulative extrinsic reward is slightly lower. When gradually transitioning to the parallel policy, the performance of the TH-MUPL cooperation policy is further improved, and the cumulative extrinsic reward also increases accordingly.
[0078] Figure 4 shows the cumulative extrinsic reward and task success rate curves of each test algorithm in Scenario 3 UAV-1 ADS, where Figure 4 (a) is the cumulative extrinsic reward curve Figure 4(b) is the task success rate curve graph. The cumulative extrinsic rewards of TH-MUPL all converge to around 0.25, and the task success rates all converge to around 95%. The number of training epochs required for TH-MUPL to achieve task completion rates of 60%, 80%, and 90% is approximately 8.0×10 3 、1.0×10 4 and 1.2×10 4 , and the training cost is significantly reduced compared to the No-Intrinsic-Reward method. In addition, the task success rate curve of the TH-MUPL algorithm grows rapidly in the scenario of 3UAV-1ADS, but the cumulative extrinsic reward index does not show an advantage. The reason is the same as that in the scenario of 2UAV-1ADS-2, that is, the TH-MUPL algorithm first reaches a higher success rate through the macroscopic cooperation strategy, so it does not have an advantage in the cumulative extrinsic reward index that comprehensively reflects the task success rate and the performance of the cooperation strategy.
[0079] Cumulative extrinsic reward and task success rate curve graph of scenario 3UAV-2ADS, where Figure 5 (a) is the cumulative extrinsic reward curve graph, Figure 5 (b) is the task success rate curve graph. In the scenario of 3UAV-2ADS, the TH-MUPL algorithm performs excellently in both the cumulative extrinsic reward and task success rate indicators. The TH-MUPL algorithm needs approximately 6.1×10 3 、7.8×10 3 、9.0×10 3 to achieve success rates of 60%, 80%, and 90%, and the training cost is significantly reduced compared to the No-Intrinsic-Reward method.
[0080] The cumulative extrinsic reward and task success rate curve graph of scenario 5UAV-2ADS is as shown in Figure 6 where Figure 6 (a) is the cumulative extrinsic reward curve graph, Figure 6 (b) is the task success rate curve graph. The bar graph of the number of training epochs required for each algorithm strategy to reach the specified success rate is as shown in Figure 7 where 7(a) is the bar graph of the number of training epochs required for each algorithm strategy to reach the specified success rate in the scenario of 2UAV-1ADS-2, Figure 7 (b) is the bar graph of the number of training epochs required for each algorithm strategy to reach the specified success rate in the scenario of 3UAV-1ADS, Figure 7 (c) is the bar graph of the number of training epochs required for each algorithm strategy to reach the specified success rate in the scenario of 3UAV-2ADS, Figure 7 (d) is the bar graph of the number of training epochs required for each algorithm strategy to reach the specified success rate in the scenario of 5UAV-2ADS. In the scenario of 5UAV-2ADS with high task difficulty, only TH-MUPL can achieve it in 5.0×104 Learn an effective multi-UAV cooperative penetration and assault strategy within the training rounds. After 5.0×10 4 training rounds, the cumulative external reward of the TH-MUPL algorithm reaches approximately -0.75, and the mission success rate increases to approximately 80%. The number of training rounds required for the TH-MUPL algorithm to reach success rates of 60% and 80% is approximately 4.0×10 4 and 4.8×10 4 . As shown in the above simulation experiment results, after the maximum number of training rounds in each scenario, the TH-MUPL algorithm achieves a success rate of over 95% in the scenarios of 2UAV-1ADS-2, 3UAV-1ADS, and 3UAV-2ADS, and a success rate of over 80% in the scenario of 5UAV-2ADS, indicating that the TH-MUPL algorithm can successfully train an effective multi-UAV cooperative penetration and assault strategy. The number of training rounds required for the TH-MUPL algorithm to reach success rates of 60%, 80%, and 90% is significantly reduced, indicating that the time-series hierarchical decision-making structure effectively improves the optimization efficiency of multi-UAV cooperative strategies under sparse reward conditions.
[0081] In one embodiment, a multi-UAV cooperative strategy learning device based on time-series hierarchical decision-making is provided, including: a UAV macro-action space construction module, a hierarchical strategy training module, a hierarchical decision-making module, and a multi-UAV cooperative strategy determination module, where: The UAV macro-action space construction module is used to perform time-series abstraction on the original sequential decision sequence by using the time-series hierarchical decision-making mechanism, and construct a UAV macro-action space capable of reaching the positive reward state; each macro-action consists of original actions.
[0082] The hierarchical strategy training module is used to train the pre-constructed hierarchical strategy according to the current observation, macro-action, and original action; the hierarchical strategy includes a macro-strategy and a micro-strategy; the macro-strategy is used to select a macro-action according to the current observation, where each UAV selects a macro-action every steps; the micro-strategy is used to select an original action every step according to the current observation and the most recently selected macro-action.
[0083] The hierarchical decision-making module is used for the UAV to perform hierarchical decision-making search in the macro-action space through the hierarchical strategy, and store the experience samples generated by the UAV hierarchical decision-making into the experience cache.
[0084] The multi-UAV cooperation strategy determination module is used to, after the UAVs search for positive external rewards using a hierarchical strategy, offline train a pre-built parallel strategy with experience samples containing original actions sampled from an experience cache, and then gradually replace the hierarchical strategy with the parallel strategy to obtain a multi-UAV cooperation strategy; the parallel strategy is used to output original actions according to the current observation, and only external rewards are used during training.
[0085] In one embodiment, the UAV macro-action space construction module is further used to analyze and summarize based on known successful cooperation behaviors, extract the original action sequences coexisting in multiple cooperation modes as macro-actions; decompose the multi-UAV cooperation action sequences of successful tasks into macro-action sequences, and abstract a reasonable macro-action space from multiple spatio-temporal cooperation modes; where the macro-action space is: ;
[0086] Wherein, is the macro-action space.
[0087] In one embodiment, the hierarchical strategy training module is further used to define the micro-strategy in a pre-trained or rule-based manner, and the macro-strategy is pre-trained using a reinforcement learning method; input the current observation and macro-action into the micro-strategy to output the original action; use the ACER method to update the UAV's macro-strategy, and the macro-strategy includes a UAV macro-value network and a macro-strategy network.
[0088] In one embodiment, the hierarchical strategy training module is further used to when the macro-action selects to approach the target, that is , the UAV micro-strategy is as shown in the UAV micro-strategy expression when the macro-action selects to approach the target; when the macro-action selects to approach threat k, that is , the UAV micro-strategy is as shown in the UAV micro-strategy expression when the macro-action selects to approach the threat; when the macro-action selects to move away from threat area k, that is , the UAV micro-strategy is as shown in the UAV micro-strategy expression when the macro-action selects to move away from the threat area; when the macro-action selects "fly clockwise around the nearest threat area", that is , the UAV micro-strategy is as shown in the UAV micro-strategy expression when the macro-action selects to fly clockwise around the nearest threat area; when the macro-action selects "fly counterclockwise around the nearest threat area", that is , the UAV micro-strategy is as shown in the UAV micro-strategy expression when the macro-action selects to fly counterclockwise around the nearest threat area; when the macro-action selects "hover in place", that is When the macro-action of the UAV selects to hover in place, the micro-strategy of the UAV is as shown in the expression of the UAV micro-strategy above.
[0089] In one embodiment, the macro-value network of the UAV is updated by using the update method of the above-mentioned UAV macro-value network.
[0090] In one embodiment, the macro-strategy network of the UAV is updated by using the update method of the above-mentioned UAV macro-strategy network.
[0091] In one embodiment, the hierarchical strategy training module is further configured to store and extract UAV experience samples by using a concurrent experience replay mechanism; the concurrent experience replay mechanism includes: the experience cache capacities of each UAV are equal, and the storage of experience samples uniformly follows the first-in-first-out principle, ensuring that each experience cache is located at the same index position The experience tuples of are derived from the same state transition process; when sampling the experience cache, the indices of the batch samples sampled by all UAVs are the same.
[0092] In one embodiment, the hierarchical strategy training module is further configured to start training the parallel strategy by using offline experience samples and optimize the parallel strategy through D3QN when the UAV stably searches for a positive external reward through the hierarchical strategy; the parallel strategy includes an action-value network and a target-value network; the action-value network and the target-value network have the same structure but different parameters; the action-value network is updated by using the above-mentioned action-value network parameter update method; the target-value network is updated by using the above-mentioned target-value network parameter update method.
[0093] In one embodiment, the parallel strategy is gradually used to replace the hierarchical strategy in the hierarchical strategy training module, specifically including: by hyperparameters Adjust the strategy replacement process, is a positive integer, ; after searching for a positive reward, start training the parallel strategy; when , start making decisions using the parallel strategy instead of the hierarchical strategy, and record the decision round number as ; where represents the success rate of the UAV completing the task through the hierarchical cooperation strategy; after transition rounds, the parallel strategy completely replaces the hierarchical strategy; the parallel strategy of the UAV is not presented in an independent parameterized form, but is defined according to the value function ; all UAVs make action selections based on the method.
[0094] For the specific limitations of the multi-UAV cooperation strategy learning device based on sequential hierarchical decision-making, reference can be made to the limitations of the multi-UAV cooperation strategy learning method based on sequential hierarchical decision-making in the foregoing text, which will not be elaborated here. Each module in the above multi-UAV cooperation strategy learning device based on sequential hierarchical decision-making can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0095] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0096] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A multi-UAV cooperative strategy learning method based on temporal hierarchical decision making, characterized in that: The method comprises: The temporal hierarchical decision mechanism is used to perform temporal abstraction on the original sequential decision sequence, and a macro action space of the drone with the ability to achieve a positive reward state is constructed; each macro action is composed of It consists of 100 primitive actions; According to the current observation, macro action and original action, the pre-built hierarchical strategy is trained; the hierarchical strategy includes macro strategy and micro strategy; the macro strategy is used to select macro action according to the current observation, wherein each drone The micro-strategy is used to select the original action at each step based on the current observation and the most recently selected macro-action; The UAV performs hierarchical decision search in the macro action space through the hierarchical strategy, and stores the experience samples generated by the UAV hierarchical decision in the experience cache; After the drone searches for a positive extrinsic reward using the hierarchical strategy, the pre-constructed parallel strategy is trained offline using experience samples containing original actions sampled from the experience cache, and then the parallel strategy is gradually used to replace the hierarchical strategy to obtain a multi-drone collaboration strategy; the parallel strategy is used to output original actions based on current observations, and only extrinsic rewards are used during training.
2. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 1 is characterized in that: The temporal hierarchical decision mechanism is used to perform temporal abstraction on the original sequential decision sequence, and a macro action space of the drone with the ability to reach a positive reward state is constructed, including: Based on the analysis and induction of known successful collaborative behaviors, the original action sequences coexisting in multiple collaborative modes are extracted as macro actions; The multi-UAV cooperative action sequence of a successful mission is decomposed into a macro action sequence, and a reasonable macro action space is abstracted from a variety of spatiotemporal coordination modes; the macro action space is: ; in, The macro action space.
3. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 1 is characterized in that: According to the current observation, macro action and original action, the pre-built hierarchical strategy is trained; the hierarchical strategy includes macro strategy and micro strategy; the macro strategy is used to select macro action according to the current observation, wherein each drone Step 1: Select a macro action. The micro-strategy is used to select the original action at each step based on the current observation and the most recently selected macro-action, including: The micro-strategy is defined by pre-training or rule-based definition, and the macro-strategy is obtained by pre-training using a reinforcement learning method; Input the current observation and macro action into the micro strategy and output the original action; The ACER method is used to update the macro strategy of the drone, and the macro strategy includes the drone macro value network and the macro strategy network.
4. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 3 is characterized in that: Input the current observation and macro action into the micro strategy and output the original action, including: When the macro action is to approach the target, the drone's micro strategy is: ; in, is the drone’s micro-strategy when the macro action is to approach the target, Indicates drone Current observations, Indicates approaching the target. Indicates drone Select the heading angle unit vector after each primitive action, From the drone The unit vector pointing to the protected target, Indicates k A primitive action; represents the original action space; When macro actions choose to approach threats k When , the drone micro-strategy is: ; in, represents the drone’s micro-strategy when the macro action is to approach the threat k, Indicates drone Point to threat area The unit vector of Indicates approaching threat k ; When macro actions choose to stay away from the threat area When , the drone micro-strategy is: ; in, Indicates that when the macro action is to move away from the threat area k Drone micro-strategy at the time, Stay away from the threat area ; When the macro action is to fly clockwise around the nearest threat area, the drone's micro strategy is: ; in, It represents the drone’s micro-strategy when the macro-action is to fly clockwise around the nearest threat area; Indicates flying clockwise around the nearest threat area; Representation and from drone A unit vector that is orthogonal to the vector closest to the threat area and points in a clockwise direction; When the macro action is to fly counterclockwise around the nearest threat area, the drone's micro strategy is: ; in, It represents the drone’s micro-strategy when the macro-action is to fly counterclockwise around the nearest threat area; It means flying counterclockwise around the nearest threat area. When the macro action is to hover in place, the drone's micro strategy is: ; in, It represents the micro-strategy of the drone when the macro-action is hovering in place; Indicates the center of the spiral; Indicates the execution of the original action Next moment drone The predicted position of Indicates circling in place.
5. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 3 is characterized in that: The update method of the drone macro value network is: ; ; ; in, represents the parameters of the drone macro value network, represents the update hyperparameters of the drone macro value network parameters, represents the gradient of the mean square error of the macro value network, Represents the value estimation of the macro strategy based on the Retrace method. Indicates drone i Current observations, Indicates drone i The macro action, represents the total reward, yes The truncated importance sampling coefficient value at time , represents the value network, represents the total number of threat areas, represents the hyperparameter, represents the value function of the observation, Respectively represent the drone at that moment i Observation and macro-action.
6. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 3 is characterized in that: The update method of the drone macro strategy network is: ; ; in, represents the parameters of the drone macro strategy network, represents the update hyperparameters of the drone macro strategy network parameters, represents the ACER policy gradient, represents the total number of threat areas, represents the value function of the observation, , Respectively t Two different values of truncated importance sampling coefficients at time instant, represents the drone macro value network, represents the drone macro strategy network, Indicates i A macro action, Indicates drone i Current observations, Represents the value estimation of the macro strategy based on the Retrace method.
7. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 1 is characterized in that: A concurrent experience playback mechanism is used to store and extract drone experience samples; The concurrent experience playback mechanism includes: the experience cache capacity of each drone is equal, the storage of experience samples follows the first-in-first-out principle, and each experience cache is located at the same index position The experience tuples of come from the same state transition process; When sampling the experience cache, all drones sample the same batch of samples.
8. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 1 is characterized in that: After the drone uses the hierarchical strategy to search for positive extrinsic rewards, it uses the experience samples containing the original actions in the cache to offline train the pre-built parallel strategy, including: After the drone stably searches for positive extrinsic rewards through the hierarchical strategy, it starts to train the parallel strategy using offline experience samples and optimizes the parallel strategy through D3QN; the parallel strategy includes an action value network and a target value network; the action value network and the target value network have the same structure but different parameters; The action value network parameters are updated as follows: ; ; ; ; in, represents the action-value network parameters, represents the action-value network, represents the loss function, For external rewards, Indicates value, represents the learning rate, express t The first truncated importance sampling coefficient value at time, represents the observation at the next moment, , Respectively represent drones i The current observation and the original action, represents the advantage function network parameters, represents the advantage function network, represents the state value function network parameters, Represents the state-value function network; The target value network parameter update method is: ; in, represents the target value network parameter, Represents the learning rate.
9. The multi-UAV cooperative strategy learning method based on temporal hierarchical decision-making according to claim 1 is characterized in that: The parallel strategy is gradually adopted to replace the hierarchical strategy, including: Through hyperparameters Adjustment strategy replacement process, , is a positive integer; After searching for positive rewards, start training the parallel strategy; when When , the parallel strategy is used instead of the hierarchical strategy to make decisions. The number of decision rounds is recorded as ;in represents the success rate of UAVs completing tasks through hierarchical collaboration strategies; go through The parallel strategy completely replaces the hierarchical strategy. The parallel strategy of the drone is defined based on the action value function. All drones are based on Method to select action: ; in , , drones The probability of selecting the action corresponding to the maximum action value function is All candidate actions are randomly selected with probability, represents the original action space, represents the action-value network, k Indicates the action number. represents the action-value network parameters, , Respectively represent drones i The current observation and the original action, Indicates k A primitive action, Indicates drone i parallel strategy.
10. A multi-UAV cooperative strategy learning device based on temporal hierarchical decision-making, characterized in that: The device comprises: The drone macro action space construction module is used to use the temporal hierarchical decision-making mechanism to perform temporal abstraction on the original sequential decision sequence and construct a drone macro action space capable of reaching a positive reward state; each macro action consists of It consists of 100 primitive actions; The hierarchical strategy training module is used to train the pre-built hierarchical strategy according to the current observation, macro action and original action; the hierarchical strategy includes macro strategy and micro strategy; the macro strategy is used to select macro action according to the current observation, wherein each drone The micro-strategy is used to select the original action at each step based on the current observation and the most recently selected macro-action; A hierarchical decision module, used for the UAV to perform hierarchical decision search in the macro action space through the hierarchical strategy, and store the experience samples generated by the UAV hierarchical decision into the experience cache; The multi-UAV cooperation strategy determination module is used to, after the UAV uses the hierarchical strategy to search for positive extrinsic rewards, use the experience samples containing the original actions sampled from the experience cache to offline train the pre-constructed parallel strategy, and then use the parallel strategy to gradually replace the hierarchical strategy to obtain the multi-UAV cooperation strategy; the parallel strategy is used to output the original action according to the current observation, and only use the extrinsic reward during training.
Citation Information
Patent Citations
Unmanned aerial vehicle group path planning method based on improved Q learning algorithm
CN109443366A
Fixed-wing unmanned aerial vehicle autonomous control cooperation strategy training method
CN112034888A
Unmanned aerial vehicle layered flight decision-making method based on SAC algorithm
CN115185288A
Unmanned aerial vehicle group path-finding system fusing block chain macroscopic task and microscopic decision instruction
CN117114325A
Unmanned aerial vehicle air combat confrontation method based on hierarchical reinforcement learning
CN119556722A