Multi-unmanned aerial vehicle cooperation strategy learning method based on inter-unmanned aerial vehicle subsequent influence internal reward
By introducing inter-aircraft successive influence intrinsic rewards in multi-drone collaboration tasks, combined with multi-agent reinforcement learning of the Dec-POMDP framework, the problem of UAV exploring complex joint state space under sparse reward conditions is solved, and the efficiency of collaboration strategy optimization is improved.
Patent Information
- Application Number
- CN202510627085.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In multi-UAV collaboration tasks, the existing technology is difficult to effectively solve the problem of drones exploring complex joint state spaces under sparse reward conditions, especially in close collaboration tasks and multi-objective search tasks, the search and utilization of positive rewards is difficult.
The multi-drone collaboration strategy learning method based on intrinsic rewards between machines is adopted. By constructing the multi-agent reinforcement learning process of the Dec-POMDP framework, weighted summing is performed by combining external rewards, intrinsic rewards with novel equivalent state values and intrinsic rewards based on successive influences to optimize the drone strategy.
This method effectively overcomes the adverse impact of the complex structure of joint state space on forward reward search, improves the optimization efficiency of multi-UAV collaboration strategy under sparse reward conditions, and encourages drones to explore inter-machine interaction states with high degree of mutual influence.
Smart Images

Figure CN120122702A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multi-UAV collaborative strategy learning, and particularly to a multi-UAV collaborative strategy learning method based on the intrinsic reward of inter-UAV successor influence. Background Art
[0002] In reinforcement learning under sparse reward conditions, intrinsic rewards based on metrics such as state novelty, curiosity, and information gain are often used to drive agents to explore states with high uncertainty, thereby improving the efficiency of agents in discovering and exploiting positive extrinsic rewards. However, using novelty as the sole incentive to drive the exploration of multi-UAVs has limitations. When the number of UAVs increases or the task area expands, the joint state space increases sharply, and the cost of traversal exploration driven by novelty intrinsic rewards increases significantly. In some close collaboration tasks, there are mutual influences among the dynamics of UAVs. The coupling of individual state transitions makes the joint state transition exhibit complex structures such as causal dependence and sequential constraints, reducing the connectivity of the joint state space. And in multi-object search tasks, since agents need to cooperate closely in specific states to complete, positive rewards usually appear in the subsequent states of the coupled transitions of individual states, further exacerbating the difficulty of searching for them. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a multi-UAV collaborative strategy learning method based on the intrinsic reward of inter-UAV successor influence.
[0004] A multi-UAV collaborative strategy learning method based on the intrinsic reward of inter-UAV successor influence, the method includes: Construct the multi-UAV collaborative strategy learning process into a multi-agent reinforcement learning process of the Dec-POMDP framework.
[0005] Each UAV selects and executes an action according to its own strategy and the current observation. After the environmental state transfers, each UAV obtains an extrinsic reward.
[0006] Each UAV calculates the intrinsic reward based on the equivalent state novelty value according to the equivalent state of the next state after the state transfer.
[0007] Each UAV determines its own UAV trajectory feature and the UAV trajectory feature affected by the states of other UAVs according to its own current state, the selected action, and the current states of other UAVs.
[0008] Determine the successor influence between two UAVs according to the UAV's own trajectory feature and the UAV trajectory feature affected by the states of other UAVs.
[0009] Determine the intrinsic reward based on the successor influence according to the successor influences among all UAVs.
[0010] The external reward, the internal reward based on the equivalent state novelty value, and the internal reward based on the successor impact are weighted and summed to obtain the overall reward.
[0011] According to the overall reward and the experience samples obtained during the UAV-environment interaction, the ACER method is used to optimize the UAV strategy, and the value network and the policy network are updated.
[0012] The above multi-UAV cooperative strategy learning method based on the internal reward of inter-aircraft successor impact. The method overcomes the adverse impact of the complex structure of the joint state space on the positive reward search by encouraging the UAVs to explore the interaction states with high mutual influence, thereby improving the optimization efficiency of the multi-UAV cooperative strategy under sparse reward conditions; adding the inter-aircraft successor impact as an internal reward on the basis of the equivalent state novelty value to drive the UAVs to cooperate and explore, so as to encourage the UAVs to explore the inter-aircraft interaction states with high mutual influence degree. Brief Description of the Drawings
[0013] Figure 1 It is a schematic flowchart of the multi-UAV cooperative strategy learning method based on the internal reward of inter-aircraft successor impact in an embodiment; Figure 2 It is a schematic diagram of the multi-UAV task scenario in another embodiment, where Figure 2 (a) is a schematic diagram of the scenario 2UAV-1ADS-2, Figure 2 (b) is a schematic diagram of the scenario 3UAV-1ADS, Figure 2 (c) is a schematic diagram of the scenario 3UAV-2ADS, Figure 2 (d) is a schematic diagram of the scenario 5UAV-2ADS; Figure 3 It is a curve graph of the cumulative external reward and the task success rate of the scenario 2UAV-1ADS-2 in another embodiment, where Figure 3 (a) is the cumulative external reward curve graph, Figure 3 (b) is the task success rate curve graph; Figure 4 It is a curve graph of the cumulative external reward and the task success rate of the scenario 3UAV-1ADS in another embodiment, where Figure 4 (a) is the cumulative external reward curve graph, Figure 4 (b) is the task success rate curve graph; Figure 5 It is a curve graph of the cumulative external reward and the task success rate of the scenario 3UAV-2ADS in another embodiment, where Figure 5 (a) is the cumulative external reward curve graph, Figure 5 (b) is the task success rate curve graph; Figure 6For the cumulative extrinsic reward and mission success rate curve graph of Scenario 5 UAV-2 ADS in another embodiment, where Figure 6 (a) is the cumulative extrinsic reward curve graph, Figure 6 (b) is the mission success rate curve graph; Figure 7 For the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in another embodiment, where 7(a) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 2 UAV-1 ADS-2 Scenario, Figure 7 (b) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 3 UAV-1 ADS Scenario, Figure 7 (c) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 3 UAV-2 ADS Scenario, Figure 7 (d) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in Scenario 5 UAV-2 ADS Scenario. Detailed implementation manners
[0014] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0015] In one embodiment, as Figure 1 shown, a multi-UAV collaborative strategy learning method based on the intrinsic reward of inter-aircraft successor influence is provided, and the method includes the following steps: Step 100: Construct the multi-UAV collaborative strategy learning process into a multi-agent reinforcement learning process of the Dec-POMDP framework.
[0016] Specifically, an agent is deployed on each UAV, and a multi-agent system is composed of multiple UAVs. The multi-UAV collaborative process is modeled as a distributed Markov decision process (Dec-POMDP), and the multi-UAV collaborative strategy learning process is constructed into a multi-agent reinforcement learning process of the Dec-POMDP framework.
[0017] The multi-agent reinforcement learning process based on the Dec-POMDP framework can be divided into autonomous decision-making, environment interaction, and policy update. Specifically, it includes: Autonomous decision-making process: It refers to the process in which each agent selects an action based on its own policy according to the current observation , which is expressed as .
[0018] Environment interaction process: It refers to the process in which the agent executes the joint action After interacting with the environment, the joint state changes from to , and the agent obtains a reward . This process is denoted as .
[0019] Policy update process: It refers to the process in which each agent uses the experience samples generated from the interaction with the environment to update the policy .
[0020] The reinforcement learning method optimizes the agent's policy by repeating the above process.
[0021] During the above three processes, explicit or implicit mutual influences may occur among the agents. Based on their own perceptions and communication among the agents, the individual observations of agent not only contain its own state, but may also contain the state information of other agents, historical observations, etc. Even in some decision-making methods, each agent sequentially selects individual actions in turn, and the individual observation also contains the action information of other agents at the current moment . Therefore, when agent selects an action , it may be affected by the information of other agents in the observation, and the specific degree of influence is determined according to the feature extraction and fitting of the network for each input composition. This influence changes continuously with the network parameters during the policy training process. The policy parameters are updated towards the optimization goal of maximizing the cumulative reward, and the feature extraction and fitting of the observation input information also change accordingly, automatically optimizing how to make more favorable decisions based on the information of other agents under different observations.
[0022] Step 102: Each UAV selects and executes an action according to its own policy and the current observation. After the environmental state changes, each UAV obtains an extrinsic reward.
[0023] Step 104: Each UAV calculates the intrinsic reward based on the novelty value of the equivalent state according to the equivalent state of the next state after the state transition.
[0024] Step 106: Each UAV determines its own UAV trajectory features and the UAV trajectory features affected by the states of other UAVs according to its own current state, the selected action, and the current states of other UAVs.
[0025] Step 108: Determine the subsequent influence between two UAVs according to the UAV's own trajectory features and the UAV trajectory features affected by the states of other UAVs.
[0026] Specifically, when multiple agents execute a joint action , causing the joint state to transfer, the individual states also transfer accordingly . In a multi-agent scenario with complex interaction relationships, the transfer of the joint state usually cannot be fully explained as a simple superposition of individual state transfers. The individual dynamics of each agent will influence each other according to the environmental characteristics. For example, in the collaborative box-pushing task, the environmental characteristic is that the box can only be pushed when two agents exert force in the same direction simultaneously. Therefore, in such scenarios, the individual state transfers of each agent may be affected by other agents, and the specific states and degrees of influence of such effects are determined by the interaction characteristics of the agents in the environment.
[0027] In addition, during the environmental interaction process, the rewards obtained by an agent are not only determined by its own state and actions but may also depend on the states and actions of other agents. The state transfer process in multi-agent reinforcement learning can be described by a Causal Graphical Model (CGM), specifically represented by a directed acyclic graph . Nodes represent all the random variables involved in the state transfer , represents the correlation relationship between random variables, and edges represent that the random variable is 's direct influencing factor.
[0028] Construct to describe the state transfer process of the multi-agent system. According to the Markov property, the joint state variable at the next moment in is only related to the joint state variable at the previous moment and the joint action variable of the agents Macroscopically describes the causal relationship between joint state variables at adjacent moments in Dec-POMDP, and does not depict the causal relationship between the state variables of each agent hidden therein. In fact, the causal relationship between agent state variables is determined by the specific task environmental characteristics.
[0029] To further analyze the mutual influence between UAVs during collaborative task execution in a threat area, taking a task scenario involving two UAVs as an example, is expanded from an individual perspective to to depict the relationship between the individual state variables of each UAV. In complex collaborative tasks, the coupling of multi-UAV individual state transfers leads to a decrease in the connectivity of the joint state space, increasing the difficulty of sparse positive reward search and utilization.
[0030] The Successor Feature (SF) is a feature that describes the subsequent state trajectory of an agent in a Markov decision process and is commonly used in transfer learning to decouple the environmental dynamics and rewards implicit in the value function.
[0031] Step 110: Determine the intrinsic reward based on the successor influence according to the successor influence among all drones.
[0032] Step 112: Perform a weighted sum of the extrinsic reward, the intrinsic reward based on the equivalent state novelty value, and the intrinsic reward based on the successor influence to obtain the overall reward.
[0033] Step 114: According to the overall reward and the experience samples obtained during the drone-environment interaction, use the ACER method to optimize the drone policy and update the value network and the policy network.
[0034] In the above multi-drone cooperative strategy learning method based on the intrinsic reward of inter-drone successor influence, the method overcomes the adverse effect of the complex structure of the joint state space on the forward reward search by encouraging drones to explore the interaction states with high mutual influence, thereby improving the optimization efficiency of the multi-drone cooperative strategy under sparse reward conditions; adding the inter-drone successor influence as an intrinsic reward on the basis of the equivalent state novelty value to drive the cooperative exploration of drones, so as to encourage drones to explore the inter-drone interaction states with a high degree of mutual influence.
[0035] In one embodiment, step 104 includes: each drone uses the random network distillation method according to the equivalent state of the next state after the state transition to determine the equivalent state novelty value, and determines the intrinsic reward based on the equivalent state novelty value according to the equivalent state novelty degree value; wherein, the prediction network parameters in the random network distillation method are updated according to the experience samples obtained during the drone-environment interaction.
[0036] In one embodiment, step 106 includes: each drone uses the first individual drone successor feature network to perform feature extraction according to its own current state and the selected action to obtain the drone's own trajectory feature; the individual drone successor feature network includes: the first fully connected layer, the rectified linear unit, the second fully connected layer, the rectified linear unit, and the third fully connected layer; each drone uses the second individual drone successor feature network to perform feature extraction according to the states of other drones, its own current state, and the selected action to obtain the drone trajectory feature affected by the states of other drones; the parameters of the first individual drone successor feature network and the second individual drone successor feature network are updated according to the experience samples obtained during the drone-environment interaction; the expressions of the drone's own trajectory feature and the drone trajectory feature affected by the states of other drones are: ;
[0037] Among them, represents the UAV at state performing an action and the subsequent self-trajectory feature, represents the UAV being in state when the UAV at state performs an action and the subsequent self-trajectory feature, respectively represent the state and action of the UAV i , represents the state of the UAV j , represents the discount factor, represents the feature representation, k represents the time step number, represents the state of the UAV i at t time step, represents the action of the UAV i at t time step, represents the state of the UAV i at t + k +1 time step.
[0038] Specifically, given the state feature representation and the agent policy , the corresponding successor feature is defined as: ;
[0039] The successor feature is the sum of the discounted state features experienced by the agent throughout the decision-making round starting from state , which can characterize the entire subsequent trajectory information. Considering the influence of each action selected by the agent in state on the subsequent trajectory feature, is expanded to: ;
[0040] In a multi-agent scenario, each agent corresponds to its own individual successor feature . Given the feature representation and all agent policies , the individual successor feature corresponding to agent is defined as: ;
[0041] Among them, is the joint state, is the joint action.
[0042] For all agents in the joint state execute the joint action After that, the sum of the discounted individual state features experienced by agent during the entire decision-making round is characterized, describing the subsequent individual trajectory information of agent . Expand into the individual successor features under two partially observable input conditions, as specifically shown in the expressions of the drone's own trajectory features and the drone's trajectory features affected by the states of other drones as described above.
[0043] Compared with characterizes the drone Under different state inputs, the subsequent trajectory features corresponding to the drone .
[0044] In one embodiment, step 108 includes: determining the subsequent influence between two drones according to the drone's own trajectory features and the drone's trajectory features affected by the states of other drones; the expression of the subsequent influence between two drones is: ;
[0045] Among them, represents the subsequent influence of drone on drone , respectively represent the current state and the corresponding action of drone i , represents the current state of drone j ; represents the own trajectory feature of drone after executing action at the current state , represents that when drone is in the current state , the own trajectory feature of drone after executing action at the current state .
[0046] Specifically, a measure of the mutual influence between agents based on successor features is proposed, defined as Successor-feature-based Influence (SI). In a multi-agent scenario, when an agent executes a joint action at the joint state , the successor influence of agent on agent is as shown in the successor influence expression between the above two drones.
[0047] Combined with the causal graph model theory, the successor influence can capture the interaction between the dynamics of drones, that is it can reflect the influence of the trajectory of the drone after executing the action by the state of the drone . In one embodiment, step 110 includes: determining an intrinsic reward based on successor influence according to the successor influence between all drones; the expression of the intrinsic reward based on successor influence is:
[0048] ; ;
[0049] wherein, represents the intrinsic reward based on successor influence of the drone , represents the influence exerted on other drones, represents the influence received from other drones; represents the successor influence of the drone k on the drone i , represents the successor influence of the drone i on the drone k , respectively represent the current state and the executed action of the drone k , represents the current state of the drone i . Refers to .
[0050] Specifically, in complex collaborative tasks, the interaction between drones is usually a prerequisite for collaboration. Therefore, an intrinsic reward is defined based on the successor influence as shown in the above expression of the intrinsic reward based on successor influence to urge drones to interact actively. The intrinsic reward based on successor influence of the drone Including the influence exerted on other drones and the influence received from other drones, it drives the drone to actively access the state with high inter-drone interaction, encourages potential collaborative behaviors, and thus improves the search efficiency for positive external rewards.
[0051] In one embodiment, step 112 includes: performing a weighted sum of the external reward, the intrinsic reward based on the equivalent state novelty value, and the intrinsic reward based on the successor influence to obtain the overall reward; the expression of the overall reward is: ;
[0052] where, represents the overall reward, represents the external reward, represents the drone 's intrinsic reward based on the successor influence, represents the intrinsic reward based on the equivalent state novelty value, , and are the weights of the external reward, the intrinsic reward based on the equivalent state novelty value, and the intrinsic reward based on the successor influence, respectively.
[0053] Specifically, to ensure that the drone has the driving force to explore new states, the intrinsic reward based on the equivalent state novelty value and the intrinsic reward based on the successor influence are jointly used for policy update, and the overall reward is defined as shown in the expression of the above overall reward.
[0054] In one embodiment, the value network update process in step 114 includes: ; ; ;
[0055] where, represents the value network parameters, and represent two hyperparameters, represents the value estimation of the current policy based on the Retrace method, represents the mean square error of the value network, represents the calculated gradient, represents the calculated expectation, represents the value of the current policy, represents the overall reward, represents t the truncated importance sampling coefficient value at the , , Denote the unmanned aerial vehicle Generate experience The behavioral strategy at that time, Denote its own strategy, Denote the value of the observation, Respectively denote the unmanned aerial vehicle i The current observation and the executed action, Respectively denote the unmanned aerial vehicle i The observation and the executed action at the next moment.
[0056] For the sake of simplicity in the text, some are represented by And Respectively replace the macro policy network of the unmanned aerial vehicle And the macro value network of the unmanned aerial vehicle .
[0057] In one embodiment, the update process of the policy network in step 114 includes: ; ;
[0058] Among them, Denote the policy network parameters, Denote the hyperparameters, Denote the policy network, Denote the ACER policy gradient, Denote the calculated gradient, k Denote the trajectory length sampled from the experience cache of the unmanned aerial vehicle, Respectively denote t The two different truncated importance sampling coefficient values at the moment, , Denote the unmanned aerial vehicle Generate experience The behavioral strategy at that time, Denote its own strategy, Denote the value of the observation, Denote the value of the current strategy, Denote the value estimation of the current strategy based on the Retrace method, Respectively denote the unmanned aerial vehicle i The current observation and the executed action.
[0059] In a specific embodiment, the multi-unmanned aerial vehicle policy learning based on the inter-aircraft successor influence intrinsic reward includes two parts: unmanned aerial vehicle-environment interaction and policy training. Specifically, it includes: During the unmanned aerial vehicle-environment interaction, each unmanned aerial vehicle selects an action according to its own strategy And the current observation and execute: After the environmental state changes, , each UAV calculates the intrinsic reward of the equivalent state novelty value and the intrinsic reward of the subsequent influence between UAVs , and then the UAV experience tuples of the internal and external rewards of the set , subsequent feature experience are added to the corresponding caches.
[0060] During the UAV policy training process, each UAV randomly samples from the experience samples to update the policy network and the value network , and at the same time, the PI-STF-RND prediction network for calculating the intrinsic rewards of the equivalent state novelty value and subsequent influence , the individual subsequent feature network and are updated. The UAV policy is optimized based on the ACER method, and the value network and the policy network are updated.
[0061] The value network parameters are updated through the above value network update process.
[0062] During the policy improvement process, the ACER policy gradient is used to update the parameters of the UAV policy network , and the parameters of the policy network are updated through the above policy network update process.
[0063] In one embodiment, the policy network in step 114 includes: a third fully connected layer, a rectified linear unit, a fourth fully connected layer, a rectified linear unit, a fifth fully connected layer, and a Softmax function.
[0064] In one embodiment, the value network in step 114 includes a sixth fully connected layer, a rectified linear unit, a seventh fully connected layer, a rectified linear unit, and an eighth fully connected layer.
[0065] It should be understood that although Figure 1 the steps in the flowchart of Figure 1At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages do not necessarily need to be executed and completed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0066] In a verification embodiment, with the background that multiple UAV missions arrive at the periphery of the threat area to identify protected targets under multiple threats, the following four scenarios are set for method verification: 2UAV-1ADS-2, 3UAV-1ADS, 3UAV-2ADS, and 5UAV-2ADS. The multi-UAV mission scenarios are as Figure 2 shown, where Figure 2 (a) is a schematic diagram of scenario 2UAV-1ADS-2, Figure 2 (b) is a schematic diagram of scenario 3UAV-1ADS, Figure 2 (c) is a schematic diagram of scenario 3UAV-2ADS, Figure 2 (d) is a schematic diagram of scenario 5UAV-2ADS. The scenario name includes the number of UAVs and the number of threats in the mission area.
[0067] In scenario 5UAV-2ADS, the target PA is deep in the overlapping area of the defense ranges of two threats. Due to the limitation of the mission area range, the UAVs need to cross a double-layer threat area when approaching the target, and the mission is more difficult. The scale of the UAV detachment is expanded to 5. The detailed configuration of the mission scenario is listed in Table 1.
[0068] Table 1 Parameter settings of multi-UAV mission scenarios
[0069] Five methods are tested and compared. The five methods include: The multi-UAV cooperative strategy learning method based on the intrinsic reward of inter-aircraft successor influence proposed in this application, abbreviated as: SI; The multi-UAV cooperative strategy learning method that only uses extrinsic rewards, abbreviated as: No-Intrinsic-Reward; The multi-UAV cooperative strategy learning method based on the intrinsic reward of inter-aircraft mutual influence of EITI, abbreviated as: EITI; The multi-UAV cooperative strategy learning method based on the intrinsic reward of inter-aircraft mutual influence of ISE, abbreviated as: ISE; The multi-UAV cooperative strategy learning method based on the intrinsic reward of inter-aircraft mutual influence of CIE, abbreviated as: CIE.
[0070] The specific reward settings of each algorithm are as follows: SI: The overall reward includes extrinsic rewards, equivalent state novelty values, and inter-agent successor influence intrinsic rewards: ;
[0071] No-Intrinsic-Reward: Only use extrinsic rewards to update the policy: ; EITI: The overall reward includes extrinsic rewards, equivalent state novelty values, and EITI inter-agent mutual influence intrinsic rewards: ; ;
[0072] ISE: The overall reward includes extrinsic rewards, equivalent state novelty values, and ISE inter-agent mutual influence intrinsic rewards: ; ;
[0073] CIE: The overall reward includes extrinsic rewards, equivalent state novelty values, and CIE inter-agent mutual influence intrinsic rewards: ; ;
[0074] For objective comparison, the EITI algorithm, ISE algorithm, and CIE algorithm all add the mutual influence intrinsic rewards measured by themselves on the basis of the equivalent state novelty value intrinsic rewards, and only differ from the SI algorithm in the mutual influence measurement method. All test algorithms are based on their own overall rewards , and use the ACER algorithm to update the UAV policy. The comparison experiment includes the SI algorithm and the No-Intrinsic-Reward algorithm, which are used to verify the influence of adding the inter-agent successor influence intrinsic rewards on the collaborative policy learning efficiency. The UAV policy network consists of a fully connected layer (1024), a rectified linear unit (ReLU), a fully connected layer (512), a rectified linear unit, a fully connected layer (7), and a softmax function from the input layer to the output layer in sequence. The UAV action value network consists of a fully connected layer (1024), a rectified linear unit, a fully connected layer (512), a rectified linear unit, and a fully connected layer (7) from the input layer to the output layer in sequence. The UAV individual successor feature network consists of a fully connected layer (1024), a rectified linear unit, a fully connected layer (512), a rectified linear unit, and a fully connected layer (20) from the input layer to the output layer in sequence. The state feature is represented as the ECV feature concatenation vector of the threatened lock / intercept duration in the individual state.
[0075] The comparative experiment includes the SI algorithm and the No-Intrinsic-Reward algorithm. By comparing the policy learning effects of the above algorithms in the task scenario, the impact of adding the inter-machine successor's influence on the intrinsic reward on the collaborative policy learning efficiency is analyzed.
[0076] The cumulative extrinsic reward and task success rate curves of Scenario 2 UAV-1 ADS-2 are as Figure 3 shown, where Figure 3 (a) is the cumulative extrinsic reward curve, Figure 3 (b) is the task success rate curve. In Scenario 2 UAV-1 ADS-2, after the SI algorithm is trained for 1.0×10^4 rounds, the cumulative extrinsic reward of the SI algorithm grows to 0.25, and the task success rate reaches about 95%.
[0077] The cumulative extrinsic reward and task success rate curves of Scenario 3 UAV-1 ADS are as Figure 4 shown, where Figure 4 (a) is the cumulative extrinsic reward curve graph, Figure 4 (b) is the task success rate curve graph. In Scenario 3 UAV-1 ADS, when the SI algorithm is trained to 3.0×10^4 rounds, the SI cumulative extrinsic reward converges to about 0.25, and the success rate is greater than 90%. The state space size in Scenario 3 UAV-1 ADS is several times that of 2 UAV-1 ADS-2. The cost of traversing and exploring only based on the equivalent state novelty value reward increases significantly. Since the SI algorithm can effectively focus on the interaction state, the improvement of the utilization efficiency of the positive reward search is more obvious. The cumulative extrinsic reward and task success rate curves of Scenario 3 UAV-2 ADS are as Figure 5 shown, where Figure 5 (a) is the cumulative extrinsic reward curve graph, Figure 5 (b) is the task success rate curve graph. In Scenario 3 UAV-2 ADS, when the SI algorithm is trained for about 1.5×10 4 rounds, the cumulative extrinsic reward reaches above 1.0.
[0078] The cumulative extrinsic reward and task success rate curves of Scenario 5 UAV-2 ADS are as Figure 6 shown, where Figure 6 (a) is the cumulative extrinsic reward curve graph, Figure 6 (b) is the task success rate curve graph. The bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate is as Figure 7 shown, where 7(a) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in the Scenario 2 UAV-1 ADS-2 scenario, Figure 7 (b) is the bar graph of the number of training rounds required for each algorithm strategy to reach the specified success rate in the Scenario 3 UAV-1 ADS scenario, Figure 7(c) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 3 UAV-2 ADS scenario. Figure 7 (d) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 5 UAV-2 ADS scenario.
[0079] The task difficulty of the scenario 5 UAV-2 ADS is relatively high. After training, the SI algorithm accumulates an external reward of about -0.5, and the task success rate converges to more than 90%. In the task scenario setting part, it has been analyzed and explained that in the scenario 5 UAV-2 ADS, the target is deep in the double-layer defense range, the task difficulty is high, and as the number of UAVs increases, the joint state space of the scenario 5 UAV-2 ADS far exceeds that of the scenario 2 UAV-1 ADS-2, scenario 3 UAV-1 ADS, and scenario 3 UAV-2 ADS, and the sparsity of positive rewards is significantly exacerbated. The experimental results show that as the task difficulty of UAV cooperation increases and the sparsity of the positive reward distribution increases, the SI algorithm is more effective in improving the optimization efficiency of UAV cooperation strategies.
[0080] The simulation results show that in the test scenario, compared with the method based only on the equivalent state novelty value-driven exploration, the inter-aircraft successor influence intrinsic reward effectively improves the cumulative external reward and the task success rate, shortening the number of training rounds required for the UAV cooperation strategy to reach 60%, 80%, and 90% success rates by 25.0%, 62.5%, 24.3%, 52.6%, 26.7%, and 53.1% respectively. Compared with other mutual influence measurement methods, the successor influence based on the SF-Retrace estimate has a more significant effect on improving the optimization efficiency of cooperation strategies, shortening the number of training rounds required to reach 60%, 80%, and 90% success rates by 13.7%, 50.5%, 20.4%, 34.0%, 19.0%, and 35.5%, fully verifying the effectiveness of the multi-UAV cooperation strategy learning method based on the inter-aircraft successor influence intrinsic reward.
[0081] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0082] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A multi-UAV cooperative strategy learning method based on the intrinsic reward of subsequent impact between machines, characterized in that: The method comprises: The multi-UAV cooperative strategy learning process is constructed as a multi-agent reinforcement learning process in the Dec-POMDP framework; Each drone selects and executes actions based on its own strategy and current observations. After the environmental state is transferred, each drone receives an external reward. Each drone calculates an intrinsic reward based on the novelty value of the equivalent state according to the equivalent state of the next state after the state transfer; Each drone determines its own trajectory characteristics and the trajectory characteristics of drones affected by the states of other drones based on its own current state and the selected action and the current states of other drones; Determine the subsequent impact between two drones based on the drone's own trajectory characteristics and the drone's trajectory characteristics affected by other drones' states; Determine the intrinsic reward based on the subsequent impact among all drones; The extrinsic reward, the intrinsic reward based on the novelty value of the equivalent state, and the intrinsic reward based on the subsequent impact are weightedly summed to obtain a total reward; According to the overall reward and the experience samples obtained during the UAV-environment interaction, the ACER method is used to optimize the UAV strategy and update the value network and the policy network.
2. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: Each drone calculates an intrinsic reward based on the novel value of the equivalent state according to the equivalent state of the next state after the state transfer, including: Each UAV adopts a random network distillation method according to the equivalent state of the next state after the state transfer to determine the equivalent state novelty value, and determines the intrinsic reward based on the equivalent state novelty value according to the equivalent state novelty value; wherein the predicted network parameters in the random network distillation method are updated according to the experience samples obtained during the UAV-environment interaction process.
3. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: Each drone determines its own trajectory characteristics and the trajectory characteristics of drones affected by the states of other drones based on its current state, the action it selects, and the current states of other drones, including: Each drone uses the first drone individual successor feature network to extract features according to its current state and selected action, and obtains the drone's own trajectory features; the drone individual successor feature network includes: a first fully connected layer, a linear rectifier unit, a second fully connected layer, a linear rectifier unit, and a third fully connected layer; Each UAV uses the second UAV individual successor feature network to extract features based on the status of other UAVs, its own current status and the selected action, and obtains the UAV trajectory features affected by the status of other UAVs; the parameters of the first UAV individual successor feature network and the second UAV individual successor feature network are updated based on the experience samples obtained during the UAV-environment interaction process; The trajectory characteristics of the drone itself and the trajectory characteristics of the drone affected by the status of other drones are: ; in, Indicates drone In Status Execute action The subsequent trajectory characteristics, Indicates drone In state When the drone In Status Execute action The subsequent trajectory characteristics, Respectively represent drones i status and actions, Indicates drone j status, represents the discount factor, Indicates feature representation, k Indicates the time sequence number. Indicates drone i exist t The state of the moment, Indicates drone i exist t The action of the moment, Indicates drone i exist t + k +1 status at the moment.
4. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: According to the trajectory characteristics of the drone itself and the trajectory characteristics of the drone affected by the status of other drones, the subsequent impact between the two drones is determined as follows: ; in, Indicates drone UAV The subsequent impact of Respectively represent drones i The current state and corresponding action, Indicates drone j The current status of Indicates drone In the current state Execute action The subsequent trajectory characteristics, Indicates drone In current state When the drone In the current state Execute action The subsequent trajectory characteristics.
5. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: According to the subsequent impact among all drones, the intrinsic reward based on the subsequent impact is determined as: ; in, Indicates drone Intrinsic rewards based on subsequent impact, Indicates the impact exerted on other drones, Indicates being affected by other drones; Indicates drone k UAV i The subsequent impact of Indicates drone i UAV k The subsequent impact of Respectively represent drones k The current state and actions performed, Indicates drone i The current status.
6. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: The extrinsic reward, the intrinsic reward based on the novelty value of the equivalent state, and the intrinsic reward based on the subsequent impact are weighted and summed to obtain the total reward: ; in, represents the total reward, Indicates external rewards, Indicates drone Intrinsic rewards based on subsequent impact, represents the intrinsic reward based on the novelty value of the equivalent state, , and They are the weights of extrinsic reward, intrinsic reward based on equivalent state novelty value, and intrinsic reward based on subsequent impact.
7. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: The value network update process includes: ; ; ; in, represents the value network parameters, and represents two hyperparameters, Represents the value estimate of the current strategy based on the Retrace method. represents the mean square error of the value network, represents the calculation of gradient, represents the computational expectation, represents the value of the current strategy, represents the total reward, express t +1 moment truncation importance sampling coefficient value, represents the value of the observation, Respectively represent drones i The current observation and the action performed, Respectively represent drones i The observation and action to be performed at the next moment.
8. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: The updating process of the policy network includes: ; ; in, represents the policy network parameters, represents the hyperparameter, represents the policy network, represents the ACER policy gradient, represents the calculation of gradient, k represents the trajectory length sampled from the UAV experience cache, Respectively t Two different values of truncated importance sampling coefficients at time instant, , Indicates drone Generate experience Behavioral strategies when Express your own strategy, represents the value of the observation, represents the value of the current strategy, Represents the value estimate of the current strategy based on the Retrace method. Respectively represent drones i The current observation and the action performed.
9. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: The strategy network includes: a third fully connected layer, a linear rectifier unit, a fourth fully connected layer, a linear rectifier unit, a fifth fully connected layer and a Softmax function.
10. The multi-UAV cooperation strategy learning method based on the intrinsic reward of subsequent impact between machines according to claim 1 is characterized in that: The value network includes a sixth fully connected layer, a linear rectifier unit, a seventh fully connected layer, a linear rectifier unit, and an eighth fully connected layer.
Citation Information
Patent Citations
Joint learning method for features and strategies based on state features and subsequent features
CN108898221A
Network parameter updating method and device for multi-agent system and terminal equipment
CN112465148A
Multi-unmanned aerial vehicle scheduling method based on hierarchical multi-agent reinforcement learning
CN118778700A
Navigation strategy optimization method in intensive obstacle environment of reinforcement learning unmanned aerial vehicle based on state entropy excitation
CN118938988A
Invertible-reasoning policy and reverse dynamics for causal reinforcement learning
WO2023167576A2