Multi-UAV Cooperative Strategy Learning Method and Device Driven by Equivalent State Novelty Value

By replacing the invariant spatiotemporal characteristics and random network distillation method, the problem of inconsistent exploration of equivalent states in multi-UAV collaboration is solved, and the strategy learning efficiency is improved.

CN120122701BActive Publication Date: 2025-07-18NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510626666.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-18
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Under the sparse reward conditions, in the learning of multi-drone cooperative strategy, traditional novel value metrics lead to inconsistent novel values of equivalent states, resulting in the drone repeatedly exploring equivalent states, affecting the efficiency of strategy learning.

Method used

The permutation of invariant spatiotemporal feature representation and random network distillation method are used to construct novel equivalent state values, and the non-equivalent state is explored through intrinsic reward-driven drones, and combined with the external reward optimization strategy network.

Benefits of technology

It effectively reduces the repeated exploration of equivalent states and improves the learning efficiency of multi-drone collaboration strategies under sparse reward conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122701B_ABST
    Figure CN120122701B_ABST
Patent Text Reader

Abstract

This application relates to a multi-UAV collaborative strategy learning method and device driven by equivalent state novelty values. The method defines a permutation-invariant spatio-temporal feature representation, aggregates spatio-temporal information in the joint state through class image feature aggregation, and realizes the unified representation of equivalent states; based on the permutation-invariant spatio-temporal features, the random network distillation method is used to quickly and effectively evaluate the access degree of the high-dimensional equivalent state set, realizes the consistent measurement of the equivalent state novelty value, and end-to-end estimates the overall access degree of the equivalent state set; drives the UAV exploration with the equivalent state novelty value; finally, uses the equivalent state novelty value as a shared intrinsic reward to jointly guide the policy update with the original extrinsic reward. This method improves the efficiency of collaborative policy optimization under sparse reward conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of multi-UAV collaborative strategy learning, and particularly to a multi-UAV collaborative strategy learning method and device driven by equivalent state novelty values. Background Art

[0002] Under sparse reward conditions, reinforcement learning is difficult to capture positive rewards through random actions, resulting in a lack of feedback signals for the policy and making it difficult to optimize. Compared with single UAVs, the size of the joint state space in multi-UAV scenarios grows exponentially with the increase in the number of agents, and specific joint states require the cooperation of each UAV to reach, leading to more severe challenges brought by sparse rewards. Therefore, there is an urgent need for an effective collaborative exploration mechanism to drive multi-UAVs to search for positive rewards in the joint state space, so as to solve the problem of multi-UAV policy learning under sparse rewards. Constructing an intrinsic reward based on state novelty is a classic exploration mechanism in reinforcement learning. Novelty is a measure of the frequency of state access. This method effectively drives UAVs to explore the state space in a traversal manner by assigning higher intrinsic reward values to states with low access frequencies. In multi-UAV scenarios, there are a large number of equivalent states in the joint state space due to the isomorphic characteristics of UAVs. Traditional novelty metrics only estimate the access degree of the state itself, resulting in inconsistent novelty metrics for equivalent states. UAVs repeatedly explore equivalent states during the search for positive (extrinsic) rewards, affecting the efficiency of policy learning. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a multi-UAV collaborative strategy learning method and device driven by equivalent state novelty values.

[0004] A multi-UAV collaborative strategy learning method driven by equivalent state novelty values, the method comprising:

[0005] In a multi-UAV task scenario, determine a joint state space according to a multi-UAV collaborative decision-making model based on a distributed partially observable Markov decision process.

[0006] Construct an equivalent state set for each state according to the joint state space.

[0007] Each UAV selects an action according to its own policy and the current observation. After all UAVs execute the joint action and undergo state transition, each UAV obtains a first experience tuple including the observation, action, and extrinsic reward.

[0008] Use permutation-invariant spatio-temporal feature representation to uniformly represent the equivalent states of the state after state transition.

[0009] The novelty value of the equivalent state is measured using the random network distillation method according to the permutation-invariant spatio-temporal characteristics of the equivalent state to determine the intrinsic reward; and the intrinsic reward is incorporated into the first experience tuple and added to the experience storage; the extrinsic rewards and intrinsic rewards of all drones are the same for each state transition.

[0010] The extrinsic reward and intrinsic reward of the drone i are weighted and summed to obtain the overall reward of the drone i ; The state transition and overall reward calculation are repeatedly executed until the decision round ends.

[0011] According to the overall reward and the experience storage, the ACER method is used to optimize the value network and policy network in the Actor-Critic architecture.

[0012] A multi-drone collaborative policy learning device driven by the novelty value of equivalent states, the device includes:

[0013] An equivalent state set construction module, which is used to determine the joint state space in the multi-drone task scenario according to the multi-drone collaborative decision-making model based on the distributed partially observable Markov decision process; and construct the equivalent state set of each state according to the joint state space.

[0014] A state transition module, which is used for each drone to select an action according to its own policy and the current observation. After all drones execute the joint action and go through state transition, each drone obtains a first experience tuple containing the observation, action, and extrinsic reward; the extrinsic rewards of all drones are the same for each state transition.

[0015] An intrinsic reward determination module, which is used to uniformly represent the equivalent states of the states after state transition using permutation-invariant spatio-temporal feature representations; measure the novelty value of the equivalent states using the random network distillation method according to the permutation-invariant spatio-temporal characteristics of the equivalent states to determine the intrinsic reward; and incorporate the intrinsic reward into the first experience tuple and add it to the experience storage; all drones share the intrinsic reward.

[0016] An overall reward determination module, which is used to weight and sum the extrinsic reward and intrinsic reward of the drone i to obtain the overall reward of the drone i ; The state transition and overall reward calculation are repeatedly executed until the decision round ends.

[0017] A drone policy optimization module, which is used to optimize the value network and policy network in the Actor-Critic architecture according to the overall reward and the experience storage using the ACER method.

[0018] The above-mentioned multi-UAV cooperative strategy learning method and device driven by the novelty value of equivalent states. The method defines a permutation-invariant spatio-temporal feature representation, aggregates spatio-temporal information in the joint state through class image feature aggregation, and realizes the unified representation of equivalent states. It designs a random network distillation method based on permutation-invariant spatio-temporal features to quickly and effectively evaluate the accessed degree of the high-dimensional equivalent state set, realizes the consistent measurement of the novelty value of equivalent states, and end-to-end estimates the overall accessed degree of the equivalent state set. It drives the UAV exploration with the novelty value of equivalent states. Finally, taking the novelty value of equivalent states as the shared intrinsic reward, it jointly guides the policy update with the original extrinsic reward. This method improves the efficiency of cooperative strategy optimization under sparse reward conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic flowchart of the multi-UAV cooperative strategy learning method driven by the novelty value of equivalent states in an embodiment;

[0020] Figure 2 It is a schematic diagram of the multi-UAV cooperative task scenario in another embodiment, where Figure 2 (a) is a schematic diagram of scenario 2UAV-1ADS-1, Figure 2 (b) is a schematic diagram of scenario 2UAV-1ADS-2, Figure 2 (c) is a schematic diagram of scenario 3UAV-1ADS, Figure 2 (d) is a schematic diagram of scenario 3UAV-2ADS;

[0021] Figure 3 It is a bar chart of the number of states covered by each algorithm in 2UAV-1ADS-1 and in scenario 3UAV-1ADS in another embodiment, where Figure 3 (a) is a bar chart of the number of states covered by each algorithm in scenario 2UAV-1ADS-1, Figure 3 (b) is a bar chart of the number of states covered by each algorithm in scenario 3UAV-1ADS;

[0022] Figure 4 It is a curve chart of the cumulative extrinsic reward and task success rate in scenario 2UAV-1ADS-1 in another embodiment, where Figure 4 (a) is a curve chart of the cumulative extrinsic reward, Figure 4 (b) is a curve chart of the task success rate;

[0023] Figure 5 It is a curve chart of the cumulative extrinsic reward and task success rate in scenario 2UAV-1ADS-2 in another embodiment, where Figure 5 (a) is a curve chart of the cumulative extrinsic reward, Figure 5 (b) is a curve chart of the task success rate;

[0024] Figure 6 Bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in different scenarios in another embodiment, where Figure 6 (a) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 2 UAVs - 1 ADS - 1, Figure 6 (b) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 2 UAVs - 1 ADS - 2, Figure 6 (c) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 3 UAVs - 1 ADS, Figure 6 (d) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 3 UAVs - 2 ADS. Detailed implementation manner

[0025] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0026] In one embodiment, as Figure 1 shown, a multi - UAV collaborative strategy learning method driven by equivalent state novelty value is provided, and the method includes the following steps:

[0027] Step 100: In a multi - UAV mission scenario, determine the joint state space according to the multi - UAV collaborative decision - making model based on the distributed partially observable Markov decision process.

[0028] Specifically, in a multi - UAV mission scenario, according to the multi - UAV collaborative decision - making model based on the distributed partially observable Markov decision process (Dec - POMDP), the joint state space includes the individual states of all UAVs and the states of other environmental elements : .

[0029] Step 102: Construct an equivalent state set for each state according to the joint state space.

[0030] Specifically, assume that there are two joint state variables in the reinforcement learning model corresponding to a multi - UAV collaborative penetration and assault mission, where , .

[0031] When the UAVs in the mission scenario are homogeneous, "UAV 1 is in the individual state , UAV 2 is in the individual state , UAV 3 is in the individual state The situation of "the drone 2 is in the individual state" , the drone 3 is in the individual state , the drone 1 is in the individual state " can be reduced to "one drone is in the individual state , one drone is in the individual state and one drone is in the individual state ".

[0032] Since the mission objectives of the drone group are the same, whether the mission is completed depends on the overall state and has nothing to do with the serial number arrangement of the drones in each individual state. Therefore and are considered equivalent at the mission level. Define the equivalent state and equivalent state set in the multi-drone mission scenario as follows:

[0033] For states and , if there exists a permutation such that and , then and are equivalent states to each other, denoted as .

[0034] For any joint state , define its equivalent state set as the set of all equivalent states, .

[0035] Step 104: Each drone selects an action according to its own strategy and the current observation. After all drones execute the joint action and go through state transition, each drone obtains a first experience tuple containing the observation, action, and external reward.

[0036] Specifically, each drone selects an action according to its own strategy and the current observation ; after all drones execute the joint action , through state transition, each drone obtains a first experience tuple containing the observation, action, and external reward, where

[0037] Step 106: Use the permutation-invariant spatio-temporal feature representation to uniformly characterize the equivalent states of the state after state transition.

[0038] Specifically, the corresponding features of different equivalent states are not the same, resulting in inconsistent estimated novelty values in the input random network distillation method for the prediction network and the target network. Therefore, by aggregating spatio-temporal information in the joint state of class image features, a permutation-invariant spatio-temporal feature representation is adopted to achieve a unified representation of equivalent states.

[0039] Permutation-Invariant (PI) is used to describe the influence of the permutation of each input in a multi-input function on the output.

[0040] Step 108: According to the permutation-invariant spatio-temporal features of the equivalent state, use the random network distillation method to measure the novelty value of the equivalent state, determine the intrinsic reward; and incorporate the intrinsic reward into the first empirical tuple and add it to the experience storage; the extrinsic and intrinsic rewards of all drones are the same for each state transition.

[0041] Specifically, the intrinsic reward method based on the state novelty value is for the state of the (approximate) number of visits is estimated to a negative correlation function of (e.g., or ) is used as the state novelty value metric, constructing an intrinsic reward to drive the agent to continuously explore states with low visit frequencies, performing a traversal search of the state space to mine the sparsely distributed positive rewards. However, in this way, the novelty value metrics of the states in the equivalent state set are inconsistent due to the different number of visits. As long as there are still individual states with relatively high novelty values, the drones will be encouraged to visit , resulting in repeated exploration of equivalent states. Under sparse reward conditions, due to the lack of reward feedback signals during the task process, the drones need to start from the initial state set to explore the state space and search for the state set containing positive rewards to provide effective feedback for policy optimization. In this process, the repeated exploration of equivalent states hinders the search for reward states , thus affecting the multi-drone cooperation policy optimization efficiency under sparse reward conditions. Therefore, the intrinsic reward method based on the state novelty value is improved to avoid the repeated exploration caused by the inconsistent measurement of the novelty values of equivalent states, converting the traversal exploration of the drones in the original joint state space into traversal exploration in the non-equivalent state space .

[0042] The original joint state space corresponding non-equivalent state space is defined as follows: , where is an arbitrary mapping that can convert equivalent states into the same state, i.e., satisfying: , .

[0043] The non - equivalent state space has the following property: for any , there is always holds, , and .

[0044] Compared with the original joint state space , traversing the non - equivalent state space to search for rewards can save of the search cost. To encourage the drone to explore the non - equivalent state space and avoid repeated exploration of equivalent states, the number of times the equivalent state set of each state is visited is measured, replacing the number of times each state is visited alone to construct an intrinsic reward based on the novelty value of equivalent states. , For the state the corresponding element in the non - equivalent state space, the number of visits is counted, and using the negative correlation function of as the novelty value intrinsic reward, encourages the drone to explore the equivalent state sets with low visit frequencies rather than the states with low visit frequencies, thereby driving the drone to traverse and explore the non - equivalent state space.

[0045] Through the method of random network distillation based on permutation - invariant spatio - temporal features, a supervised learning training set of historical state - label vectors is constructed using a fixed target mapping, and the label vector error is predicted based on permutation - invariant spatio - temporal state features to achieve a consistent measure of the novelty value of high - dimensional equivalent states, and end - to - end estimate the overall degree of visit of the equivalent state set; driving the drone exploration with the novelty value of equivalent states.

[0046] The experience sample in the experience storage is , where represents the current observation, represents the selected action, represents the next - moment observation, represents the overall reward, is the policy, represents whether this round is a termination flag.

[0047] Step 110: Weighted - sum the extrinsic reward and the intrinsic reward of the drone i to obtain the overall reward of the drone i ; Loop to execute state transition and overall reward calculation until the decision - making round ends.

[0048] Step 112: Optimize the value network and policy network in the Actor-Critic architecture according to the overall reward and experience storage, using the ACER method.

[0049] Specifically, based on the Actor-Critic structure, each UAV maintains its own policy network and value network , where and are network parameters. For the sake of simplicity in the following text, some are represented by and instead.

[0050] Since the multi-UAV cooperative penetration and assault is a fully cooperative task, the external rewards (i.e., the original reward signals feedback by the environment) of all UAVs are the same for each state transition, that is , , . In addition, to drive the UAVs to cooperate in exploring the state space, all UAVs share the intrinsic reward based on the equivalent state novelty value, that is .

[0051] The UAV uses the weighted sum of the external reward and the intrinsic reward as the overall reward to update the value function and the policy network.

[0052] In the above multi-UAV cooperative strategy learning method driven by the equivalent state novelty value, the method defines a permutation-invariant spatio-temporal feature representation, aggregates the spatio-temporal information in the joint state through class image feature aggregation, and realizes the unified representation of the equivalent state; designs a random network distillation method based on the permutation-invariant spatio-temporal feature to quickly and effectively evaluate the accessed degree of the high-dimensional equivalent state set, realizes the consistent measurement of the equivalent state novelty value, and estimates the overall accessed degree of the equivalent state set end-to-end; drives the UAVs to explore with the equivalent state novelty value; finally, uses the equivalent state novelty value as the shared intrinsic reward to jointly guide the policy update with the original external reward, and this method improves the efficiency of cooperative strategy optimization under sparse reward conditions.

[0053] In one embodiment, the joint state space in step 100 includes the individual states of all UAVs and the states of other environmental factors; the individual state of the UAV is:

[0054] ;

[0055] where are respectively the x coordinate, y coordinate, yaw angle variable, and speed variable of the UAV i ; is the UAV Threatened k Locking or interception duration, is an integer.

[0056] The environmental state includes the radar antenna angle, channel occupancy, and target state.

[0057] In one embodiment, step 102 includes: extracting a key observable part from the state variables in the joint state space to form sub-state variables; determining the sub-states of each state according to the sub-state variables; constructing the equivalent sub-states and equivalent sub-state sets of each sub-state; and when constructing the intrinsic reward based on the novel value of the equivalent state, using the equivalent sub-states and equivalent sub-state sets to replace the equivalent states and equivalent state sets of the corresponding states in the original joint state space.

[0058] In one embodiment, step 106 includes: using a permutation-invariant spatio-temporal feature representation to uniformly characterize the equivalent state of the state after state transition, obtaining the permutation-invariant spatio-temporal feature of the equivalent state; the expression of the permutation-invariant spatio-temporal feature is:

[0059] ;

[0060] where is the equivalent state of the permutation-invariant spatio-temporal feature, n is the total number of UAVs, , , is the number of threatening UAVs i ; , , , , represents the range of the rectangular task area; is used to indicate the position of the UAV. If the UAV is located in the grid , then , otherwise ; compared with characterizes the spatial information in the state, then embeds the time information in the state into , is the UAV threatened locking or interception time of the linear mapping, , and are respectively the mapping values of the upper and lower bounds of the locking or interception duration of the UAV threatened k .

[0061] Specifically, when measuring the state novelty value based on the random network distillation method, the feature forms of the state variables input to the prediction network and the target network are one of the keys to novelty value measurement. The state variables As high-dimensional real-valued variables, they can be directly input into the prediction network and the target network after normalization processing, or the values of each dimension can be transformed through a method similar to one-hot encoding and then concatenated as input features. The prediction network and the target network use a fully-connected layer as the input layer to process the features. However, there is a significant problem with the above feature construction, that is, different states with the same equivalence correspond to different features, resulting in inconsistent estimated novelty values after input into the prediction network and the target network.

[0062] Therefore, a method for constructing permutation-invariant spatio-temporal features (PI-STF) of states is proposed. Permutation invariance (PI) is used to describe the influence of the permutation of each input in a multi-input function on the output. For the function , , if the permutation of the input elements is permuted without changing the output, that is, for any permutation matrix , there is established, then the function satisfies permutation invariance.

[0063] The permutation-invariant spatio-temporal feature maps the state variable into an image-like feature with a size of , and the number of channels is , denoted as represents the range of the rectangular task area, and through discrete intervals are transformed into grids, is the number of threats, represents rounding up.

[0064] The permutation-invariant spatio-temporal feature is specifically defined as shown in the expression of the above permutation-invariant spatio-temporal feature.

[0065] Compared with directly using , mapping to and then constructing the permutation-invariant spatio-temporal feature has two advantages. On the one hand, numerical normalization processing is performed on this variable, and on the other hand, the ambiguity in the feature construction process is avoided. If directly using , when , it is impossible to distinguish "drone Located in the grid , but not tracked and intercepted by the threat system , that is but " and "drone Not located in the grid , that is " are two completely different situations; while using ensures that once the drone is located in the grid , regardless of whether it is currently tracked and intercepted is greater than 0, effectively distinguishing the above two situations.

[0066] In one embodiment, step 108 includes: constructing a random distillation network; the random distillation network includes a target network and a prediction network, the target network and the prediction network have the same structure but different initialized network parameters; the parameters of the target network remain fixed after self-initialization, and the parameters of the prediction network are updated according to the mean square error between its own output and the output of the target network using the gradient direction; input the permutation-invariant spatio-temporal features of the equivalent state into the target network and the prediction network, and determine the equivalent state novelty value metric according to the outputs of the target network and the prediction network as:

[0067] ;

[0068] Construct the intrinsic reward based on the equivalent state novelty value as:

[0069] ;

[0070] wherein is the intrinsic reward based on the equivalent state novelty value, are the state at the current moment, the action at the current moment, and the state at the next moment respectively, is the equivalent state novelty value metric of the state at the next moment.

[0071] Specifically, in the scenario of multi-drone collaborative penetration and assault mission, the joint state is , where the individual state of drone is , and the environmental state includes non-real-time observable factors such as the radar antenna angle and the channel occupancy situation. To improve the sparse reward search efficiency and method feasibility, extract the key observable part from the state variable to form the sub-state variable , reducing the size of the search space. Under threat countermeasures, multiple UAVs need to complete collaborative penetration and assault tasks through close spatio-temporal coordination. Therefore, intrinsic rewards are needed to encourage the exploration of diverse UAV spatio-temporal coordination patterns, so as to discover effective collaborative strategies. Compared with relative position information and the duration of being locked or intercepted by the countermeasure system, UAV attitude information and speed information are relatively less important when exploring diverse spatio-temporal coordination patterns.

[0072] For this reason, define the sub-state variable Ignore the yaw angle variable when and the speed variable , , , .

[0073] For the sub-states and , if there exists a permutation such that , then and are equivalent sub-states, denoted as .

[0074] For any sub-state , define its equivalent state set as the set of all equivalent sub-states, , analogous to equivalent states and equivalent state sets, define equivalent sub-states and equivalent sub-state sets based on .

[0075] Use the sub-state to replace to count the number of times the equivalent state set is visited , which is used to construct an intrinsic reward based on the novelty value of equivalent states. , when the state is a low-dimensional variable, the number of UAVs is small, and the state space is small, it can be calculated by counting, and can be simply accumulated based on the equivalent state set on this basis. As the dimension of the state variable increases and the number of UAVs increases, the size of the state space increases rapidly. At this time, counting the number of times each state is visited based on the counting method requires a large amount of memory, and the size of the equivalent state set increases with the number of permutations and combinations of UAV serial numbers. Calculating and accumulating the number of times each element in it is visited separately is too cumbersome.

[0076] Based on the above analysis, aiming at the high-dimensional continuous-discrete hybrid state space, an end-to-end equivalent state novelty value consistency measurement method is proposed to construct an intrinsic reward to encourage the UAV to efficiently explore the non-equivalent state space and improve the collaborative strategy learning efficiency under sparse reward conditions. Using the sub-state variable to replace the state variable for the design of the intrinsic reward. Without affecting the elaboration and analysis of the method, for the sake of simplicity, the state variable related to the novelty value measurement in the following text is default to refer to . Random Network Distillation (RND) is a method for measuring the novelty value of high-dimensional states or observations. The random network distillation method estimates the novelty value of a state or observation based on the following fact: in the supervised learning paradigm, if a test sample has more similar trained samples, the corresponding prediction error is smaller.

[0077] The random network distillation method contains two neural networks: the target network and the predictor network. The parameters of the target network are fixed after random initialization, and the predictor network is trained with the experienced states or observations as samples. Taking the measurement of the novelty value of the observation variable as an example, the target network and the predictor network map the observation into features of the same dimension. The predictor network takes the output of the target network as the target value, and based on the observations experienced by the agent, constructs a training sample set containing observations-labels, and updates its own network parameters by minimizing the mean square error . Under this training mechanism, for a given observation variable , if it is accessed more frequently, the number of times the predictor network parameters are updated based on its similar observation samples is also more, and the corresponding prediction error is smaller; conversely, if the observation is accessed less frequently, the number of times the predictor network parameters are updated based on its similar observation samples is also less, and the corresponding prediction error is larger. Therefore, the prediction error can be used to measure the novelty value of the observation . If the inputs of the target network and the predictor network are changed from observations to states, the novelty value of the state variable can be measured. Compared with other high-dimensional state or observation novelty value measurement methods, such as the method based on density estimation, the random network distillation method has the significant advantages of wide application range, simple structure, and small computational amount, and only needs to additionally train the predictor network.

[0078] The Permutation-Invariant Spatio-Temporal Feature-based Random Networks Distillation (PI-STF-RND) method aims to achieve an end-to-end high-dimensional equivalent state novelty metric and further construct an intrinsic reward to guide the efficient exploration of drones. As a type of image-like feature, when permutation-invariant spatio-temporal features are used for state novelty metric, the target network in the random network distillation method based on permutation-invariant spatio-temporal features and the prediction network take the Convolution Layer as the input layer, and use multiple convolutional kernels to extract and fuse features of permutation-invariant spatio-temporal features from different perspectives. Then, the output of the convolutional layer is processed by the Flatten operation and input into the subsequent fully connected layer. , where is the dimension of the final output of the network in the random network distillation method.

[0079] Given the state , the target network and the prediction network respectively output and . Both the target network and the prediction network in the constructed random network distillation contain a convolutional layer, a Rectified Linear Unit (ReLU), a flattening layer, and a fully connected layer. Among them, the convolutional kernel size of the convolutional layer is 5×5, the number of convolutional kernels is 32, the stride is 2, and the padding size is 2. The size of the fully connected layer FC1 is 512, and the size of the fully connected layer FC2 is 10, that is, N = 10. The target network and the prediction network have the same structure, but the initialized network parameters are different. The parameters of the target network remain fixed after self-initialization, while the parameters of the prediction network need to be updated according to the mean square error between its own output and the output of the target network. The mean square error between its own output and the output of the target network is:

[0080] ;

[0081] The parameters of the prediction network are updated by the gradient method, aiming to achieve ; the update process is:

[0082] ;

[0083] ;

[0084] where is the set of states experienced in this decision-making episode.

[0085] After each episode ends, the set generated during the multi-UAV decision-making process is used as a training sample containing features-labels to update the parameters of the prediction network. In the supervised learning paradigm, the neural network has a better fitting effect on training samples with higher occurrence frequencies.

[0086] Therefore, when the new state variable is converted into a permutation-invariant spatio-temporal feature and then input into the PI-STF-RND network, the difference between the output of the prediction network and the output of the target network can reflect the number of times the permutation-invariant feature or its approximate feature appears in the training samples . The smaller the difference, the larger.

[0087] According to having permutation invariance, , for all it holds that . The closer is to , the more times the equivalent state of is visited during the decision-making process, and the lower the overall novelty value of its equivalent state set

[0088] .;

[0089] The PI-STF-RND method uses a permutation-invariant feature representation to map equivalent states to the same spatio-temporal features, and then estimates the occurrence frequencies of each spatio-temporal feature in the training samples through the random network distillation method, thereby directly reflecting the access degree of the corresponding equivalent state set, saving the cumbersome operation of separately estimating and accumulating the access times of each equivalent state, and realizing the fast and effective evaluation of the novelty value of high-dimensional equivalent states.

[0090] To drive the UAV to actively explore the state space, the novelty value measured by the PI-STF-RND method is used to construct the intrinsic reward of the equivalent state novelty value:

[0091] .;

[0092] After each state transition, the novelty value measurement of the next moment state based on PI-STF-RND , as an intrinsic reward, all drones share this intrinsic reward. Since it actually measures not only the degree of access of the state but the degree of access of the entire equivalent state set , it avoids repeated access to equivalent states caused by inconsistent measurement of novelty values, effectively encourages drones to explore non-equivalent states with low access frequencies, and improves the search efficiency for sparse positive rewards.

[0093] In one embodiment, both the target network and the prediction network include: a convolutional layer, a rectified linear unit, a flattening layer, and two fully connected layers.

[0094] In one embodiment, step 110 includes: weighting and summing the extrinsic reward and the intrinsic reward of the drone i to obtain the total reward of the drone i as:

[0095] ;

[0096] where is the total reward of the drone i , and are the weights of the extrinsic reward and the intrinsic reward based on the novelty value of the equivalent situation respectively, is the extrinsic reward, is the intrinsic reward based on the novelty value of the equivalent situation.

[0097] In one embodiment, the second experience tuple in the experience store is: , where is the total reward, is the policy, is the current observation, is the action, is the observation at the next moment, is the extrinsic reward, is the flag indicating whether the episode terminates; step 112 includes: each drone randomly samples from the experience store, and uses the sampled experience data and the total reward to update the value network and the policy network using the ACER method; the ACER method includes: a policy evaluation link and a policy improvement link; in the policy evaluation link: the value network parameters are updated by minimizing the mean square error; the update expression of the value network parameters is:

[0098] ;

[0099] ;

[0100] ;

[0101] Among them, is the mean square error of the value network, is the parameter of the value network, is the learning rate of the UAV value network, is the value estimation of the current policy based on the Retrace method, is the overall reward, is the value network, respectively represent the current observation and action of the UAV i , is the UAV i parameter of the policy network of, is a hyperparameter, is the truncated importance sampling coefficient, , , is the UAV generates experience The historical policy at that time is recorded as the behavioral policy.

[0102] In the policy improvement step, the parameters of the UAV policy network are updated using the ACER policy gradient; the update expression of the policy network parameters is:

[0103] ;

[0104] ;

[0105] Among them, is the parameter of the policy network, is the learning rate of the UAV policy network, is the ACER policy gradient, is the current policy of the UAV, is the value of the current observation, is the cost.

[0106] In one embodiment, the overall framework of multi-UAV policy learning based on the equivalent state novelty value intrinsic reward includes: UAV-environment interaction and UAV policy training:

[0107] (1) UAV-environment interaction process:

[0108] Step 1: Each UAV selects an action according to its own policy and the current observation ;

[0109] Step 2: All UAVs execute the joint action After that, through state transition, each UAV obtains an experience tuple containing observations, actions, and external rewards. Indicates whether the round terminates;

[0110] Step 3: Then, use the random network distillation method based on permutation-invariant spatio-temporal features to measure the novelty value. Calculate the corresponding intrinsic reward. The experience tuple combining the intrinsic reward and the external reward is added to the experience storage, and each UAV performs the above operations after each state transition.

[0111] (2) UAV policy training process:

[0112] Step 1: After the decision-making round ends, the prediction network of the random network distillation method based on permutation-invariant spatio-temporal features is updated using all the state variables experienced in this round.

[0113] Step 2: Each UAV randomly samples from the experience storage and updates its own policy and value function . The UAV policy is optimized based on the ACER method, which includes a policy evaluation and a policy improvement link, and the UAV value network and policy network are updated respectively.

[0114] In the policy evaluation link, the value network parameters are updated by minimizing the mean square error .

[0115] In the policy improvement link, the parameters of the UAV policy network are updated using the ACER policy gradient . .

[0116] If , , otherwise . The ACER method is a typical off-policy training method that trains the current policy by sampling previous experience samples from the experience cache.

[0117] In the policy evaluation link and the policy improvement link, the value network parameters and the policy network parameters are both updated according to the sampled historical experience samples. Given the historical decision-making trajectory of the UAV sampled from the experience cache , , the gradient of the corresponding value network parameters is calculated as follows:

[0118] ;

[0119] Value network parameters are updated in the same way as shown in the above update expression of the value network parameters.

[0120] The corresponding gradient of the policy network parameters is , and is updated in the same way as shown in the above update expression of the policy network parameters.

[0121] Each UAV interacts with the environment by executing the current policy, and then continuously optimizes the policy according to the above policy update rules using the interaction experience, repeating this process until the maximum number of training rounds is reached.

[0122] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0123] In a verification embodiment, with the background that a UAV task unit arrives at the periphery of the defense area and identifies the protected target under multiple threats, a multi-UAV cooperative penetration and assault task scenario under threat countermeasures is set. Under the given threats and UAV system performance parameters, the deployment plan in the task scenario satisfies:

[0124] (1) The protected target is within the effective defense range that a single UAV cannot enter, and the task needs to be completed by the cooperation of the UAV unit;

[0125] (2) The UAV unit needs to continuously cooperate closely in space and time to complete the task, which requires a high level of cooperation, resulting in greater difficulty in policy training under sparse reward conditions.

[0126] Based on the different relative positions of the target and the threat system, the following four scenarios are set for algorithm verification: 2UAV-1ADS-1, 2UAV-1ADS-2, 3UAV-1ADS, and 3UAV-2ADS. The scenario names include the number of UAVs and the number of threat systems in the mission area. As Figure 2 shown, where Figure 2(a) is Scenario 2UAV-1ADS-1, Figure 2 (b) is Scenario 2UAV-1ADS-2, Figure 2 (c) is Scenario 3UAV-1ADS, Figure 2 (d) is Scenario 3UAV-2ADS. Both Scenario 2UAV-1ADS-1 and 2UAV-1ADS-2 include 2 UAVs and 1 threat in a 15km×15km mission area, and the distance from the target PA to the threat system ADS1 is the same. However, in Scenario 2UAV-1ADS-2, the target is located northwest of the threat system. Due to the mission area restriction, the point closest to the target outside the red line range deviates significantly from the UAV incoming direction, that is, the range of penetration points that meet the conditions is narrow. Therefore, the mission difficulty is greater than that of Scenario 2UAV-1ADS-1; Scenario 3UAV-1ADS further reduces the distance between the target and the threat system, and the mission difficulty is further increased. Correspondingly, the scale of the UAV detachment is expanded to 3; Scenario 3UAV-2ADS tests the learning of the multi-UAV cooperation strategy when the target is in the overlapping defense range of multiple threat systems.

[0127] The multi-UAV cooperation strategy learning method proposed in this application based on the equivalent state novelty value intrinsic reward is denoted as PI-STF-RND. In addition, two benchmark methods are selected for comparison, denoted as No-Intrinsic-Reward and ECV-RND methods respectively.

[0128] No-Intrinsic-Reward: No additional intrinsic reward is calculated, and the UAV strategy is updated only with extrinsic rewards based on the ACER algorithm; under this method, the UAV randomly explores the state space according to the current strategy to search for positive rewards. The overall reward is set as: .

[0129] ECV-RND: Each part of the state variable is encoded by ECV (Efficient Coordinated Vector, ECV) and then directly concatenated, and used as the state feature to be input into the RND network. The situation novelty value estimated by ECV-RND is used as the intrinsic reward ECV-RND, and the weighted sum of the ECV-RND intrinsic reward and the extrinsic reward is used to update the UAV policy based on the ACER algorithm. ECV encoding is an extension of the classical one-hot encoding method, which introduces weights to represent the distance between continuous real numbers and the surrounding discrete real numbers, so as to realize the representation of continuous variables 1. Different from PI-STF-RND, since the state feature is a one-dimensional array at this time, the input layers of the target network and the prediction network in ECV-RND are set as fully connected layers. Since the state feature based on ECV encoding and concatenation does not have permutation invariance, and the features corresponding to different equivalent states are different, ECV-RND does not consider the overall access degree of the equivalent state set, but only considers the given state itself. The overall reward is set as: .

[0130] By comparing PI-STF-RND and No-Intrinsic-Reward, the impact of the intrinsic reward based on the equivalent state novelty value on the UAV state space exploration efficiency and the UAV cooperative strategy learning efficiency can be analyzed; by comparing PI-STF-RND and ECV-RND, it can be analyzed whether the intrinsic reward based on the equivalent state novelty value can effectively reduce the repeated access to equivalent states compared with the general state novelty value intrinsic reward, so as to improve the cooperative strategy learning efficiency.

[0131] Table 1 Reward setting parameters of UAV policy learning algorithm

[0132]

[0133] The UAV policy network consists of a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, a fully connected layer (with 7 layers), and a softmax function, from the input layer to the output layer. The UAV action value network consists of a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 7 layers), from the input layer to the output layer. The PI-STF-RND target network and prediction network consist of a convolutional layer with 32 5×5 convolutional kernels, a rectified linear unit, a flattening layer, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 10 layers). The ECV-RND target network and prediction network consist of a fully connected layer (with m layers, where m is a positive integer), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 10 layers). In the 2UAV-1ADS-1, 2UAV-1ADS-2, and 3UAV-1ADS scenarios, m is 2048, and in the 3UAV-2ADS scenario, m is 3872.

[0134] In the multi-UAV collaborative exploration experiment, the exploration effects of the exploration mechanism driven by the intrinsic reward of the PI-STF-RND equivalent state novelty value, the exploration mechanism driven by the intrinsic reward of the ECV-RND state novelty value, and the No-Intrinsic-Reward random exploration mechanism were compared by excluding the influence of external rewards.

[0135] Table 2 Comparison of the exploration effects of each algorithm in the 2UAV-1ADS-1 and 3UAV-1ADS scenarios

[0136]

[0137] The bar charts of the number of covered states of each algorithm in the 2UAV-1ADS-1 and 3UAV-1ADS scenarios are as Figure 3 shown, where Figure 3 (a) is the bar chart of the number of covered states of each algorithm in the 2UAV-1ADS-1 scenario, Figure 3 (b) is the bar chart of the number of covered states of each algorithm in the 3UAV-1ADS scenario.

[0138] Table 2 and Figure 3The data in was obtained after each test algorithm trained for 2.0×104 rounds in the scenario of 2UAV-1ADS-1 and 5.0×104 rounds in the scenario of 3UAV-1ADS. From the experimental results, it can be seen that the number of original state coverages and non-equivalent state coverages of the PI-STF-RND algorithm are significantly higher than those of the ECV-RND and No-Intrinsic-Reward algorithms in both scenarios. The number of original state coverages and non-equivalent state coverages of the No-Intrinsic-Reward algorithm are the least, indicating that the corresponding random exploration mechanism has the lowest efficiency. In the scenario of 2UAV-1ADS-1, the number of original state coverages of PI-STF-RND is about 17.6% higher than that of the ECV-RND algorithm and about 148.1% higher than that of the No-Intrinsic-Reward algorithm. The number of non-equivalent state coverages is about 30.1% higher than that of the ECV-RND algorithm and about 165.3% higher than that of the No-Intrinsic-Reward algorithm. In the scenario of 3UAV-1ADS, the number of original state coverages of PI-STF-RND is about 19.8% higher than that of the ECV-RND algorithm and about 47.4% higher than that of the No-Intrinsic-Reward algorithm. The number of non-equivalent state coverages is about 37.5% higher than that of the ECV-RND algorithm and about 70.6% higher than that of the No-Intrinsic-Reward algorithm. The experimental results of multi-UAV cooperative exploration show that the PI-STF-RND algorithm effectively reduces the repeated access of UAVs to the equivalent state set through the intrinsic reward of equivalent state novelty value, and improves the exploration efficiency of the original state space and non-equivalent state space (RQ1).

[0139] Figure 4 shows the performance of each algorithm in the scenario of 2UAV-1ADS-1, where PI-STF-RND is significantly better than other algorithms in multiple evaluation metrics such as cumulative extrinsic reward, task success rate, and the number of training rounds required to reach 60%, 80%, and 90% success rates. As Figure 4 (a) shows, the cumulative extrinsic reward curve of PI-STF-RND starts to rise at about 2.0×103 training rounds and finally converges basically at 7.0×103 training rounds, and is higher than the cumulative extrinsic reward curves of other algorithms throughout the training process. As Figure 4As shown in (b), after training for about 6.0×10³ rounds, the task success rate of the PI-STF-RND algorithm approaches 100%. At this time, the task success rate of the ECV-RND algorithm is only about 60%, and the task success rate of the No-Intrinsic-Reward algorithm is only about 25%. According to the experimental results, since the tasks in Scenario 2UAV-1ADS-1 are relatively simple, the No-Intrinsic-Reward algorithm can also achieve an accumulated extrinsic reward of about 0.5 after training for 1.0×10⁴ rounds through random exploration, and the task success rate exceeds 70%. As Figure 4 As shown in (a), in Scenario 2UAV-1ADS-1, when the task success rate reaches 60%, 80%, and 90%, the number of training rounds required by the PI-STF-RND algorithm is reduced by 23.7%, 34.2%, and 34.5% compared with the ECV-RND algorithm, and is reduced by more than 43.75%, 50.0%, and 45.0% compared with the No-Intrinsic-Reward algorithm.

[0140] As Figure 5 shown, compared with Scenario 2UAV-1ADS-1, Scenario 2UAV-1ADS-2 is more challenging. Therefore, the No-Intrinsic-Reward algorithm can achieve a task success rate of more than 70% after 1.0×10⁴ training rounds in Scenario 2UAV-1ADS-1 through random exploration. However, in Scenario 2UAV-1ADS-2, even after 2.0×10⁴ rounds of training, this algorithm has not learned the cooperative strategy for successful penetration and assault. This result is consistent with the analysis of Scenario 2UAV-1ADS-1 and Scenario 2UAV-1ADS-2 during the test scenario setting. The change in the relative position of the target PA and the threat system ADS1 in the scenario has a greater impact on the task difficulty. Despite the increase in task difficulty, the PI-STF-RND algorithm is significantly better than other comparison algorithms in multiple evaluation indicators such as accumulated extrinsic reward, task success rate, and the number of training rounds required to reach 60%, 80%, and 90% success rates in Scenario 2UAV-1ADS-2. As Figure 5 shown in (a), during the entire training process, the accumulated extrinsic reward curve of the PI-STF-RND algorithm is always above that of the ECV-RND algorithm. As Figure 5 shown in (b), the ECV-RND algorithm reaches a task success rate of about 83% after 2.0×10⁴ rounds of training. The PI-STF-RND algorithm approaches an 80% task success rate after 1.0×10⁴ training rounds and reaches a task success rate of about 95% after 2.0×10⁴ training rounds. The bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in different scenarios is as Figure 6 shown, where Figure 6(a) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 2UAV-1ADS-1. Figure 6 (b) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 2UAV-1ADS-2. Figure 6 (c) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 3UAV-1ADS. Figure 6 (d) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 3UAV-2ADS. As Figure 6 shown in (a), in the scenario 2UAV-1ADS-1, when the task success rate of the strategy reaches 60%, 80%, and 90%, the number of training rounds required by the PI-STF-RND algorithm is reduced by more than 34.3%, 47.2%, and 42.0% compared with the ECV-RND algorithm, and by more than 54.0%, 48.5%, and 42.0% compared with the No-Intrinsic-Reward algorithm.

[0141] The simulation results show that in the test scenario, compared with the existing algorithms, the exploration efficiency of this method for the state space and the non-equivalent state space is improved by more than 17.6% and 30.1% respectively, indicating that this method effectively reduces the repeated access of the UAV to the equivalent state during the exploration process.

[0142] In one embodiment, a multi-UAV cooperative strategy learning device driven by the equivalent state novelty value is provided, including: an equivalent state set construction module, a state transition module, an intrinsic reward determination module, an overall reward determination module, and a UAV policy optimization module, where:

[0143] The equivalent state set construction module is used to determine the joint state space according to the multi-UAV collaborative decision-making model based on the distributed partially observable Markov decision process in the multi-UAV task scenario; and construct the equivalent state set of each state according to the joint state space.

[0144] The state transition module is used for each UAV to select an action according to its own policy and the current observation. After all UAVs execute the joint action and go through state transition, each UAV obtains a first experience tuple including the observation, action, and extrinsic reward; the extrinsic rewards of all UAVs are the same for each state transition.

[0145] The intrinsic reward determination module is used to uniformly represent the equivalent state of the state after state transition by using the permutation-invariant spatio-temporal feature representation; measure the equivalent state novelty value by using the random network distillation method according to the permutation-invariant spatio-temporal feature of the equivalent state to determine the intrinsic reward; and add the intrinsic reward to the experience storage after integrating it into the first experience tuple; all UAVs share the intrinsic reward.

[0146] The overall reward determination module is used to perform weighted summation on the extrinsic reward and intrinsic reward of the drone i to obtain the overall reward of the drone i ; continuously execute state transition and overall reward calculation until the decision round ends.

[0147] The drone policy optimization module is used to optimize the value network and policy network in the Actor-Critic architecture by using the ACER method according to the overall reward and experience storage.

[0148] In one embodiment, the combined state space in the equivalent state set construction module includes the individual states of all drones and the states of other environmental factors; the individual state of the drone is as shown in the expression of the individual state of the above-mentioned drone ; the environmental state includes the radar antenna angle, channel occupancy, and target state.

[0149] In one embodiment, the equivalent state set construction module is further configured to extract the key observable part from the state variables in the combined state space to form sub-state variables; determine the sub-states of each state according to the sub-state variables; construct the equivalent sub-states and equivalent sub-state sets of each sub-state; when constructing the intrinsic reward based on the novelty value of the equivalent state, use the equivalent sub-states and equivalent sub-state sets to replace the equivalent states and equivalent state sets of the corresponding states in the original combined state space.

[0150] In one embodiment, the intrinsic reward determination module is further configured to use permutation-invariant spatio-temporal feature representation to uniformly represent the equivalent state of the state after state transition to obtain the permutation-invariant spatio-temporal feature of the equivalent state; the permutation-invariant spatio-temporal feature of the equivalent state is as shown in the expression of the permutation-invariant spatio-temporal feature of the above-mentioned equivalent state.

[0151] In one embodiment, the intrinsic reward determination module is further configured to construct a stochastic distillation network; the stochastic distillation network includes a target network and a prediction network, the target network and the prediction network have the same structure but different initialized network parameters; the parameters of the target network remain fixed after self-initialization, and the parameters of the prediction network are updated according to the mean square error between its own output and the output of the target network using the gradient direction; input the permutation-invariant spatio-temporal feature of the equivalent state into the target network and the prediction network, and determine the novelty value metric of the equivalent state according to the outputs of the target network and the prediction network; the novelty value metric of the equivalent state is as described in the expression of the novelty value metric of the above-mentioned equivalent state; construct the intrinsic reward based on the novelty value of the equivalent state, and the intrinsic reward based on the novelty value of the equivalent state is as shown in the expression of the intrinsic reward based on the novelty value of the above-mentioned equivalent state.

[0152] In one embodiment, both the target network and the prediction network include: a convolutional layer, a rectified linear unit, a flattening layer, and two fully connected layers.

[0153] In one embodiment, the overall reward determination module is configured to perform a weighted sum of the extrinsic reward and the intrinsic reward of the drone i to obtain the overall reward of the drone i The overall reward is as described in the expression of the overall reward of the above-mentioned drone i

[0154] In one embodiment, the second experience tuple in the experience store is: , where is the overall reward, is the policy, is the current observation, is the action, is the next moment observation, is the extrinsic reward, is the flag indicating whether the round terminates; the drone policy optimization module is further configured to randomly sample from the experience store for each drone, and use the sampled experience data and the overall reward to update the value network and the policy network by using the ACER method; the ACER method includes: a policy evaluation link and a policy improvement link; in the policy evaluation link: update the value network parameters by minimizing the mean square error; the update manner of the value network parameters is as described in the above update expression of the value network parameters.

[0155] In the policy improvement link, use the ACER policy gradient to update the parameters of the drone policy network; the update manner of the policy network parameters is as described in the above update expression of the policy network parameters.

[0156] For the specific limitations of the multi-drone cooperation strategy learning device driven by the equivalent state novelty value, reference can be made to the limitations of the multi-drone cooperation strategy learning method driven by the equivalent state novelty value in the above text, which will not be elaborated here. Each module in the above-mentioned multi-drone cooperation strategy learning device driven by the equivalent state novelty value can be implemented in whole or in part through software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0158] ​The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A multi-UAV cooperative strategy learning method driven by equivalent state novelty value, characterized in that The method includes: In a multi-UAV mission scenario, determine the joint state space according to a multi-UAV collaborative decision-making model based on a distributed partially observable Markov decision process; Construct an equivalent state set for each state according to the joint state space; Each UAV selects an action according to its own policy and current observation. After all UAVs execute the joint action and undergo state transition, each UAV obtains a first experience tuple containing observation, action, and extrinsic reward; Use permutation-invariant spatio-temporal feature representation to uniformly represent the equivalent state of the state after state transition, and the permutation-invariant spatio-temporal feature of the equivalent state is: ; Among them, is the permutation-invariant spatio-temporal feature of the equivalent state , n is the total number of drones , , is the number of drones of the threatening drones ; , , , , represents the range of the rectangular mission area; is used to indicate the position of the drone. If the drone is located in the grid , then , otherwise ; is the linear mapping of the duration when the drone is threatened to be locked or intercepted , and are the mapping values of the upper and lower bounds of the duration when the drone is threatened to be locked or intercepted respectively; Measure the novelty value of the equivalent state using the random network distillation method according to the permutation-invariant spatio-temporal feature of the equivalent state, and determine the intrinsic reward; and incorporate the intrinsic reward into the first experience tuple and add it to the experience storage; the extrinsic reward and the intrinsic reward of all UAVs are the same for each state transition; the novelty value measurement of the equivalent state is: ; Among them, is the equivalent state novelty value metric, is the output of the prediction network, are the parameters of the prediction network, is the output of the target network, is the equivalent state 's permutation-invariant spatio-temporal feature; Sum the extrinsic and intrinsic rewards of the drone i with weights to obtain the total reward of the drone i ; repeatedly execute state transition and total reward calculation until the decision round ends; Optimize the value network and policy network in the Actor-Critic architecture using the ACER method according to the overall reward and the experience storage.

2. The multi-UAV collaborative strategy learning method driven by equivalent state novelty value according to claim 1, characterized in that The combined state space includes the individual states of all unmanned aerial vehicles and the states of other environmental factors; the unmanned aerial vehicle has the following individual state: ; Among them, are respectively the x-coordinate, y-coordinate, yaw angle variable, and speed variable of the UAV ; is the duration during which the UAV is threatened locked or intercepted, , and is the number of UAVs threatening the UAV The environmental state includes radar antenna angle, channel occupancy, and target state.

3. The multi-UAV collaborative strategy learning method driven by the equivalent state novelty value according to claim 1, wherein Construct an equivalent state set for each state according to the joint state space, including: Extract the key observable part from the state variables in the joint state space to form sub-state variables; Determine the sub-state of each state according to the sub-state variables; Construct the equivalent sub-state and equivalent sub-state set for each sub-state; When constructing the intrinsic reward based on the novelty value of the equivalent state, use the equivalent sub-state and equivalent sub-state set to replace the equivalent state and equivalent state set of the corresponding state in the original joint state space.

4. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1, characterized in that Measure the novelty value of the equivalent state using the random network distillation method according to the permutation-invariant spatio-temporal feature of the equivalent state, and determine the intrinsic reward, including: Construct a random distillation network; the random distillation network includes a target network and a prediction network, the target network and the prediction network have the same structure but different initialized network parameters; the parameters of the target network remain fixed after self-initialization, and the parameters of the prediction network are updated according to the mean square error between its own output and the output of the target network using the gradient direction; Input the permutation-invariant spatio-temporal feature of the equivalent state into the target network and the prediction network, and determine the novelty value measurement of the equivalent state according to the outputs of the target network and the prediction network; Construct the intrinsic reward based on the novelty value of the equivalent state as: ; Among them, is the intrinsic reward based on the equivalent state novelty value, are the state at the current moment, the action at the current moment, and the state at the next moment respectively, is the equivalent state novelty value metric of the state at the next moment.

5. The multi-UAV collaborative strategy learning method driven by equivalent state novelty value according to claim 4, wherein Both the target network and the prediction network include: a convolutional layer, a rectified linear unit, a flattening layer, and two fully connected layers.

6. The multi-UAV collaborative strategy learning method driven by equivalent state novelty value according to claim 1, characterized in that Sum the extrinsic rewards and intrinsic rewards of the drone i with weights to obtain the total reward of the drone i as follows: ; Among them, is the overall reward for the drone , and are the weights of the extrinsic reward and the intrinsic reward based on the equivalent situation novelty value respectively, is the extrinsic reward, and is the intrinsic reward based on the equivalent situation novelty value.

7. The multi-UAV cooperative strategy learning method driven by the equivalent state novelty value according to claim 1, characterized in that The second experience tuple in the experience storage is as follows: , where is the overall reward, is the policy, is the current observation, is the action, is the observation at the next moment, is the extrinsic reward, is the termination flag for this round; Optimize the value network and policy network in the Actor-Critic architecture using the ACER method according to the overall reward and the experience storage, including: Each UAV randomly samples from the experience storage, and uses the sampled experience data and the overall reward to update the value network and policy network using the ACER method; the ACER method includes: a policy evaluation link and a policy improvement link; In the policy evaluation phase: the value network parameters are updated by minimizing the mean squared error; the update expression of the value network parameters is: ; ; ; Among them, is the mean squared error of the value network, are the parameters of the value network, is the learning rate of the UAV value network, is the value estimation of the current policy based on the Retrace method, is the total reward, is the value network, are respectively the current observation and action of the UAV , is the UAV parameters of the policy network, is a hyperparameter, is the truncated importance sampling coefficient, , , is the UAV generates experience when the historical policy, denoted as the behavior policy; In the policy improvement phase, the parameters of the UAV policy network are updated using the ACER policy gradient; the update expression of the policy network parameters is: ; ; Among them, are the parameters of the policy network, is the learning rate of the UAV policy network, is the ACER policy gradient, is the current policy of the UAV, is the value of the current observation, is the cost.

8. A multi-UAV collaborative strategy learning device driven by an equivalent state novelty value, characterized in that The device includes: An equivalent state set construction module, which is used to determine the joint state space according to the multi-UAV collaborative decision-making model based on the distributed partially observable Markov decision process in the multi-UAV mission scenario; and construct an equivalent state set for each state according to the joint state space; A state transition module, which is used for each UAV to select an action according to its own policy and the current observation. After all UAVs execute the joint action and go through state transition, each UAV obtains a first experience tuple containing the observation, action, and extrinsic reward; the extrinsic reward of all UAVs is the same for each state transition; An intrinsic reward determination module, which is used to uniformly represent the equivalent state of the state after state transition using the permutation-invariant spatio-temporal feature representation to obtain the permutation-invariant spatio-temporal feature of the equivalent state; measure the novelty value of the equivalent state using the random network distillation method according to the permutation-invariant spatio-temporal feature of the equivalent state to determine the intrinsic reward; and add the intrinsic reward to the first experience tuple and then add it to the experience storage; all UAVs share the intrinsic reward; where the permutation-invariant spatio-temporal feature of the equivalent state is: ; Among them, is the permutation-invariant spatio-temporal feature of the equivalent state ; n is the total number of drones, , , is the number of threatening drones ; , , , , represents the range of the rectangular mission area; is used to indicate the position of the drone. If the drone is located in the grid , then , otherwise ; is the linear mapping of the duration when the drone is threatened and locked or intercepted ; , and are the mapping values of the upper and lower bounds of the duration when the drone is threatened and locked or intercepted, respectively; The novelty value measurement of the equivalent state is: ; Among them, is the equivalent state novelty value metric, is the output of the prediction network, are the parameters of the prediction network, is the output of the target network, is the equivalent state of the permutation-invariant spatio-temporal feature; Overall Reward Determination Module, which is used to perform weighted summation of the extrinsic reward and the intrinsic reward of the drone i to obtain the overall reward of the drone i ; repeatedly execute state transition and overall reward calculation until the decision-making round ends; A UAV policy optimization module, which is used to optimize the value network and policy network in the Actor-Critic architecture using the ACER method according to the total reward and the experience storage.

Citation Information

Patent Citations

  • Bit rate optimization algorithm based on DDPG for energy collectible communication

    CN109548044A

  • Multi-energy cooperative control method, device and equipment of power grid system and storage medium

    CN118040788A