Multi-unmanned aerial vehicle cooperation strategy learning method and device driven by novel value in equivalent state

By using the permutation of invariant spatiotemporal characteristics and random network distillation method to measure the novel value of equivalent state in multi-drone collaborative strategy learning, the problem of drones repeatedly exploring equivalent state under sparse reward conditions is solved, and the strategy learning efficiency is improved.

CN120122701AActive Publication Date: 2025-06-10NAT UNIV OF DEFENSE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510626666.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-10
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the learning of multi-drone collaborative strategy, under the sparse reward conditions, traditional novel value metrics cannot effectively distinguish the equivalent state, resulting in the drone repeatedly exploring the equivalent state, affecting the efficiency of strategy learning.

Method used

The permutation invariant spatiotemporal feature representation is used to uniformly characterize the equivalent state, and the novel value of equivalent state is measured by the random network distillation method, and the intrinsic reward is determined to drive the drone to explore the non-equivalent state space.

Benefits of technology

It effectively reduces repeated access to equivalent states, improves the search efficiency of sparse positive rewards, and improves the efficiency of multi-drone collaboration strategy optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122701A_ABST
    Figure CN120122701A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-unmanned aerial vehicle cooperation strategy learning method and device driven by an equivalent state novel value. According to the method, replacement invariant spatial-temporal feature representation is defined, and uniform representation of an equivalent state is realized through spatial-temporal information in a similar image feature aggregation joint state; a random network distillation method is adopted to quickly and effectively evaluate the accessed degree of the high-dimensional equivalent state set on the basis of the permutation invariant spatial-temporal characteristics, consistent measurement of equivalent state novel values is achieved, and the overall accessed degree of the equivalent state set is estimated end to end; driving the unmanned aerial vehicle to explore by using the equivalent state novel value; and finally, the equivalent state novel value is used as a shared internal reward, the shared internal reward and an original external reward guide strategy updating together, and the method improves the cooperation strategy optimization efficiency under the sparse reward condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multi-UAV cooperative strategy learning, and particularly to a multi-UAV cooperative strategy learning method and device driven by equivalent state novelty values. Background Art

[0002] Under sparse reward conditions, it is difficult for reinforcement learning to capture positive rewards through random actions, resulting in a lack of feedback signals for the policy and making it difficult to optimize. Compared with single UAVs, the size of the joint state space in the multi-UAV scenario increases exponentially with the increase in the number of agents, and specific joint states require the cooperation of each UAV to reach, resulting in more severe challenges brought by sparse rewards. Therefore, there is an urgent need for an effective cooperative exploration mechanism to drive multi-UAVs to search for positive rewards in the joint state space, so as to solve the problem of multi-UAV policy learning under sparse rewards. Constructing an intrinsic reward based on state novelty is a classic exploration mechanism in reinforcement learning. Novelty is a measure of the frequency of state access. This method effectively drives the UAV to explore the state space in a traversal manner by assigning higher intrinsic reward values to states with low access frequencies. In the multi-UAV scenario, there are a large number of equivalent states in the joint state space due to the isomorphic characteristics of UAVs. Traditional novelty metrics only estimate the access degree of the state itself, resulting in inconsistent novelty metrics for equivalent states. UAVs repeatedly explore equivalent states during the search for positive (extrinsic) rewards, affecting the efficiency of policy learning. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a multi-UAV cooperative strategy learning method and device driven by equivalent state novelty values.

[0004] A multi-UAV cooperative strategy learning method driven by equivalent state novelty values, the method includes: In a multi-UAV task scenario, determine the joint state space according to a multi-UAV collaborative decision-making model based on a distributed partially observable Markov decision process.

[0005] Construct an equivalent state set for each state according to the joint state space.

[0006] Each UAV selects an action according to its own policy and the current observation. After all UAVs execute the joint action and undergo state transition, each UAV obtains a first experience tuple including observation, action, and extrinsic reward. Use permutation-invariant spatio-temporal feature representation to uniformly represent the equivalent states of the state after state transition.

[0007] The novelty value of the equivalent state is measured by using the random network distillation method according to the permutation-invariant spatio-temporal characteristics of the equivalent state to determine the intrinsic reward; and after incorporating the intrinsic reward into the first experience tuple, it is added to the experience storage; the extrinsic rewards and intrinsic rewards of all drones are the same for each state transition. The extrinsic rewards and intrinsic rewards of the drones i are weighted and summed to obtain the overall reward of the drones i ; the state transition and overall reward calculation are repeatedly executed until the decision round ends.

[0008] According to the overall reward and the experience storage, the ACER method is used to optimize the value network and policy network in the Actor-Critic architecture.

[0009] A multi-UAV collaborative strategy learning device driven by the novelty value of equivalent states, the device includes: An equivalent state set construction module, which is used to determine the joint state space according to the multi-UAV collaborative decision-making model based on the distributed partially observable Markov decision process in the multi-UAV task scenario; and construct the equivalent state set of each state according to the joint state space.

[0010] A state transition module, which is used for each UAV to select an action according to its own policy and the current observation. After all UAVs execute the joint action and go through state transition, each UAV obtains a first experience tuple containing the observation, action, and extrinsic reward; the extrinsic rewards of all UAVs are the same for each state transition.

[0011] An intrinsic reward determination module, which is used to uniformly represent the equivalent states of the states after state transition by using the permutation-invariant spatio-temporal feature representation; measure the novelty value of the equivalent state by using the random network distillation method according to the permutation-invariant spatio-temporal characteristics of the equivalent state to determine the intrinsic reward; and after incorporating the intrinsic reward into the first experience tuple, add it to the experience storage; all UAVs share the intrinsic reward.

[0012] An overall reward determination module, which is used to weight and sum the extrinsic rewards and intrinsic rewards of the drones i to obtain the overall reward of the drones i ; the state transition and overall reward calculation are repeatedly executed until the decision round ends.

[0013] A UAV policy optimization module, which is used to optimize the value network and policy network in the Actor-Critic architecture according to the overall reward and the experience storage by using the ACER method.

[0014] The above-mentioned multi-UAV cooperative strategy learning method and device driven by the novelty value of equivalent states define a permutation-invariant spatio-temporal feature representation, aggregate spatio-temporal information in the joint state through class image feature aggregation, and achieve a unified representation of equivalent states; design a random network distillation method based on permutation-invariant spatio-temporal features to quickly and effectively evaluate the access degree of the high-dimensional equivalent state set, achieve a consistent measurement of the novelty value of equivalent states, and estimate the overall access degree of the equivalent state set end-to-end; drive the UAV exploration with the novelty value of equivalent states; finally, use the novelty value of equivalent states as a shared intrinsic reward to jointly guide the policy update with the original extrinsic reward. This method improves the efficiency of cooperative strategy optimization under sparse reward conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic flowchart of a multi-UAV cooperative strategy learning method driven by the novelty value of equivalent states in an embodiment; Figure 2 It is a schematic diagram of a multi-UAV cooperative task scenario in another embodiment, where Figure 2 (a) is a schematic diagram of scenario 2UAV-1ADS-1, Figure 2 (b) is a schematic diagram of scenario 2UAV-1ADS-2, Figure 2 (c) is a schematic diagram of scenario 3UAV-1ADS, Figure 2 (d) is a schematic diagram of scenario 3UAV-2ADS; Figure 3 It is a bar chart of the number of states covered by each algorithm in 2UAV-1ADS-1 and in scenario 3UAV-1ADS in another embodiment, where Figure 3 (a) is a bar chart of the number of states covered by each algorithm in scenario 2UAV-1ADS-1, Figure 3 (b) is a bar chart of the number of states covered by each algorithm in scenario 3UAV-1ADS; Figure 4 It is a curve chart of the cumulative extrinsic reward and task success rate in scenario 2UAV-1ADS-1 in another embodiment, where Figure 4 (a) is a curve chart of the cumulative extrinsic reward, Figure 4 (b) is a curve chart of the task success rate; Figure 5 It is a curve chart of the cumulative extrinsic reward and task success rate in scenario 2UAV-1ADS-2 in another embodiment, where Figure 5 (a) is a curve chart of the cumulative extrinsic reward, Figure 5 (b) is a curve chart of the task success rate; Figure 6 It is a bar chart of the number of training rounds required for each algorithm strategy to reach a specified success rate in different scenarios in another embodiment, where Figure 6(a) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 2UAV-1ADS-1, Figure 6 (b) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 2UAV-1ADS-2, Figure 6 (c) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 3UAV-1ADS, Figure 6 (d) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario 3UAV-2ADS. Specific implementation manner

[0016] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0017] In one embodiment, as Figure 1 shown, a multi-UAV cooperative strategy learning method driven by an equivalent state novelty value is provided, and the method includes the following steps: Step 100: In a multi-UAV mission scenario, determine the joint state space according to a multi-UAV collaborative decision-making model based on a distributed partially observable Markov decision process.

[0018] Specifically, in a multi-UAV mission scenario, according to a multi-UAV collaborative decision-making model based on a distributed partially observable Markov decision process (Dec-POMDP), the joint state space includes the individual states of all UAVs and the states of other environmental elements : .

[0019] Step 102: Construct an equivalent state set for each state according to the joint state space.

[0020] Specifically, assume that there are two joint state variables in the reinforcement learning model corresponding to a multi-UAV cooperative penetration and assault mission , where , .

[0021] When the UAVs in the mission scenario are homogeneous, the situation of "UAV 1 is in the individual state , UAV 2 is in the individual state , UAV 3 is in the individual state " is the same as the situation of "UAV 2 is in the individual state , UAV 3 is in the individual state , UAV 1 is in the individual state " can be attributed to "a drone in an individual state , a drone is in individual state and a drone in individual state ”.

[0022] Since the mission objectives of the drone group are consistent, whether the mission is completed or not depends on the overall status, and has nothing to do with the sequence number of the drones in each individual state. and They can be considered equivalent at the task level. The equivalent states and equivalent state sets in the multi-UAV task scenario are defined as follows: For status and , if there is a permutation , so that ,and ,but and are equivalent to each other, denoted by .

[0023] For any joint state , define its equivalent state set for The set of all equivalent states, .

[0024] Step 104: Each drone selects an action based on its own strategy and current observation. After all drones perform the joint action, after state transfer, each drone obtains the first experience tuple containing observation, action and external reward.

[0025] Specifically, each drone has its own strategy. and current observations Select Action ; All drones perform joint actions After the state transfer, each drone obtains the first experience tuple containing observation, action, and external reward. ,in Indicates whether the round is terminated.

[0026] Step 106: Use the permutation-invariant spatiotemporal feature representation to uniformly characterize the equivalent states of the states after the state transfer.

[0027] Specifically, the features corresponding to different equivalent states are different, resulting in inconsistent novelty values ​​estimated after inputting the prediction network and the target network in the random network distillation method. Therefore, the spatiotemporal information in the joint state is aggregated by image-like features, and the permutation-invariant spatiotemporal feature representation is used to achieve a unified representation of equivalent states.

[0028] Permutation-Invariant (PI) is used to describe the impact of the permutation of each input in a multi-input function on the output.

[0029] Step 108: Use the random network distillation method to measure the novelty value of the equivalent state according to the permutation-invariant spatio-temporal characteristics of the equivalent state, and determine the intrinsic reward; and add the intrinsic reward to the first experience tuple and then add it to the experience storage; the extrinsic and intrinsic rewards of all drones are the same for each state transition.

[0030] Specifically, based on the intrinsic reward method of state novelty value for the state of the (approximate) number of visits is estimated to the negative correlation function of (e.g., or ) is used as the state novelty value metric, and an intrinsic reward is constructed to drive the agent to continuously explore states with low visit frequencies, and perform a traversal search of the state space to mine the sparsely distributed positive rewards. However, in this way, for each state in the equivalent state set due to the different number of visits, the novelty value metrics are inconsistent. As long as there are still individual states with relatively high novelty values, the drones will be encouraged to visit , resulting in repeated exploration of the equivalent states. Under sparse reward conditions, due to the lack of reward feedback signals during the task process, the drones need to start from the initial state set to explore the state space and search for the state set containing positive rewards to provide effective feedback for policy optimization. In this process, the repeated exploration of the equivalent states hinders the search for the reward states , thereby affecting the multi-drone cooperation policy optimization efficiency under sparse reward conditions. Therefore, the intrinsic reward method based on state novelty value is improved to avoid the repeated exploration caused by the inconsistent measurement of the equivalent state novelty value, and the traversal exploration of the drones in the original joint state space is converted into traversal exploration in the non-equivalent state space .

[0031] The original joint state space corresponding non-equivalent state space is defined as follows: , where is any mapping that can convert the equivalent state into the same state, that is, it satisfies: , .

[0032] The non-equivalent state space has the following property: for any , there is Established , and .

[0033] Compared with the original joint state space , traversing the non-equivalent state space to search for rewards can save search costs. To encourage the drone to explore the non-equivalent state space and avoid repeated exploration of equivalent states, the number of times the equivalent state set of each state is visited is measured and replaces the number of times each state is visited individually to construct an intrinsic reward based on the novelty value of equivalent states. , For the state the corresponding element in the non-equivalent state space the number of visits was counted, and the negative correlation function of

[0034] is used as the intrinsic reward of the novelty value to encourage the drone to explore the equivalent state set with low visit frequency rather than the state with low visit frequency, thereby driving the drone to traverse and explore the non-equivalent state space.

[0035] The experience samples in the experience storage are , where represents the current observation, represents the selected action, represents the next moment observation, represents the overall reward, is the policy, represents whether this round is a termination flag.

[0036] Step 110: Perform weighted summation of the extrinsic reward and the intrinsic reward of the drone i to obtain the overall reward of the drone i ; repeatedly execute state transition and overall reward calculation until the decision-making round ends.

[0037] Step 112: Optimize the value network and the policy network in the Actor-Critic architecture using the ACER method according to the overall reward and the experience storage.

[0038] Specifically, based on the Actor-Critic structure, each UAV maintains its own policy network and value network , where and are network parameters. For the sake of simplicity in the following text, part of them are represented by and .

[0039] Since the multi-UAV collaborative penetration and assault is a fully cooperative task, the external rewards (i.e., the original reward signals feedback by the environment) of all UAVs are the same for each state transition, that is , , . In addition, to drive the UAVs to collaboratively explore the state space, all UAVs share the intrinsic reward based on the equivalent state novelty value, that is .

[0040] The UAV uses the weighted sum of the external reward and the intrinsic reward as the overall reward to update the value function and the policy network.

[0041] In the above multi-UAV collaborative policy learning method driven by the equivalent state novelty value, the method defines a permutation-invariant spatio-temporal feature representation, aggregates the spatio-temporal information in the joint state through class image feature aggregation, and realizes the unified representation of the equivalent state; designs a random network distillation method based on the permutation-invariant spatio-temporal feature to quickly and effectively evaluate the accessed degree of the high-dimensional equivalent state set, realizes the consistent measurement of the equivalent state novelty value, and estimates the overall accessed degree of the equivalent state set end-to-end; drives the UAVs to explore with the equivalent state novelty value; finally, uses the equivalent state novelty value as the shared intrinsic reward to jointly guide the policy update with the original external reward. This method improves the efficiency of collaborative policy optimization under sparse reward conditions.

[0042] In one embodiment, the joint state space in step 100 includes the individual states of all UAVs and the states of other environmental factors; the individual state of the UAV is:[[]] ; where are the x coordinate, y coordinate, yaw angle variable, and speed variable of the UAV i respectively; is the duration that the UAV is threatened k locked or intercepted, is an integer.

[0043] The environmental state includes the radar antenna angle, channel occupancy, and target state.

[0044] In one embodiment, step 102 includes: extracting a key observable part from the state variables in the joint state space to form sub-state variables; determining the sub-states of each state according to the sub-state variables; constructing equivalent sub-states and equivalent sub-state sets for each sub-state; when constructing the intrinsic reward based on the novel value of the equivalent state, using the equivalent sub-states and equivalent sub-state sets to replace the equivalent states and equivalent state sets of the corresponding states in the original joint state space.

[0045] In one embodiment, step 106 includes: using a permutation-invariant spatio-temporal feature representation to uniformly represent the equivalent states of the states after state transition, obtaining the permutation-invariant spatio-temporal features of the equivalent states; the expression of the permutation-invariant spatio-temporal features is: ;

[0046] where, is the permutation-invariant spatio-temporal feature of the equivalent state , n is the total number of UAVs, , , is the number of UAVs of the threatening UAV i ; , , , , represents the range of the rectangular mission area; is used to indicate the position of the UAV. If the UAV is located in the grid , then , otherwise ; compared with characterizes the spatial information in the state, then embeds the time information in the state into , is the UAV being threatened locking or intercepting time of the linear mapping, , and are respectively the mapping values of the upper and lower bounds of the locking or intercepting duration of the UAV being threatened k .

[0047] Specifically, when measuring the state novelty value based on the random network distillation method, the feature forms of the state variables input to the prediction network and the target network are one of the keys to the novelty value measurement. The state variables As a high-dimensional real-valued variable, it can be directly input into the prediction network and the target network after normalization processing, or the values of each dimension can be transformed through a method similar to one-hot encoding and then concatenated as input features. The prediction network and the target network use a fully-connected layer as the input layer to process the features. However, there is a significant problem with the above feature construction, that is, different states with the same equivalence correspond to different features, resulting in inconsistent estimated novelty values after being input into the prediction network and the target network.

[0048] Therefore, a method for constructing permutation-invariant spatio-temporal features (PI-STF) of states is proposed. Permutation invariance (PI) is used to describe the influence of the permutation of each input in a multi-input function on the output. For the function , , if the permutation of the input elements is changed without changing the output, that is, for any permutation matrix , there is holds, then the function satisfies permutation invariance.

[0049] The permutation-invariant spatio-temporal feature maps the state variable into an image-like feature with a size of , and the number of channels is , denoted as represents the range of the rectangular task area, and is transformed into through discrete intervals grid, is the number of threats, represents rounding up.

[0050] The permutation-invariant spatio-temporal feature is specifically defined as shown in the expression of the above permutation-invariant spatio-temporal feature.

[0051] Compared with directly using , mapping to and then constructing the permutation-invariant spatio-temporal feature has two advantages. On the one hand, numerical normalization is performed on this variable, and on the other hand, ambiguity in the feature construction process is avoided. If is directly used, when , it is impossible to distinguish between "the drone is located in the grid , but is not tracked and intercepted by the threat system , that is but " and "the drone Not on grid ,Right now "Two completely different situations; and using This ensures that once the drone Located in the grid In, regardless of whether it is currently being tracked or intercepted, Both are greater than 0, which effectively distinguishes the above two situations.

[0052] In one embodiment, step 108 includes: constructing a random distillation network; the random distillation network includes a target network and a prediction network, the target network and the prediction network have the same structure, and the initialized network parameters are different; the parameters of the target network remain fixed after initialization, and the parameters of the prediction network are updated according to the mean square error between its own output and the target network output using the gradient direction; the permutation invariant spatiotemporal features of the equivalent state are input into the target network and the prediction network, and according to the outputs of the target network and the prediction network, the equivalent state novelty value metric is determined as: ; The intrinsic reward based on equivalent state novelty value is constructed as: ; in, is the intrinsic reward based on the novelty value of the equivalent state, They are the current state, the current action, and the next state. It is the equivalent state novelty value measure of the state at the next moment.

[0053] Specifically, in the multi-UAV collaborative penetration assault mission scenario, the joint state is , including drones The individual status is , environmental status In order to improve the efficiency of sparse reward search and the feasibility of the method, the state variables Extract the key observable parts and form sub-state variables , reducing the size of the search space. Under threat countermeasures, multiple drones need to complete collaborative penetration and assault missions through close spatiotemporal coordination. Therefore, it is necessary to encourage the exploration of diverse drone spatiotemporal coordination modes through intrinsic rewards, so as to explore effective collaborative strategies. Relative position information and the duration of being locked or intercepted by the countermeasure system, drone attitude information and speed information are relatively secondary when exploring diverse spatiotemporal coordination modes.

[0054] To do this, define the substate variable Ignore the yaw angle variable and speed variables , , , .

[0055] For sub - states and , if there exists a permutation such that , then and are equivalent sub - states, denoted as .

[0056] For any sub - state , define its equivalent state set as the set of all equivalent sub - states. , analogous to equivalent states and equivalent state sets, define equivalent sub - states and equivalent sub - state sets based on .

[0057] Use the sub - state to replace to count the number of times the equivalent state set is visited , which is used to construct an intrinsic reward based on the novelty value of equivalent states. , when the state is a low - dimensional variable, the number of UAVs is small, and the state space is small, it can be calculated by counting, and can be simply accumulated based on the equivalent state set on this basis. As the dimension of the state variable increases, the number of UAVs increases, and the size of the state space increases sharply. At this time, counting the number of times each state is visited based on the counting method requires a large amount of memory, and the size of the equivalent state set increases with the number of permutations and combinations of UAV serial numbers. Calculating and accumulating the number of times each element in it is visited separately is too cumbersome.

[0058] Based on the above analysis, aiming at a high - dimensional continuous - discrete hybrid state space, an end - to - end consistent measure method for the novelty value of equivalent states is proposed to construct an intrinsic reward to encourage UAVs to efficiently explore the non - equivalent state space and improve the learning efficiency of cooperative strategies under sparse reward conditions. Design the intrinsic reward by using the sub - state variable to replace the state variable . Without affecting the method description and analysis, for simplicity, the state variables related to novelty value measurement in the following text are default to refer to . Random Network Distillation (RND) is a method for measuring the novelty value of high-dimensional states or observations. The random network distillation method estimates the novelty value of a state or observation based on the following fact: in the supervised learning paradigm, if a test sample has more similar trained samples, the corresponding prediction error is smaller.

[0059] The random network distillation method consists of two neural networks: the Target Network and the Predictor Network. The parameters of the target network are fixed after random initialization, and the predictor network is trained with the experienced states or observations as samples. Taking the measurement of the novelty value of the observed variable as an example, the target network and the predictor network map the observation into features of the same dimension. The predictor network takes the output of the target network as the target value, and based on the observations experienced by the agent, constructs a training sample set containing observation-label , and updates its own network parameters by minimizing the mean squared error . Under this training mechanism, for a given observed variable , if it is accessed more frequently, the predictor network parameters are updated more times based on its similar observation samples, and the corresponding prediction error is smaller; conversely, if the observation is accessed less frequently, the predictor network parameters are updated fewer times based on its similar observation samples, and the corresponding prediction error is larger. Therefore, the prediction error can be used to measure the novelty value of the observation . If the inputs of the target network and the predictor network are changed from observations to states, the novelty value of the state variable can be measured. Compared with other methods for measuring the novelty value of high-dimensional states or observations, such as density estimation-based methods, the random network distillation method has significant advantages of wide applicability, simple structure, and small computational cost, and only needs to additionally train the predictor network.

[0060] The Permutation-Invariant Spatio-Temporal Feature-based Random Networks Distillation (PI-STF-RND) method aims to achieve end-to-end high-dimensional equivalent state novelty value measurement and further construct an intrinsic reward to guide the efficient exploration of drones. As a kind of image-like feature, when permutation-invariant spatio-temporal features are used for state novelty value measurement, the target network in the permutation-invariant spatio-temporal feature-based random network distillation method and the prediction network Taking the convolution layer as the input layer, multiple convolutional kernels are used to extract and fuse features of the permutation-invariant spatio-temporal features from different perspectives. Then, the output of the convolution layer is processed by the flatten operation and input into the subsequent fully connected layer. , where is the dimension of the final output of the network in the random network distillation method.

[0061] Given the state , the target network and the prediction network respectively output and . In the constructed random network distillation, both the target network and the prediction network include a convolution layer, a rectified linear unit (ReLU), a flatten layer, and a fully connected layer. Among them, the convolutional kernel size of the convolution layer is 5×5, the number of convolutional kernels is 32, the stride is 2, and the padding size is 2. The size of the fully connected layer FC1 is 512, and the size of the fully connected layer FC2 is 10, that is, N = 10. The target network and the prediction network have the same structure, but the initialized network parameters are different. The parameters of the target network remain fixed after self-initialization, while the parameters of the prediction network need to be updated according to the mean square error between its own output and the output of the target network. The mean square error between its own output and the output of the target network is:[[]] ; The prediction network parameter is updated by the gradient method, aiming to achieve ; The update process is:[[]] ; ; where is the set of states experienced in this decision-making episode (Episode).

[0062] After each episode ends, the set generated during the multi-UAV decision-making process is used as a training sample containing features - labels to update the parameters of the prediction network. In the supervised learning paradigm, the neural network has a better fitting effect on the training samples with higher frequencies of occurrence.

[0063] Therefore, when the new state variable is converted into the permutation-invariant spatio-temporal feature and then input into the PI-STF-RND network, the prediction network outputs and the output of the target network The difference between them can reflect the permutation-invariant features or the number of occurrences of its approximate features in the training samples , the smaller the difference, then the larger.

[0064] According to having permutation invariance, , for all it holds, and we can obtain . The closer is to , the more times the equivalent state of is accessed during the decision-making process, and the lower the overall novelty value of its equivalent state set . Based on the above analysis, the following method for measuring the novelty value of the equivalent state is proposed: ;

[0065] The PI-STF-RND method uses a feature representation with permutation invariance to map equivalent states to the same spatio-temporal features, and then estimates the occurrence frequency of each spatio-temporal feature in the training samples through the random network distillation method, thus directly reflecting the access degree of the corresponding equivalent state set, saving the cumbersome operation of separately estimating and accumulating the access times of each equivalent state, and realizing the fast and effective evaluation of the novelty value of high-dimensional equivalent states.

[0066] To drive the drone to actively explore the state space, the novelty value measured by the PI-STF-RND method is used to construct the intrinsic reward of the equivalent state novelty value : ;

[0067] After each state transition, the novelty value measurement of the next moment state based on PI-STF-RND is used as the intrinsic reward, and all drones share this intrinsic reward. Since actually measures not only the access degree of state , but the overall access degree of the equivalent state set , it avoids the repeated access to equivalent states caused by inconsistent novelty value measurement, effectively encourages the drone to explore non-equivalent states with low access frequency, and improves the search efficiency for sparse positive rewards.

[0068] In one embodiment, both the target network and the prediction network include: a convolutional layer, a rectified linear unit, a flattening layer, and two fully connected layers.

[0069] In one embodiment, step 110 includes: the drone i ​​​​​The extrinsic reward and the intrinsic reward are weighted and summed to obtain the overall reward of the drone i The overall reward of the drone is: ;

[0070] Wherein, is the overall reward of the drone i and and are the weights of the extrinsic reward and the intrinsic reward based on the equivalent situation novelty value respectively, is the extrinsic reward, is the intrinsic reward based on the equivalent situation novelty value.

[0071] In one embodiment, the second experience tuple in the experience store is: wherein, is the overall reward, is the policy, is the current observation, is the action, is the observation at the next moment, is the extrinsic reward, is the flag indicating whether the round is terminated; Step 112 includes: each drone randomly samples from the experience store, and uses the sampled experience data and the overall reward to update the value network and the policy network by using the ACER method; The ACER method includes: a policy evaluation link and a policy improvement link; In the policy evaluation link: the value network parameters are updated by minimizing the mean square error; The update expression of the value network parameters is: ; ; ;

[0072] Wherein, is the mean square error of the value network, are the parameters of the value network, is the learning rate of the drone value network, is the value estimate of the current policy based on the Retrace method, is the overall reward, is the value network, are respectively the current observation and action of the drone i , is the drone i parameters of the policy network, is a hyperparameter, is the truncated importance sampling coefficient, , , is the drone Generate experience The historical strategy at this time is denoted as the behavioral strategy.

[0073] In the policy improvement step, the parameters of the UAV policy network are updated using the ACER policy gradient; the update expression for the policy network parameters is as follows: ; ; Among them, are the parameters of the policy network, is the learning rate of the UAV policy network, is the ACER policy gradient, is the current policy of the UAV, is the value of the current observation, is the cost.

[0074] In one embodiment, the overall framework of multi-UAV policy learning based on the equivalent state novelty value intrinsic reward includes two parts: UAV-environment interaction and UAV policy training: (1) UAV-environment interaction process: Step 1: Each UAV selects an action according to its own policy and the current observation ; Step 2: After all UAVs execute the joint action , through state transition, each UAV obtains an experience tuple containing observation, action, and extrinsic reward indicating whether the round terminates; Step 3: Then, the novelty value is measured using the random network distillation method based on permutation-invariant spatio-temporal features to calculate the corresponding intrinsic reward , and the experience tuple combining the intrinsic reward and the extrinsic reward is added to the experience storage. Each UAV performs the above operations after each state transition.

[0075] (2) UAV policy training process: Step 1: After the decision-making round ends, the prediction network of the random network distillation method based on permutation-invariant spatio-temporal features is updated using all the state variables experienced in this round; Step 2: Each UAV randomly samples from the experience storage and updates its own policy and the value function . The UAV policy is optimized based on the ACER method, including the policy evaluation and policy improvement steps, and the UAV value network and the policy network Perform an update.

[0076] In the policy evaluation phase, the value network parameters are updated by minimizing the mean squared error Perform an update.

[0077] In the policy improvement phase, use the ACER policy gradient to update the parameters of the UAV policy network .

[0078] If , otherwise . The ACER method is a typical off-policy training method that trains the current policy by sampling past experience samples from the experience buffer.

[0079] In the policy evaluation phase and the policy improvement phase, the value network parameters and the policy network parameters are both updated according to the sampled historical experience samples. Given the historical decision trajectory of the UAV sampled from the experience buffer , , the gradient calculation of the corresponding value network parameters is as follows: ; The value network parameters are updated in the same way as the update expression of the above value network parameters.

[0080] The corresponding gradient of the policy network parameters is , and the update method of is as shown in the update expression of the above policy network parameters.

[0081] Each UAV interacts with the environment by executing the current policy, and then continuously optimizes the policy according to the above policy update rules using the interaction experience, repeating this process until the maximum number of training episodes is reached.

[0082] It should be understood that although Figure 1 the steps in the flowchart of are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least some of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed and completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0083] In a verification embodiment, with the background that a UAV task unit arrives at the periphery of the defense area and identifies the protected target under multiple threats, a multi-UAV cooperative penetration and assault mission scenario under threat countermeasures is set up. Under the given threats and UAV system performance parameters, the deployment plan of the mission scenario satisfies: (1) The protected target is within the effective defense range that a single UAV cannot enter, and the mission needs to be completed by the cooperation of the UAV unit; (2) The UAV unit needs to continuously carry out close cooperation in space and time to complete the mission, which requires a high level of cooperation, resulting in greater difficulty in policy training under sparse reward conditions.

[0084] Four scenarios are set for algorithm verification based on the different relative positions of the target and the threat system: 2UAV-1ADS-1, 2UAV-1ADS-2, 3UAV-1ADS, and 3UAV-2ADS. The scenario names include the number of UAVs and the number of threat systems in the mission area. As Figure 2 shown, among them Figure 2 (a) is the scenario 2UAV-1ADS-1, Figure 2 (b) is the scenario 2UAV-1ADS-2, Figure 2 (c) is the scenario 3UAV-1ADS, Figure 2 (d) is the scenario 3UAV-2ADS. Both the scenarios 2UAV-1ADS-1 and 2UAV-1ADS-2 include 2 UAVs and 1 threat in a 15 km × 15 km mission area, and the distance from the target PA to the threat system ADS1 is the same. However, in the scenario 2UAV-1ADS-2, the target is located in the northwest direction of the threat system. Due to the mission area limitation, the point closest to the target outside the red line range has a large deviation from the UAV incoming direction, that is, the range of penetration points that meet the conditions is narrow. Therefore, the mission difficulty is greater than that of the scenario 2UAV-1ADS-1; the scenario 3UAV-1ADS further reduces the distance between the target and the threat system, and the mission difficulty is further increased. Accordingly, the scale of the UAV unit is expanded to 3; the scenario 3UAV-2ADS tests the learning situation of the multi-UAV cooperation strategy when the target is located in the overlapping defense range of multiple threat systems.

[0085] The multi - UAV collaborative policy learning method proposed in this application based on the intrinsic reward of equivalent state novelty value is denoted as PI - STF - RND. In addition, two benchmark methods are selected for comparison, denoted as No - Intrinsic - Reward and ECV - RND methods respectively.

[0086] No - Intrinsic - Reward: No additional intrinsic reward is calculated, and the UAV policy is updated only with extrinsic rewards based on the ACER algorithm. Under this method, the UAV randomly explores the state space according to the current policy to search for positive rewards. The overall reward is set as: 。

[0087] ECV - RND: Each part of the state variable is encoded by ECV (Efficient Coordinated Vector, ECV) and then directly concatenated as the state feature input into the RND network. The situation novelty value estimated by ECV - RND is used as the intrinsic reward ECV - RND, and the weighted sum of the ECV - RND intrinsic reward and the extrinsic reward is used to update the UAV policy based on the ACER algorithm. ECV encoding is an extension of the classical one - hot encoding method, introducing weights to represent the distance between continuous real numbers and the surrounding discrete real numbers, thus realizing the representation of continuous variables 1. Different from PI - STF - RND, since the state feature is a one - dimensional array at this time, the input layers of the target network and the prediction network in ECV - RND are set as fully - connected layers. Since the state feature based on ECV encoding and concatenation does not have permutation invariance, the features corresponding to different equivalent states are different. Therefore, ECV - RND does not consider the overall access degree of the equivalent state set, but only considers the given state itself. The overall reward is set as: 。

[0088] Comparing PI - STF - RND with No - Intrinsic - Reward can analyze the influence of the intrinsic reward based on the equivalent state novelty value on the UAV state - space exploration efficiency and the UAV collaborative policy learning efficiency; comparing PI - STF - RND and ECV - RND can analyze whether the intrinsic reward based on the equivalent state novelty value effectively reduces the repeated access to equivalent states compared with the general intrinsic reward of state novelty value, thus improving the collaborative policy learning efficiency.

[0089] Table 1 Reward - setting parameters of UAV policy - learning algorithms

[0090] The UAV policy network consists of, from the input layer to the output layer, a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, a fully connected layer (with 7 layers), and a softmax function. The UAV action value network consists of, from the input layer to the output layer, a fully connected layer (with 1024 layers), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 7 layers). The PI-STF-RND target network and the prediction network consist of a convolutional layer with 32 5×5 convolutional kernels, a rectified linear unit, a flattening layer, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 10 layers). The ECV-RND target network and the prediction network consist of a fully connected layer (with m layers, where m is a positive integer), a rectified linear unit, a fully connected layer (with 512 layers), a rectified linear unit, and a fully connected layer (with 10 layers). In the 2UAV-1ADS-1, 2UAV-1ADS-2, and 3UAV-1ADS scenarios, m is 2048, and in the 3UAV-2ADS scenario, m is 3872.

[0091] In the multi-UAV cooperative exploration experiment, the exploration effects of the exploration mechanism driven by the intrinsic reward of the PI-STF-RND equivalent state novelty value, the exploration mechanism driven by the intrinsic reward of the ECV-RND state novelty value, and the No-Intrinsic-Reward random exploration mechanism are compared by excluding the influence of external rewards.

[0092] Table 2 Comparison of the exploration effects of each algorithm in the 2UAV-1ADS-1 and 3UAV-1ADS scenarios

[0093] The bar charts of the number of covered states of each algorithm in the 2UAV-1ADS-1 and 3UAV-1ADS scenarios are as Figure 3 shown, where Figure 3 (a) is the bar chart of the number of covered states of each algorithm in the 2UAV-1ADS-1 scenario, Figure 3 (b) is the bar chart of the number of covered states of each algorithm in the 3UAV-1ADS scenario.

[0094] Table 2 and Figure 3The data in [description] was obtained after each test algorithm was trained for 2.0×10^4 rounds in the scenario of 2UAV-1ADS-1 and 5.0×10^4 rounds in the scenario of 3UAV-1ADS. From the experimental results, it can be seen that the number of original state covers and non-equivalent state covers of the PI-STF-RND algorithm are significantly higher than those of the ECV-RND and No-Intrinsic-Reward algorithms in both scenarios. The number of original state covers and non-equivalent state covers of the No-Intrinsic-Reward algorithm are the least, indicating that the corresponding random exploration mechanism has the lowest efficiency. In the scenario of 2UAV-1ADS-1, the number of original state covers of PI-STF-RND is about 17.6% higher than that of the ECV-RND algorithm and about 148.1% higher than that of the No-Intrinsic-Reward algorithm. The number of non-equivalent state covers is about 30.1% higher than that of the ECV-RND algorithm and about 165.3% higher than that of the No-Intrinsic-Reward algorithm. In the scenario of 3UAV-1ADS, the number of original state covers of PI-STF-RND is about 19.8% higher than that of the ECV-RND algorithm and about 47.4% higher than that of the No-Intrinsic-Reward algorithm. The number of non-equivalent state covers is about 37.5% higher than that of the ECV-RND algorithm and about 70.6% higher than that of the No-Intrinsic-Reward algorithm. The experimental results of multi-UAV collaborative exploration show that the PI-STF-RND algorithm effectively reduces the repeated access of UAVs to the equivalent state set through the intrinsic reward of equivalent state novelty, and improves the exploration efficiency of the original state space and non-equivalent state space (RQ1).

[0095] Figure 4 shows the performance of each algorithm in the scenario of 2UAV-1ADS-1, where PI-STF-RND is significantly better than other algorithms in multiple evaluation indicators such as cumulative extrinsic reward, task success rate, and the number of training rounds required to reach 60%, 80%, and 90% success rates. As Figure 4 (a) shows, the cumulative extrinsic reward curve of PI-STF-RND starts to rise at about 2.0×10^3 training rounds and finally converges basically at 7.0×10^3 training rounds, and is higher than the cumulative extrinsic reward curves of other algorithms throughout the training process. As Figure 4As shown in (b), after training for about 6.0×103 rounds, the task success rate of the PI-STF-RND algorithm approaches 100%. At this time, the task success rate of the ECV-RND algorithm is only about 60%, and the task success rate of the No-Intrinsic-Reward algorithm is only about 25%. According to the experimental results, since the tasks in scenario 2UAV-1ADS-1 are relatively simple, the No-Intrinsic-Reward algorithm can also achieve an accumulated extrinsic reward of about 0.5 after training for 1.0×104 rounds through random exploration, and the task success rate exceeds 70%. As Figure 4 As shown in (a), in scenario 2UAV-1ADS-1, when the task success rate reaches 60%, 80%, and 90%, the number of training rounds required by the PI-STF-RND algorithm is reduced by 23.7%, 34.2%, and 34.5% compared with the ECV-RND algorithm, and is reduced by more than 43.75%, 50.0%, and 45.0% compared with the No-Intrinsic-Reward algorithm.

[0096] As Figure 5 shown, compared with scenario 2UAV-1ADS-1, scenario 2UAV-1ADS-2 is more challenging. Therefore, the No-Intrinsic-Reward algorithm can achieve a task success rate of more than 70% after 1.0×104 training rounds in scenario 2UAV-1ADS-1 through random exploration. However, in scenario 2UAV-1ADS-2, even after 2.0×104 rounds of training, the algorithm has not learned the cooperative strategy for successful penetration and assault. This result is consistent with the analysis of scenarios 2UAV-1ADS-1 and 2UAV-1ADS-2 during the test scenario setting. The change in the relative position of the target PA and the threat system ADS1 in the scenario has a greater impact on the task difficulty. Despite the increase in task difficulty, the PI-STF-RND algorithm is significantly better than other comparison algorithms in multiple evaluation indicators such as cumulative extrinsic reward, task success rate, and the number of training rounds required to reach 60%, 80%, and 90% success rates in scenario 2UAV-1ADS-2. As Figure 5 As shown in (a), during the entire training process, the cumulative extrinsic reward curve of the PI-STF-RND algorithm is always above that of the ECV-RND algorithm. As Figure 5 shown in (b), the ECV-RND algorithm reaches a task success rate of about 83% after 2.0×104 rounds of training. The PI-STF-RND algorithm approaches an 80% task success rate after 1.0×104 training rounds and reaches a task success rate of about 95% after 2.0×104 training rounds. The bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in different scenarios is as Figure 6 shown, where Figure 6(a) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 2UAV-1ADS-1. Figure 6 (b) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 2UAV-1ADS-2. Figure 6 (c) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 3UAV-1ADS. Figure 6 (d) is a bar chart of the number of training rounds required for each algorithm strategy to reach the specified success rate in the scenario of 3UAV-2ADS. As Figure 6 shown in (a), when the task success rate of the strategy in the scenario of 2UAV-1ADS-1 reaches 60%, 80%, and 90%, the number of training rounds required by the PI-STF-RND algorithm is reduced by more than 34.3%, 47.2%, and 42.0% compared with the ECV-RND algorithm, and by more than 54.0%, 48.5%, and 42.0% compared with the No-Intrinsic-Reward algorithm.

[0097] The simulation results show that in the test scenario, compared with the existing algorithms, the exploration efficiency of this method for the state space and the non-equivalent state space is improved by more than 17.6% and 30.1% respectively, indicating that this method effectively reduces the repeated access of the UAV to the equivalent state during the exploration process.

[0098] In one embodiment, a multi-UAV cooperative strategy learning device driven by the equivalent state novelty value is provided, including: an equivalent state set construction module, a state transition module, an intrinsic reward determination module, an overall reward determination module, and a UAV policy optimization module, where: The equivalent state set construction module is used to determine the joint state space according to the multi-UAV collaborative decision-making model based on the distributed partially observable Markov decision process in the multi-UAV task scenario; and construct the equivalent state set of each state according to the joint state space.

[0099] The state transition module is used for each UAV to select an action according to its own policy and the current observation. After all UAVs execute the joint action and go through state transition, each UAV obtains a first experience tuple containing the observation, action, and extrinsic reward; the extrinsic reward of all UAVs is the same for each state transition.

[0100] The intrinsic reward determination module is used to uniformly represent the equivalent state of the state after state transition by using the permutation-invariant spatio-temporal feature representation; measure the equivalent state novelty value by using the random network distillation method according to the permutation-invariant spatio-temporal feature of the equivalent state to determine the intrinsic reward; and add the intrinsic reward to the experience storage after integrating it into the first experience tuple; all UAVs share the intrinsic reward.

[0101] The overall reward determination module is used for the UAVi Perform a weighted sum of the extrinsic rewards and intrinsic rewards to obtain the overall reward of the drone i ; repeatedly execute state transition and overall reward calculation until the decision-making round ends.

[0102] The drone policy optimization module is used to optimize the value network and policy network in the Actor-Critic architecture using the ACER method based on the overall reward and experience storage.

[0103] In one embodiment, the joint state space in the equivalent state set construction module includes the individual states of all drones and the states of other environmental factors; the individual state of the drone is as shown in the expression of the individual state of the above-mentioned drone ; the environmental state includes the radar antenna angle, channel occupancy, and target state.

[0104] In one embodiment, the equivalent state set construction module is further configured to extract the key observable part from the state variables in the joint state space to form sub-state variables; determine the sub-states of each state according to the sub-state variables; construct the equivalent sub-states and equivalent sub-state sets of each sub-state; when constructing the intrinsic reward based on the novelty value of the equivalent state, use the equivalent sub-states and equivalent sub-state sets to replace the equivalent states and equivalent state sets of the corresponding states in the original joint state space.

[0105] In one embodiment, the intrinsic reward determination module is further configured to uniformly represent the equivalent state of the state after state transition using permutation-invariant spatio-temporal feature representation to obtain the permutation-invariant spatio-temporal feature of the equivalent state; the permutation-invariant spatio-temporal feature of the equivalent state is as shown in the expression of the permutation-invariant spatio-temporal feature of the above-mentioned equivalent state.

[0106] In one embodiment, the intrinsic reward determination module is further configured to construct a stochastic distillation network; the stochastic distillation network includes a target network and a prediction network, the target network and the prediction network have the same structure and different initialized network parameters; the parameters of the target network remain fixed after self-initialization, and the parameters of the prediction network are updated according to the mean square error between its own output and the output of the target network using the gradient direction; input the permutation-invariant spatio-temporal feature of the equivalent state into the target network and the prediction network, and determine the novelty value metric of the equivalent state according to the outputs of the target network and the prediction network; the novelty value metric of the equivalent state is as described in the expression of the novelty value metric of the above-mentioned equivalent state; construct the intrinsic reward based on the novelty value of the equivalent state, and the intrinsic reward based on the novelty value of the equivalent state is as shown in the expression of the intrinsic reward based on the novelty value of the above-mentioned equivalent state.

[0107] In one embodiment, both the target network and the prediction network include: a convolutional layer, a rectified linear unit, a flattening layer, and two fully connected layers.

[0108] In one embodiment, the overall reward determination module is configured to perform a weighted sum of the extrinsic reward and the intrinsic reward of the drone i to obtain the overall reward of the drone i ; the overall reward is as described in the expression of the overall reward of the above-mentioned drone i .

[0109] In one embodiment, the second experience tuple in the experience store is: , where is the overall reward, is the policy, is the current observation, is the action, is the observation at the next moment, is the extrinsic reward, is the flag indicating whether the round terminates; the drone policy optimization module is further configured to randomly sample each drone from the experience store, and use the sampled experience data and the overall reward to update the value network and the policy network by using the ACER method; the ACER method includes: a policy evaluation link and a policy improvement link; in the policy evaluation link: the value network parameters are updated by minimizing the mean square error; the update method of the value network parameters is as described in the above update expression of the value network parameters.

[0110] In the policy improvement link, the parameters of the drone policy network are updated by using the ACER policy gradient; the update method of the policy network parameters is as described in the above update expression of the policy network parameters.

[0111] For the specific limitations of the multi-drone cooperation strategy learning device driven by the equivalent state novelty value, reference can be made to the limitations of the multi-drone cooperation strategy learning method driven by the equivalent state novelty value in the above text, which will not be elaborated here. Each module in the above-mentioned multi-drone cooperation strategy learning device driven by the equivalent state novelty value can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of brief description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0113] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A multi-UAV cooperative strategy learning method driven by equivalent state novelty value, characterized in that: The method comprises: In the multi-UAV mission scenario, the joint state space is determined according to the multi-UAV collaborative decision model based on distributed partially observable Markov decision process; According to the joint state space, construct an equivalent state set of each state; Each drone selects an action based on its own strategy and current observations. After all drones perform a joint action, after state transfer, each drone obtains the first experience tuple containing observations, actions, and external rewards. The equivalent states after state transfer are uniformly represented by using permutation-invariant spatiotemporal feature representation; According to the permutation-invariant spatiotemporal features of equivalent states, the novelty value of equivalent states is measured using the random network distillation method to determine the intrinsic reward; the intrinsic reward is incorporated into the first experience tuple and added to the experience storage; the extrinsic and intrinsic rewards of all drones are the same for each state transition; Drone i The weighted sum of the external reward and intrinsic reward of the drone is obtained. i The total reward of ; loop execution of state transfer and total reward calculation until the end of the decision round; According to the total reward and the experience storage, the ACER method is used to optimize the value network and the policy network in the Actor-Critic architecture.

2. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1 is characterized in that: The joint state space includes the individual states of all drones and the states of other environmental factors; The individual status is: ; in, UAV i x-coordinate, y-coordinate, yaw angle variable, speed variable; For drones Being threatened k Lock or block duration, , Threat drones i The number of drones; Environmental status includes radar antenna angle, channel occupancy, and target status.

3. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1 is characterized in that: According to the joint state space, an equivalent state set of each state is constructed, including: Extracting key and significant parts from the state variables in the joint state space to form sub-state variables; Determine the sub-state of each state according to the sub-state variables; Construct equivalent substates and equivalent substate sets for each substate; When constructing the intrinsic reward based on the novelty value of the equivalent state, the equivalent sub-state and the equivalent sub-state set are used to replace the equivalent state and the equivalent state set of the corresponding state in the original joint state space.

4. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1 is characterized in that: The equivalent states after state transfer are uniformly characterized by using permutation-invariant spatiotemporal features, and the permutation-invariant spatiotemporal features of the equivalent states are obtained as follows: ; in, Equivalent state The permutation-invariant spatiotemporal characteristics of n is the total number of drones, , , Threat drones i The number of drones; , , , , Represents the scope of the rectangular task area; Used to indicate the location of the drone. Located in the grid In , otherwise, ; For drones Being threatened Lock or block duration The linear mapping of , and UAV Being threatened A mapping of values ​​to the lower and upper bounds of the lock or intercept duration.

5. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1 is characterized in that: According to the equivalent state permutation invariant spatiotemporal features, the random network distillation method is used to measure the equivalent state novelty value and determine the intrinsic reward, including: Constructing a random distillation network; the random distillation network includes a target network and a prediction network, the target network and the prediction network have the same structure, and the initialized network parameters are different; the parameters of the target network remain fixed after initialization, and the parameters of the prediction network are updated according to the mean square error between its own output and the output of the target network using the gradient direction; The permutation-invariant spatiotemporal features of the equivalent state are input into the target network and the prediction network, and the novelty value metric of the equivalent state is determined according to the outputs of the target network and the prediction network: ; The intrinsic reward based on the equivalent state novelty value is constructed according to the equivalent state novelty value: ; in, is the intrinsic reward based on the novelty value of the equivalent state, They are the current state, the current action, and the next state. It is the equivalent state novelty value measure of the state at the next moment.

6. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 5 is characterized in that: The target network and the prediction network both include: a convolutional layer, a linear rectification function, a flattening layer, and two fully connected layers.

7. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1 is characterized in that: Drone i The weighted sum of the external reward and intrinsic reward of the drone is obtained. i The total reward is: ; in, For drones i The total reward, and are the weights of the extrinsic reward and the intrinsic reward based on the equivalent situation novelty value, For external rewards, is an intrinsic reward based on the equivalent situation novelty value.

8. The multi-UAV cooperative strategy learning method driven by equivalent state novelty value according to claim 1 is characterized in that: The second experience tuple in the experience storage is: ,in, For the overall reward, For strategy, For the current observation, For action, For the next moment observation, For external rewards, It is a flag indicating whether the round is terminated; According to the total reward and the experience storage, the ACER method is used to optimize the value network and the policy network in the Actor-Critic architecture, including: Each drone randomly samples from the experience storage, and uses the sampled experience data and the overall reward to update the value network and the strategy network using the ACER method; the ACER method includes: a strategy evaluation link and a strategy improvement link; In the strategy evaluation phase: the value network parameters are updated by minimizing the mean square error; the update expression of the value network parameters is: ; ; ; in, is the mean square error of the value network, are the parameters of the value network, is the learning rate of the drone value network, It is an estimate of the value of the current strategy based on the Retrace method. For the overall reward, For the value network, UAV i The current observation and action of For drones i The parameters of the policy network, is a hyperparameter, is the truncated importance sampling coefficient, For drones Generate experience The historical strategy at that time is recorded as the behavioral strategy; In the strategy improvement phase, the ACER policy gradient is used to update the parameters of the drone policy network; the update expression of the policy network parameters is: ; ; in, are the parameters of the policy network, is the learning rate of the drone strategy network, is the ACER policy gradient, The current strategy for drones, is the value of the current observation, For the cost.

9. A multi-UAV cooperative strategy learning device driven by equivalent state novelty value, characterized in that: The device comprises: An equivalent state set construction module is used to determine a joint state space in a multi-UAV mission scenario according to a multi-UAV collaborative decision model based on a distributed partially observable Markov decision process; and to construct an equivalent state set of each state according to the joint state space; The state transfer module is used for each drone to select an action based on its own strategy and current observation. After all drones perform a joint action, after the state transfer, each drone obtains the first experience tuple containing observations, actions and external rewards. The external rewards of all drones are the same for each state transfer. An intrinsic reward determination module is used to uniformly characterize the equivalent states of the states after the state transfer by using a permutation-invariant spatiotemporal feature representation; to measure the novelty value of the equivalent states by using a random network distillation method according to the permutation-invariant spatiotemporal features of the equivalent states, and to determine the intrinsic reward; and to incorporate the intrinsic reward into the first experience tuple and add it to the experience storage; all drones share the intrinsic reward; The overall reward determination module is used to i The weighted sum of the external reward and intrinsic reward of the drone is obtained. i The total reward of ; loop execution of state transfer and total reward calculation until the end of the decision round; The drone strategy optimization module is used to optimize the value network and the strategy network in the Actor-Critic architecture using the ACER method according to the overall reward and the experience storage.

Citation Information

Patent Citations

  • Bit rate optimization algorithm based on DDPG for energy collectible communication

    CN109548044A

  • Unmanned driving training method for incomplete information scene in sparse high-dimensional state

    CN115965879A

  • Multi-unmanned aerial vehicle air combat strategy generation method and device and computer equipment

    CN116430888A

  • Multi-energy cooperative control method, device and equipment of power grid system and storage medium

    CN118040788A