A Multi-Agent Cooperative Navigation Method Based on Deep Reinforcement Learning
By introducing R-Drop and self-attention mechanisms into the collision avoidance policy network and combining the priority experience replay mechanism, the problems of overfitting and insufficient sample utilization in multi-agent collaborative navigation are solved, and the generalization ability and navigation performance of the model are improved.
Patent Information
- Application Number
- CN202310550857.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-05-16
AI Technical Summary
The existing deep reinforcement learning algorithms are prone to overfitting, insufficient global strategy optimization, and low sample utilization in multi-agent collaborative navigation, resulting in insufficient generalization capabilities of the model and affecting the collaborative navigation performance.
The R-Drop mechanism and self-attention mechanism are introduced in the collision avoidance policy network, combined with the priority experience playback mechanism, and optimize the policy network by randomly deleting hidden layer nodes, filtering important environmental information, and improving sample utilization.
It improves the generalization ability of the model, enhances the collaborative navigation performance of the agent in complex environments, improves the navigation success rate and reduces navigation time.
Smart Images

Figure CN116579372B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent cooperative navigation, and particularly relates to a multi-agent cooperative navigation method based on deep reinforcement learning. Background Art
[0002] Multi-agent cooperative navigation is an important basis for multi-agent systems to complete cooperative tasks and has received extensive attention in recent years. It requires agents to have the ability to coordinate with each other to execute tasks in a complex environment and avoid collisions during the task process to ensure their own safety. Compared with single agents, multi-agent systems that achieve cooperative navigation can complete tasks more efficiently, improve the fault tolerance of the system and the adaptability to the environment. Multi-agent cooperative navigation has a wide range of application scenarios, and some applications include multi-robot formation control, multi-robot target search, and autonomous mobile service robots. The deep reinforcement learning algorithm inherits the superiority of the deep learning algorithm in perception and feature extraction, and realizes end-to-end learning by mapping the state information of the agent to the feature space without artificial feature design. As for the large amount of data required for training, it is generated at low cost through interaction with the simulation environment, thus easily solving the sample problem. These characteristics make deep reinforcement learning one of the hottest research and application directions in the fields of multi-agent and artificial intelligence.
[0003] Existing patented technologies, such as "A Multi-agent Navigation Algorithm Based on Deep Reinforcement Learning", with the authorization publication number "CN113218400B". This method integrates the A* algorithm into the PPO algorithm. The former is a path planning method, and the latter is a deep reinforcement learning method. This method uses a designed reward and punishment function to achieve a deep integration of the two algorithms. The agent inputs the original image data of the sensor and decides and plans the best action path to reach the target point. Since this method needs to extract the features of the image information obtained by the scanner and obtain the low-dimensional environmental features through convolutional neural network training, the process of image processing is time-consuming and also places relatively high requirements on the performance of the device, and the training process is relatively long.
[0004] The hierarchical stable multi-agent deep reinforcement learning algorithm can well learn the end-to-end solution for multi-agent cooperative navigation. This algorithm directly maps the original sensor data to control signals instead of using a planning-based method. Specifically, the training phase of this algorithm is carried out in a random environment, during which the agents can learn cooperation strategies. Once the strategy is learned, it is deployed to each agent to complete cooperative navigation in an unknown environment without time-consuming planning and information exchange operations regarding target selection. However, there are problems in this algorithm model such as being prone to overfitting, insufficient global optimality of the strategy, and low utilization rate of samples during the training process. Therefore, it is necessary to improve the algorithm model network to enhance the model generalization ability and improve the cooperative navigation performance. Summary of the Invention
[0005] In order to overcome the above problems existing in the prior art, the purpose of the present invention is to provide a multi-agent cooperative navigation method based on deep reinforcement learning. The self-attention mechanism is used in the collision avoidance policy network to process the state sequence information of other agents, enabling the agents to selectively screen out important environmental information, thereby achieving the purpose of optimizing the policy. And the R-Drop mechanism is added during its training process to improve the overfitting problem of the model by randomly deleting hidden layer nodes and improving the loss function. At the same time, the prioritized experience replay mechanism is used to give a higher sampling rate to samples with greater importance to improve the utilization rate of samples. Thus, this method improves the generalization ability of the model and thereby enhances the cooperative navigation performance.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A multi-agent cooperative navigation method based on deep reinforcement learning, where a motion device is taken as an agent or controlled by an agent, and each agent performs the following steps;
[0008] Step 1, observe the global state, where the global state refers to the relative position coordinates of all targets detected by the current motion device and other agents; construct a target selection policy network and a collision avoidance policy network;
[0009] Step 2, select a target according to the target selection policy network, where the target refers to the target location that the current motion device needs to navigate to;
[0010] Step 3, observe the local state, where the local state refers to the distance between the current motion device and surrounding obstacles detected;
[0011] Step 4: Determine whether there is an obstacle ahead. If not, the current moving device moves one step towards the selected target and returns to Step 1; if so, obtain an angle according to the collision avoidance policy network, the current moving device turns to this angle and moves one step forward, and returns to Step 1; the "forward" refers to the direction of the angle to which the current moving device turns.
[0012] The target selection policy network, this neural network has an input layer, two hidden layers and an output layer; this network is trained in an obstacle-free environment: store the state information observed in each round into the experience replay pool of the experience replay mechanism module, and extract samples from the experience replay pool and input them into the target selection policy network, then input the predicted value and the true value of the network into the loss function to obtain the loss value, and its expression is And update the neural network parameters according to the stochastic gradient descent method, where represents taking the expectation, i represents the i-th agent, t represents the t-th time step, represents the actual reward value, G ts represents the action value function value output by the corresponding target selection policy network.
[0013] The collision avoidance policy network in the above Step 1 specifically includes:
[0014] (1) Train and test the collision avoidance policy network to evaluate the performance metrics of the model; the collision avoidance policy network is used for collision avoidance.
[0015] (2) Add an R-Drop mechanism module to the collision avoidance policy network; when an obstacle is observed in the direction within the perception range of the agent, the network outputs a turning angle to guide the agent to avoid obstacles.
[0016] The R-Drop is a regularization method, which is a variant of the Dropout method, used to alleviate the overfitting problem of the model and improve the generalization ability of the model.
[0017] (3) On the basis of step (2), add a self-attention mechanism module to the collision avoidance policy network. The self-attention mechanism module preprocesses the state sequence information of other agents and is used to screen environmental information of different importance levels, so as to improve the collaborative navigation ability of the model.
[0018] (4) On the basis of step (3), replace the experience replay mechanism module used in the training process of the collision avoidance policy network with a prioritized experience replay mechanism module. The prioritized experience replay mechanism module performs sample extraction based on the priority of samples during the training process of the network model. The priority represents the importance of the samples. The larger the priority of a sample, the higher its corresponding sampling rate, the higher the utilization rate of the sample, and the probability that the model learns a good policy is increased, thus obtaining an improved network model;
[0019] (5) Train and test the improved network model for navigation, and evaluate the performance metrics of the model.
[0020] (6) Perform navigation with the improved network model after training and testing.
[0021] In steps (1) and (5), the process of training and testing the network model and the improved algorithm network model includes the following steps:
[0022] (1) Train the target selection strategy of the algorithm network model in a barrier-free environment; store the state information observed in each episode into the experience replay pool of the experience replay mechanism module, and extract samples from the experience replay pool and input them into the target selection strategy network. Then, input the predicted value and the true value of the network into the loss function to obtain the loss value, and its expression is And update the neural network parameters according to the stochastic gradient descent method, where, represents taking the expectation, i represents the i-th agent, t represents the t-th time step, represents the actual reward value, G ts represents the action value function value output by the corresponding target selection strategy network;
[0023] (2) Repeat step (1) until episode reaches 10,000 rounds and then end. At this time, the target selection strategy network has converged;
[0024] (3) Use the trained target selection strategy as a warm start to train the collision avoidance strategy in an environment where obstacles are unknown and randomly set; store the state information observed in each episode into the experience replay pool, and extract samples from the experience replay pool and input them into the collision avoidance strategy network. Then, input the predicted value and the true value of the network into the loss function to obtain the loss value, and update the neural network parameters according to the stochastic gradient descent method;
[0025] (4) Repeat step (3) until episode reaches 10,000 rounds and then end. At this time, the collision avoidance strategy network has converged;
[0026] (5) Test the performance of the algorithm model in an environment where obstacles are unknown and randomly set, generate 1000 test tasks, and evaluate with success rate and normalized average maximum navigation time as performance metrics.
[0027] In step (2), in each training step, given the input data pair (x i , y i ), input x i into the forward channel of the network twice, thus obtaining two distributions predicted by the model, denoted as P1 and P2 respectively; the two forward passes are indeed based on two different sub-models, and the neurons randomly deleted when the sample x i passes through the model with Dropout twice are different. Therefore, for the same input data pair (x i , y i ), the two distributions P1 and P2 predicted by the model are different;
[0028] Then, during training, the R-Drop mechanism module regularizes the model prediction by minimizing the bidirectional KL divergence between these two output distributions of the same sample, thereby reducing the difference between the training and test models. The bidirectional KL divergence usually takes half of the sum of the two KL divergences of the two distributions, and its expression is where D KL represents the KL divergence between the two output distributions;
[0029] After adding the R-Drop mechanism module, the loss function of the network becomes
[0030]
[0031] where, and are the losses of the sample's two forward passes in the algorithm collision avoidance policy network respectively, and are the output values of the sample's two forward passes in the algorithm collision avoidance policy network respectively, D KL represents the KL divergence between the two output distributions, and α is the weight coefficient for controlling this divergence.
[0032] Step (3) is specifically as follows:
[0033] Use the self-attention mechanism to preprocess the observation sequence information of other agents, enabling the agent to selectively filter out important environmental information and ignore unimportant information;
[0034] The loss of the sample during the forward pass in the algorithm collision avoidance policy network is
[0035]
[0036] Among them, represents taking the expectation, i represents the i-th agent, and t represents the t-th time step. represents the actual reward value, G ca represents the action value function value output by the corresponding collision avoidance policy network. represents the observation value. represents the environmental information from other agents, and this value is output by the self-attention mechanism module. represents the action value output by the corresponding collision avoidance policy network.
[0037] The prioritized experience replay module in step (4) is a special binary tree, where the value of each node is the sum of the values of its child nodes, and the priority of the sample is used as the leaf node. The observed state information is the sample data. Before storing the sample in the experience replay pool, it is necessary to calculate the priority of each sample, so that a corresponding relationship can be established between the priority of the leaf node of the binary tree and the sample data. The sum of the priorities of all samples in the sample experience pool is the priority of the root node.
[0038] In step (6), during the cooperative navigation process, each agent moves at a constant speed of v = 1 meter per time step and has a steering angle that varies within the range of [-(π / 2), (π / 2)]. If all agents reach different targets respectively, the navigation is successful; if a collision occurs, the navigation fails. Both situations end this round. In each time step, each agent first selects a target based on its observation of the global state, then rotates its field of view center to the selected target and observes the local state. If no obstacle is observed in the direction within the sensing range, the agent will move directly towards the target; otherwise, the lower-level policy is activated to output an angle, and the agent will turn to this angle and move forward. Step (1) is to train and test the original algorithm network model, and step (5) is to train and test the improved network model, and the results need to be compared.
[0039] Advantages of the present invention:
[0040] This method adds an R-Drop mechanism module to the collision avoidance policy network to alleviate the overfitting phenomenon of network learning. The self-attention mechanism module is used to preprocess the sequence of state information about other agents input into the collision avoidance policy network to enhance the agent's ability to screen important environmental information. The prioritized experience replay method is used for the storage and extraction of samples during the training process to improve the utilization rate of samples and increase the probability of learning good policies.
[0041] The present invention solves the following technical problems: First, the original deep reinforcement learning algorithm model is prone to overfitting during the training process; second, the difference in the influence degree of the states of other agents on the current agent is not considered, resulting in insufficient global optimality of the policy; third, the utilization rate of samples during the network training process is not high. Description of the Drawings
[0042] Figure 1 It is a schematic flowchart of a multi-agent collaborative navigation method based on deep reinforcement learning according to the present invention.
[0043] Figure 2 It is a schematic structural diagram of the R-Drop module.
[0044] Figure 3 It is a schematic structural diagram of a collision avoidance policy network with a self-attention mechanism module added.
[0045] Figure 4 It is a structure diagram of the storage and sampling data of the prioritized experience replay module.
[0046] Figure 5 It is a comparison chart of the collaborative navigation performance based on the improved deep reinforcement learning network model. Detailed Embodiment
[0047] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0048] To deepen the understanding of the present invention, the present invention will be further described in detail below with reference to the drawings. The present invention proposes a multi-agent collaborative navigation method based on deep reinforcement learning, including the following steps:
[0049] S1: Train and test the collision avoidance policy network, and evaluate the performance indicators of the model;
[0050] S2: Add an R-Drop mechanism module to the collision avoidance policy network;
[0051] S3: Add a self-attention mechanism module to the collision avoidance policy network;
[0052] S4: Replace the experience replay mechanism module used in the training process of the collision avoidance policy network with a prioritized experience replay mechanism module;
[0053] S5: Train and test the improved network model, and evaluate the performance indicators of the model.
[0054] The R-Drop mechanism module performs random deletion operations on the hidden layer nodes of the collision avoidance policy network, and adds the bidirectional KL divergence between the output distributions of two sub-models randomly sampled from the same sample to the loss function to regularize and alleviate the inconsistency between models, improve the overfitting problem of the model, and enhance the generalization ability of the model.
[0055] The self-attention mechanism module is used to preprocess the state sequence information of other agents, and replaces the original feature vector directly input into the collision avoidance policy network with the feature vector output by the self-attention mechanism module, enabling the agent to selectively filter out important environmental information and ignore unimportant information, thereby learning a better collaborative navigation performance strategy.
[0056] The prioritized experience replay mechanism module samples based on the importance of samples, gives a higher sampling rate to samples with greater importance, improves the utilization rate of samples, and also increases the probability of the model learning a good strategy, thereby enhancing the collaborative navigation performance of the model.
[0057] Furthermore, the process of training and testing the network model includes the following steps:
[0058] S51: Train the target selection strategy of the algorithm model in a barrier-free environment. Store the state information observed in each episode in the experience replay pool, and extract samples from the experience replay pool and input them into the target selection policy network. Then, input the predicted value and the true value of the network into the loss function to obtain the loss value, and update the neural network parameters according to the stochastic gradient descent method;
[0059] S52: Repeat step (1) until the episode reaches 10,000 rounds and then ends. At this time, the target selection policy network has converged.
[0060] S53: Use the trained target selection strategy as a warm start to train the collision avoidance strategy in an environment with unknown and randomly placed obstacles. Store the state information observed in each episode in the experience replay pool, and extract samples from the experience replay pool and input them into the collision avoidance policy network. Then, input the predicted value and the true value of the network into the loss function to obtain the loss value, and update the neural network parameters according to the stochastic gradient descent method;
[0061] S54: Repeat step (3) until the episode reaches 10,000 rounds and then ends. At this time, the collision avoidance policy network has converged.
[0062] S55: Test the performance of the algorithm model in an environment with unknown and randomly placed obstacles. Generate 1,000 test tasks and evaluate them using the success rate and the normalized average maximum navigation time as performance metrics.
[0063] The present invention will be further described below in conjunction with relevant background technologies and implementation steps:
[0064] In step S1 of the present invention, performance evaluation indicators are recorded, including the success rate and the normalized average maximum navigation time, which are used as the criteria for subsequent performance evaluation.
[0065] The hierarchical stable multi-agent deep reinforcement learning algorithm can well learn the end-to-end solution of multi-agent collaborative navigation. This algorithm directly maps the original sensor data to the control signal instead of using a planning-based method. Specifically, the training phase of this algorithm is carried out in a random environment, during which the agents can learn cooperation strategies. Once the strategy is learned, the strategy is deployed to each agent to complete collaborative navigation in an unknown environment without time-consuming planning and the exchange operation of target selection information. However, there are problems in this algorithm model such as being prone to overfitting, insufficient global optimality of the strategy, and low utilization rate of samples during the training process. Therefore, it is necessary to improve the algorithm model network to improve the model generalization ability and enhance the collaborative navigation performance.
[0066] (1) R-Drop mechanism module
[0067] In step S2, an R-Drop mechanism module is added to the collision avoidance policy network. The structural schematic diagram of the R-Drop module is as Figure 2 shown. In each training step, given the input data pair (x i , y i ), x i is input into the forward channel of the network twice. Thus, two distributions predicted by the model can be obtained, denoted as P1 and P2 respectively. Although it is in the same model, since Dropout randomly deletes neurons in the model, the two forward passes are indeed based on two different sub-models. When the sample x i passes through the model with Dropout twice, the randomly deleted neurons are different. Therefore, for the same input data pair (x i , y i ), the two distributions P1 and P2 predicted by the model are different.
[0068] Then, during the training process, the R-Drop method regularizes the model prediction by minimizing the bidirectional KL divergence between these two output distributions of the same sample, thereby narrowing the difference between the training and test models. The bidirectional KL divergence usually takes half of the sum of the two KL divergences of the two distributions, and its expression is where D KL represents the KL divergence between the two output distributions.
[0069] After adding the R-Drop mechanism module, the loss function of the network becomes
[0070]
[0071] Among them, and are the losses of the two forward passes of the sample in the algorithm collision avoidance policy network respectively, and are the output values of the two forward passes of the sample in the algorithm collision avoidance policy network respectively. D KL represents the KL divergence between the two output distributions, and α is the weight coefficient for controlling this divergence.
[0072] The R-Drop mechanism processes both the hidden layer neurons and the output of the sub-model with Dropout sampling, which not only prevents the overfitting of the model but also alleviates the inconsistency between models that Dropout may bring, and has excellent performance.
[0073] (2) Self-attention mechanism module
[0074] In step S3, a self-attention mechanism module is added to the collision avoidance policy network. The schematic diagram of the structure of the collision avoidance policy network with the self-attention mechanism module added is as Figure 3 shown. The self-attention mechanism is used to preprocess the observation sequence information of other agents, and the feature vector output by the self-attention module is used to replace the original feature vector directly input into the collision avoidance policy network, enabling the agent to selectively filter out important environmental information and ignore unimportant information.
[0075] The loss of the sample during the forward pass in the algorithm collision avoidance policy network is
[0076]
[0077] Among them, represents the expectation, i represents the i-th agent, t represents the t-th time step, represents the actual reward value, G ca represents the action value function value output by the corresponding collision avoidance policy network, represents the observation value, represents the environmental information from other agents, and this value is obtained from the output of the self-attention mechanism module, represents the action value output by the corresponding collision avoidance policy network. (3) Prioritized experience replay module
[0078] In step S4, a prioritized experience replay module is added to the collision avoidance policy network. The structure diagram of the prioritized experience replay module for storing and sampling data is as Figure 4As shown. This is a special binary tree where the value of each node is the sum of the values of its child nodes, and the priority of the sample is used as the leaf node. The observed state information is the sample data. Before storing the sample in the experience replay pool, it is necessary to calculate the priority of each sample, so that a corresponding relationship can be established between the priority of the leaf nodes of the binary tree and the sample data. The sum of the priorities of all samples in the sample experience pool is the priority of the root node.
[0079] The specific sampling process is as follows: During sampling, divide the total priority corresponding to the root node by the size of the mini-batch sampling into several intervals, and perform uniform random sampling in each of the divided intervals. If the priority of the root node is 42 and 6 samples are drawn, the priorities of each interval are [0-7], [7-14], [14-21], [21-28], [28-35], [35-42]. For example, if 24 is selected in the interval [21-28], then the entire extraction process starts from the top 42 and searches downwards. At this time, there are two child nodes under the top 42. First, compare 24 with 29 at the lower left corner of 42. Since the value at the lower left corner is larger than the value in hand, take the left path. Then compare it with 13 at the lower left corner of 29. Since 24 in hand is larger than 13, take the right path, and modify the value in hand according to 13, becoming 24–13 = 11. Then compare 11 with 12 at the lower left corner of 13. The result is that 12 is larger than 11, and 12 corresponds to a leaf node, so select 12 as the priority selected this time, and also select the sample data corresponding to 12. The sampling process for each interval is carried out in this way, and finally 6 samples are obtained.
[0080] The prioritized experience replay mechanism samples based on the importance of the samples, gives a higher sampling rate to the more important samples, improves the utilization rate of the samples, and also increases the probability that the model learns good strategies, thereby enhancing the cooperative navigation performance of the model.
[0081] The present invention relates to three parts: the R-Drop regularization method, the self-attention mechanism, and the prioritized experience replay mechanism. Add an R-Drop mechanism module to the collision avoidance policy network to alleviate the overfitting phenomenon of network learning. Use the self-attention mechanism module to preprocess the sequence of state information about other agents input into the collision avoidance policy network to enhance the agent's ability to screen important environmental information. Use the prioritized experience replay method for storing and extracting samples during the training process to improve the utilization rate of samples and increase the probability of learning good strategies. Compared with the original algorithm, the cooperative navigation performance of the improved deep reinforcement learning algorithm model has been effectively improved.
[0082] The comparison chart of the cooperative navigation performance of the network model before and after improvement is as Figure 5As shown, the method proposed by the present invention is superior to the original algorithm network model in terms of both navigation success rate and normalized average maximum navigation time. The normalized average maximum navigation time refers to the average of the time steps taken by the last agent to reach the target divided by the maximum number of time steps in a round, and the maximum number of time steps in a round is set to twice the side length of the environment. When the number of agents is 4, the success rate of the method proposed by the present invention is 95.9%, with the largest difference in improvement compared to the original algorithm, which is 5.6%. When the number of agents is 5, the normalized average maximum navigation time of the method proposed by the present invention is 0.492, with the largest reduction amplitude compared to the original algorithm, which is 10.71%.
[0083] In summary, the present invention mainly solves three technical problems. First, the original deep reinforcement learning algorithm model is prone to overfitting during the training process. Second, the problem that the global optimality of the strategy is insufficient because the differences in the influence degrees of the states of other agents on the current agent are not considered. Third, the problem of low utilization rate of samples during the network training process.
[0084] The above-disclosed is only one example of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand the process of implementing the above example and make equivalent changes according to the claims of the present invention.
Claims
1. A multi-agent collaborative navigation method based on deep reinforcement learning, characterized in that, Including the following steps; Taking a motion device as an agent or being controlled by an agent, and each agent performs the following steps; Step 1, observing the global state, where the global state refers to all the targets detected by the current motion device and the relative position coordinates of other agents; constructing a target selection policy network and a collision avoidance policy network; Step 2, selecting a target according to the target selection policy network, where the target refers to the target location that the current motion device needs to navigate to; Step 3, observing the local state, where the local state refers to the distance between the current motion device and the surrounding obstacles detected; Step 4, judging whether there is an obstacle in front. If not, the current motion device moves one step towards the selected target and returns to Step 1; if so, obtaining an angle according to the collision avoidance policy network, the current motion device turns to this angle and moves forward one step, and returns to Step 1; the "forward" refers to the direction of the angle to which the current motion device turns; The collision avoidance policy network in Step 1 specifically includes: (1) Training and testing the collision avoidance policy network, and evaluating the performance indicators of the model; (2) Adding an R-Drop mechanism module to the collision avoidance policy network; when an obstacle is observed in the direction within the perception range of the agent, the network outputs a turning angle to guide the agent to avoid obstacles; (3) On the basis of (2), adding a self-attention mechanism module to the collision avoidance policy network, and the self-attention mechanism module preprocesses the state sequence information of other agents, which is used to screen environmental information of different importance levels, so as to improve the collaborative navigation ability of the model; (4) On the basis of (3), replacing the experience replay mechanism module used in the training process of the collision avoidance policy network with a prioritized experience replay mechanism module to obtain an improved network model; (5) Training and testing the improved network model for navigation, and evaluating the performance indicators of the model; (6) Conducting navigation with the improved network model after training and testing.
2. The multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, wherein In step (1), training and testing the collision avoidance policy network includes the following steps: (1) The target selection strategy for training the algorithm network model in the barrier-free environment; storing the state information observed in each round into the experience replay pool of the experience replay mechanism module, and extracting samples from the experience replay pool and inputting them into the target selection strategy network, then inputting the predicted value and the true value of the network into the loss function to obtain the loss value, and its expression is And updating the neural network parameters according to the stochastic gradient descent method, where represents taking the expectation, i represents the i-th agent, and t represents the t-th time step. represents the actual reward value, G ts represents the action value function value output by the corresponding target selection strategy network. (2) Repeating step (1) until the episode reaches 10,000 rounds and then ending. At this time, the target selection policy network has converged; (3) Using the trained target selection policy as a warm start, training the collision avoidance policy in an environment where the obstacles are unknown and randomly set; storing the state information observed in each episode into the experience replay pool, and extracting samples from the experience replay pool and inputting them into the collision avoidance policy network, then inputting the predicted value and the true value of the network into the loss function to obtain the loss value, and updating the neural network parameters according to the stochastic gradient descent method; (4) Repeating step (3) until the episode reaches 10,000 rounds and then ending. At this time, the collision avoidance policy network has converged; (5) Testing the performance of the algorithm model in an environment where the obstacles are unknown and randomly set, generating 1,000 test tasks, and evaluating with the success rate and the normalized average maximum navigation time as the performance indicators.
3. A multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, characterized in that In step (5), training and testing of the improved network model for navigation are carried out, including the following steps: (1) The target selection strategy for training the algorithm network model in a barrier-free environment; storing the state information observed in each round into the experience replay pool of the experience replay mechanism module, and extracting samples from the experience replay pool and inputting them into the target selection strategy network. Then, inputting the predicted value and the true value of the network into the loss function to obtain the loss value, and its expression is And updating the neural network parameters according to the stochastic gradient descent method, where represents taking the expectation, i represents the i-th agent, and t represents the t-th time step. represents the actual reward value, and G ts represents the action value function value output by the corresponding target selection strategy network. (2) Repeat step (1) until the episode reaches 10,000 rounds and then ends. At this time, the target selection policy network has converged; (3) Using the trained target selection policy as a warm start, train the collision avoidance policy in an environment where obstacles are unknown and randomly set; store the state information observed in each episode into the experience replay pool, and extract samples from the experience replay pool and input them into the collision avoidance policy network. Then, input the predicted value and the true value of the network into the loss function to obtain the loss value, and update the neural network parameters according to the stochastic gradient descent method; (4) Repeat step (3) until the episode reaches 10,000 rounds and then ends. At this time, the collision avoidance policy network has converged; (5) Test the performance of the algorithm model in an environment where obstacles are unknown and randomly set, generate 1,000 test tasks, and evaluate using the success rate and the normalized average maximum navigation time as performance metrics.
4. A multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, characterized in that, In the step (2), in each training step, given the input data pair (x i , y i ), x i is input into the forward channel of the network twice, thereby obtaining two distributions predicted by the model, denoted as P1 and P2 respectively; the two forward passes are indeed based on two different sub-models, and the neurons randomly deleted when the sample x i passes through the model with Dropout twice are different. Therefore, for the same input data pair (x i , y i ), the two distributions P1 and P2 predicted by the model are different; Then, during the training process, the R-Drop mechanism module regularizes the model prediction by minimizing the bidirectional KL divergence between the two output distributions of the same sample, thereby reducing the difference between the training and test models. The bidirectional KL divergence usually takes half of the sum of the two KL divergences of the two distributions, and its expression is where D KL represents the KL divergence between the two output distributions.
5. A multi-agent collaborative navigation method based on deep reinforcement learning according to claim 4, characterized in that, After adding the R-Drop mechanism module, the loss function of the collision avoidance policy network becomes Among them, and are the losses of the two forward passes of the sample in the algorithm collision avoidance policy network respectively, and are the output values of the two forward passes of the sample in the algorithm collision avoidance policy network respectively. D KL represents the KL divergence between the two output distributions, and α is the weight coefficient for controlling this divergence.
6. The multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, wherein The specific step (3) is as follows: Use the collision avoidance policy network with a self-attention mechanism module added to preprocess the observation sequence information of other agents, enabling the agent to selectively filter out important environmental information and ignore unimportant information; The loss of the sample during forward propagation in the algorithm collision avoidance policy network is Among them, represents the expectation, i represents the i-th agent, and t represents the t-th time step. represents the actual reward value, G ca represents the action-value function value output by the corresponding collision avoidance policy network. represents the observation value. represents the environmental information from other agents, and this value is obtained by the output of the self-attention mechanism module. represents the action value output by the corresponding collision avoidance policy network.
7. A multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, characterized in that The prioritized experience replay module in step (4) is a special binary tree, where the value of each node is the sum of the values of its child nodes, and the priority of the sample is used as the leaf node. The observed state information is the sample data. Before storing the sample into the experience replay pool, it is necessary to calculate the priority of each sample, so that a correspondence can be established between the priority of the leaf node of the binary tree and the sample data. The sum of the priorities of all samples in the sample experience pool is the priority of the root node.
8. A multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, characterized in that, In step (6), during the cooperative navigation process, each agent moves at a constant speed v = 1 meter per time step and has a steering angle that varies within the range of [-(π / 2), (π / 2)]. If all agents reach different targets respectively, it is considered a successful navigation. If a collision occurs, the navigation fails. Both situations end this episode. In each time step, each agent first selects a target according to its observation of the global state, then rotates its field of view center to the selected target and observes the local state. If no obstacle is observed in the direction within the sensing range, the agent will move directly towards the target. Otherwise, the lower-level policy is activated to output an angle, and the agent will turn to this angle and move forward.
9. A multi-agent collaborative navigation method based on deep reinforcement learning according to claim 1, characterized in that The target selection policy network has an input layer, two hidden layers and an output layer; the network is trained in an obstacle-free environment: the state information observed in each round is stored in the experience replay pool of the experience replay mechanism module, and samples are drawn from the experience replay pool and input into the target selection policy network. Then, the predicted value and the true value of the network are input into the loss function to obtain the loss value, and its expression is and the neural network parameters are updated according to the stochastic gradient descent method, where represents the expectation, i represents the i-th agent, and t represents the t-th time step. represents the actual reward value, and G ts represents the action value function value output by the corresponding target selection policy network.
Citation Information
Patent Citations
A multi-agent navigation algorithm based on deep reinforcement learning
CN113218400B
Mobile robot social navigation method based on reinforcement learning
CN115456851A
Unmanned aerial vehicle autonomous obstacle avoidance method and system based on self-attention mechanism and 2DPCA
CN115984722A