Unmanned vehicle-unmanned aerial vehicle cluster cooperative hunting method and system in urban roadway environment
By introducing a multi-headed attention mechanism and event-triggered communication strategy into the collaborative roundup system of unmanned vehicle-drone clusters, the problems of communication load and delay in traditional D3PG are solved, and a more efficient and stable collaborative roundup effect is achieved to adapt to changes in complex environments.
Patent Information
- Application Number
- CN202510427140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
In the existing unmanned vehicle-drone cross-domain cluster collaborative roundup technology, traditional D3PG relies on frequent communication exchanges to synchronize state information, increase network load and delay, and is difficult to adapt to unprecedented scenarios, weakening the robustness and flexibility of the system.
The multi-head attention mechanism and event-triggered communication strategy are introduced, and the target network with the same initial weight as the main network is implemented to achieve accurate information sharing and state updates at necessary moments, combining meta-learning and soft update mechanisms to enhance the adaptability and stability of the system.
It improves communication efficiency, enhances the generalization ability of the system and the speed of adapting to new environments, reduces unnecessary information redundancy, reduces resource consumption, and improves the system's real-time response and scalability.
Smart Images

Figure CN120276491A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV cooperation, and particularly to a method and system for collaborative pursuit of unmanned vehicles - UAV clusters in an urban alley environment. Background Art
[0002] With the acceleration of the urbanization process and the change of complex combat environments, traditional single - soldier or single unmanned systems are difficult to meet the requirements of modern tasks. Especially in the urban alley environment, where the terrain is intricate and buildings are dense, traditional manual combat methods can no longer meet the high requirements of modern warfare. In recent years, the cross - domain cooperation technology of intelligent unmanned cluster systems has become a research hotspot globally. This technology can achieve information sharing, collaborative decision - making, and efficient execution among intelligent devices such as unmanned vehicles and UAVs, greatly improving combat efficiency and safety.
[0003] As an emerging combat mode, intelligent unmanned cluster systems demonstrate great potential in multiple fields such as public security monitoring, disaster relief, and military reconnaissance due to their flexibility, efficiency, and adaptability. Compared with the independent operation of traditional single unmanned systems, multi - agent collaborative work can better handle diverse requirements in complex environments and provide more comprehensive task support. Especially in the dynamically changing urban environment, the cross - domain cluster collaborative pursuit technology of unmanned vehicles - UAVs has become a key means to solve this problem.
[0004] In the existing cross - domain cluster collaborative pursuit technology of unmanned vehicles - UAVs, a Distributed Deep Deterministic Policy Gradient algorithm (D3PG) based on the Actor - Critic architecture is currently widely used for collaborative pursuit. In each agent of D3PG, an Actor network is equipped to generate actions, and a Critic network is used to evaluate the value of the current state - action pair. However, traditional D3PG relies on frequent communication exchanges to synchronize the state information of each agent, increasing network load and potentially leading to increased latency, which limits the real - time response ability and scalability of the system. Since each agent makes decisions only based on its own local observations, it is difficult to adapt to scenarios that have not been encountered before, weakening the robustness and flexibility of the system. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for collaborative pursuit of unmanned vehicles - UAV clusters in an urban alley environment to solve the problems in the existing technology in view of the above - mentioned deficiencies of the prior art.
[0006] The present invention specifically provides the following technical solutions:
[0007] A method for collaborative hunting of unmanned vehicle - unmanned aerial vehicle clusters in an urban alley environment, comprising:
[0008] Construct a target network with the same initial weights as the main network, and use unmanned vehicles and unmanned aerial vehicles as agents. Set a target network in each agent, and set an experience replay buffer in the main network; both the main network and the target network include a decision-making Actor network and an evaluation Critic network;
[0009] At each time step t, obtain the current state of each agent, and according to the current state of each agent, use the decision-making Actor network in the target network to generate actions, and execute the actions to obtain a new state, a reward, and a flag indicating whether the task is completed;
[0010] Collect information of all agents including the current state, actions, rewards, new states, and flags indicating whether the task is completed, and construct a trajectory through the information;
[0011] Add the trajectory to the experience replay buffer. When the target network reaches a set interval or a predetermined condition, randomly extract a batch of samples from the experience replay buffer, obtain the maximum expected return of the next state of each sample, obtain a loss with the maximum expected return and backpropagate to update the evaluation Critic network, and guide the decision-making Actor network to learn until the target network model parameters are updated when the termination condition is met;
[0012] Obtain the current state of each agent through the target network model with updated parameters, pass the current state of each agent to other agents through the main network, and obtain a trajectory after processing the current state through the decision-making Actor network and the evaluation Critic network in the main network, for cross-domain cluster collaborative hunting of unmanned vehicles - unmanned aerial vehicles in an urban alley environment.
[0013] Preferably, the obtaining the maximum expected return of the next state of each sample, obtaining a loss with the maximum expected return and backpropagating to update the evaluation Critic network, and guiding the decision-making Actor network to learn until the target network model parameters are updated when the termination condition is met is specifically:
[0014] Calculate the maximum expected return in the next state, and the specific expression is:
[0015] y j =r j +γQ'(s j+1 ,μ'(s j+1 |θ' μ )|θ' Q );
[0016] Among them, y j is the target value, r jFor immediate reward, γ is the discount factor, Q′ is the Q-function for evaluating the Critic network, and s j+1 is the state at the next time step, μ′ is the policy function of the decision-making Actor network, and θ ′μ and θ′Q are the parameters of the decision-making Actor network and the evaluation Critic network respectively;
[0017] The mean squared error loss function is used to backpropagate and update the evaluation Critic network. The specific expression is:
[0018] L = (y j - Q(s j , a j |θ Q )) 2 ;
[0019] where L is the mean squared error loss function, Q is the Q-function of the current evaluation Critic network, s j is the state at the current time step, a j is the action taken in state s j and θ Q is the parameter of the current Critic network;
[0020] The learning of the decision-making Actor network is guided by the gradient provided by the evaluation Critic network. The specific expression is:
[0021] J(θ μ ) = E[Q(s j , μ(s j |θ μ )|θ Q )];
[0022] where J(θ μ ) is the objective function of the decision-making Actor network, E[] is the expected value, μ is the policy function of the current decision-making Actor network, and θ μ is the parameter of the current decision-making Actor network;
[0023] The soft update method is used to smoothly transition the parameters of the target network. The specific expression is:
[0024] θ' k = τ·θ k + (1 - τ)·θ' k ;
[0025] where k ∈ {μ, Q}, θ ′k is the parameter of the updated target network, θ k is the parameter of the current target network, and τ is the soft update coefficient.
[0026] Preferably, the specific expression for generating an action is: a t,i = μ(s t,i | θ μ ), where a t,i is the output action, μ is the policy function, s t,i is the input state, and θ μ is the parameter of the policy function μ.
[0027] Preferably, the structure of the decision-making Actor network includes a first fully connected layer fc1, a second fully connected layer fc2, and a final output layer fc3. The first fully connected layer fc1 takes the state dimension as the input and outputs 128 nodes. The second fully connected layer fc2 maintains 128 nodes, and the final output layer fc3 maps the 128 nodes to the action space dimension.
[0028] Preferably, the decision-making Actor network uses the ReLU activation function to activate the first fully connected layer fc1 and the second fully connected layer fc2, and uses the Tanh function to ensure that the output range of the final output layer fc3 is between [-1, 1].
[0029] Preferably, the structure of the evaluation Critic network includes a first fully connected layer fc1, a second fully connected layer fc2, and a final output layer fc3. The first fully connected layer fc1 takes the vector obtained by concatenating the state and the action as the input and outputs 128 nodes. The second fully connected layer fc2 maintains 128 nodes, and the final output layer fc3 outputs a single scalar value, representing the long-term cumulative reward Q value of the current state-action pair.
[0030] Preferably, the evaluation Critic network uses ReLU to activate the first fully connected layer fc1 and the second fully connected layer fc2, and the final output layer directly gives the prediction of the long-term cumulative reward Q value.
[0031] Preferably, the state includes the self-state and the environmental state. The self-state includes the real-time position, speed, and heading angle parameters of the agent, and the environmental state includes the current position and motion trajectory of the target. The action representation generates discrete or continuous actions based on the environmental state. The reward representation is the feedback when the target is surrounded.
[0032] The present invention provides an unmanned vehicle - UAV cluster collaborative surrounding system in an urban alley environment, including:
[0033] A construction module for constructing a target network with the same initial weights as the main network, using unmanned vehicles and UAVs as agents, setting a target network in each agent, and setting an experience replay buffer in the main network. Both the main network and the target network include a decision-making Actor network and an evaluation Critic network;
[0034] An action module, which is used to obtain the current state of each agent within each time step t, and based on the current state of each agent, generate an action using the decision-making Actor network in the target network and execute the action to obtain a new state, a reward, and a task completion flag;
[0035] A trajectory acquisition module, which is used to collect information of all agents including the current state, action, reward, new state, and task completion flag, and construct a trajectory through the information;
[0036] An update module, which is used to add the trajectory to the experience replay buffer. The target network randomly samples a batch of samples from the experience replay buffer at a set interval or when a predetermined condition is met, obtains the maximum expected return of the next state of each sample, obtains a loss based on the maximum expected return and performs backpropagation to update the evaluation Critic network, and guides the decision-making Actor network to learn until the target network model parameters are updated when the termination condition is met;
[0037] A cooperative hunting module, which is used to obtain the current state of each agent through the target network model with updated parameters, transfer the current state of each agent to other agents through the main network, and obtain a trajectory after processing the current state through the decision-making Actor network and the evaluation Critic network in the main network, so as to perform cross-domain cluster cooperative hunting of unmanned vehicles and drones in the urban alley environment.
[0038] The present invention provides a computer device, including a memory and a processor. A program is stored in the memory. When the program is executed by the processor, the processor executes the steps of the above-mentioned method for cross-domain cluster cooperative hunting of unmanned vehicles and drones in the urban alley environment.
[0039] Compared with the prior art, the present invention has the following remarkable advantages:
[0040] The present invention constructs a target network with the same initial weight as the main network, takes unmanned vehicles and drones as agents, obtains the state of each agent, and each agent transfers its own state to other agents, so that each agent can dynamically focus on other agents crucial to its decision-making, realizes accurate information sharing, reduces resource waste and increased latency caused by ineffective communication, enhances the adaptability of agents to unexperienced scenarios through the interaction between agents, can improve the generalization performance of the system. At the same time, the target network starts a state update process at a certain interval or when a predetermined condition is met until the target network model parameters are updated when the termination condition is met. Based on this, a communication strategy based on event triggering is proposed, and state updates are only performed at necessary moments, further optimizing the communication process. In this way, unnecessary information redundancy is reduced and the data transmission efficiency is improved. Description of the Drawings
[0041] Figure 1 It is the interaction logic diagram between components in the present invention;
[0042] Figure 2 It is the analysis diagram of the experimental results of the present invention;
[0043] Figure 3 It is the working principle diagram of the present invention;
[0044] Figure 4 It is the flowchart of a method for collaborative capture of an unmanned vehicle - unmanned aerial vehicle cluster in an urban alley environment according to the present invention. Specific implementation manners
[0045] Next, in combination with the accompanying drawings in the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] In the field of multi - agent systems (MAS), although existing methods based on deep reinforcement learning (DRL), such as the Distributed Deep Deterministic Policy Gradient (D3PG), have achieved certain success, some limitations and challenges have emerged in the actual application process. These problems not only affect the overall performance of the system but also limit its promotion and use in a wider range of scenarios. Therefore, the present invention aims to propose improvement measures for these defects and is committed to solving the following key technical problems:
[0047] 1. Improve communication efficiency, reduce latency and network load.
[0048] Problem description: Traditional D3PG relies on frequent communication exchanges to synchronize the state information of each agent, which not only increases the network load but also may lead to an increase in latency, especially in large - scale or rapidly changing dynamic environments. This high - frequency data transmission requirement limits the real - time response ability and scalability of the system.
[0049] Solution: ES-HADDPG introduces a multi-head attention mechanism, which enables each agent to dynamically pay attention to other members that are critical to its decision-making, achieve accurate information sharing, and reduce resource waste caused by invalid communication. In addition, an event-triggered communication strategy is proposed to update the status only when necessary, further optimizing the communication process. In this way, unnecessary information redundancy is reduced and data transmission efficiency is improved.
[0050] 2. Enhance the generalization ability of the model and increase the speed of adapting to new environments.
[0051] Problem description: Since existing methods mainly rely on local observations to make decisions, they may not adapt well to new situations when faced with scenarios they have never encountered, resulting in poor overall performance. This is because the existing model structure is difficult to capture the potential correlations and interaction patterns between different agents, which weakens the robustness and flexibility of the system.
[0052] Solution: ES-HADDPG uses the idea of meta-learning to build a basic model that can quickly adapt to new environments. The model can quickly adjust parameters under a small number of samples to meet the task requirements under unknown conditions. At the same time, by sharing the weights of some neural network layers, it promotes knowledge transfer between different agents and enhances the versatility and adaptability of the system. Figure 2 As shown in the figure, the experimental results show that the pre-trained ES-HADDPG framework can converge to high-quality strategies faster in the new test environment and show stronger generalization ability.
[0053] 3. Accelerate the convergence process and ensure stable learning results.
[0054] Problem description: When dealing with complex nonlinear relationships, the standard D3PG is prone to falling into local optimal solutions and it is difficult to ensure the acquisition of the global optimal solution. In addition, due to the lack of effective design of the exploration mechanism, the algorithm may converge to a suboptimal strategy prematurely, affecting the final learning effect.
[0055] Solution: To overcome this problem, ES-HADDPG adopts an adaptive step size adjustment strategy (Adaptive Step Size Adjustment) to dynamically adjust the learning rate according to the difficulty of the task, avoiding overfitting caused by fixed step size settings.
[0056] Meanwhile, combined with Prioritized Experience Replay, more valuable experience samples are preferentially selected for learning, effectively accelerating the convergence speed. In addition, a Soft Updates mechanism and various innovative exploration methods (such as noisy networks, random search, etc.) are introduced to gradually replace the old target network parameters, maintaining a smooth transition during the policy iteration process, improving the training stability, and preventing premature convergence.
[0057] 4. Optimize the experience replay buffer to improve learning efficiency.
[0058] Problem description: The traditional Experience Replay Buffer fails to fully consider the impact of time factors, which may lead to the neglect of some important but relatively new experiences, thus affecting the learning efficiency.
[0059] Solution: ES-HADDPG ensures that different types of experiences can be fully learned by introducing a time decay factor and a diversity-based sampling strategy. At the same time, combined with PER, valuable experience samples are preferentially selected for learning, further improving the learning efficiency and algorithm performance.
[0060] 5. Adjust the reward function to promote a more effective learning process.
[0061] Problem description: Sparse reward signals make it difficult for the agent to obtain sufficient feedback to guide its behavior, thus prolonging the learning cycle.
[0062] Solution: ES-HADDPG designs a reasonable auxiliary reward function to guide the agent to learn the expected behavior faster. For example, a certain reward is given for successfully approaching the target but not being captured, and the collision penalty is reduced. This method is particularly helpful in alleviating the challenges brought by sparse rewards and promoting a more effective learning process.
[0063] 6. Reduce the consumption of computing resources and lower the cost threshold.
[0064] Problem description: In order to maintain high precision, existing technologies often require powerful computing devices to support the training of large-scale neural networks, which places high demands on hardware facilities and also brings problems such as energy waste.
[0065] Solution: By optimizing and simplifying the network architecture, the computational load and memory occupancy are significantly reduced without sacrificing performance. Specifically, ES-HADDPG adopts a lightweight convolutional neural network (CNN) as the feature extractor and minimizes the number of fully connected layers to reduce the model complexity. In addition, model compression techniques such as pruning and quantization are implemented to further reduce the storage space requirements. These improvement measures enable the ES-HADDPG framework to run on resource-constrained embedded devices, reducing the deployment cost and technical threshold, and facilitating its popularization and application in more practical scenarios.
[0066] In summary, the ES-HADDPG framework proposed in the present invention provides effective solutions to multiple problems existing in the prior art. It is particularly outstanding in improving communication efficiency, enhancing generalization ability, accelerating convergence stability, and reducing resource consumption, providing new ideas and technical means for solving the collaborative decision-making problem of multi-agent systems. Through these innovative improvements, not only the collaborative decision-making level of multi-agent systems is improved, but also new ideas and technical means are provided for future research and development, bringing considerable economic benefits and social value at the same time.
[0067] As Figure 1 shown, ES-HADDPG is a deep reinforcement learning algorithm designed for multi-agent systems. It inherits the advantages of DDPG and introduces a series of improvement measures to solve the problems existing in traditional methods. The framework aims to improve communication efficiency, enhance generalization ability, accelerate convergence stability, and reduce resource consumption to achieve a more efficient and robust collaborative decision-making process.
[0068] Interaction logic between components:
[0069] The entire ES-HADDPG system consists of multiple agents with the same configuration, and each agent has its own Actor and Critic networks. These agents cooperate by sharing environmental information:
[0070] State transfer: Each agent observes the local environmental state it is in and transfers it to other agents. Action selection: Based on the current state, each agent uses its internal Actor network to calculate the best action and execute it. Reward feedback: After executing the action, the environment returns an immediate reward to the corresponding agent, which is used as the basis for subsequent learning. Value evaluation: All agents jointly maintain a global experience replay buffer to record the results of each interaction. When an update is needed, a batch of samples is randomly selected for the Critic network to evaluate the value and adjust the behavior of the Actor accordingly.
[0071] Table 1 ES-HADDPG Algorithm
[0072]
[0073] As Figure 4 shown, based on this, this embodiment provides a collaborative surrounding and capturing method for unmanned vehicle - UAV clusters in urban alley environments, including:
[0074] Step S1: Construct a target network with the same initial weights as the main network, and regard the unmanned vehicle and the UAV as agents. Set a target network in each agent, and set an experience replay buffer in the main network. Both the main network and the target network include a decision-making Actor network and an evaluation Critic network.
[0075] Initialization: Define hyperparameters, construct Actor and Critic network instances and initialize the corresponding Adam optimizers, create a copy of the target network with the same initial weights as the main network, and prepare the experience replay buffer.
[0076] Definition of Actor network: A neural network responsible for generating action policies, which receives the state as input and outputs continuous action values. The structure of the decision-making Actor network includes a first fully connected layer fc1, a second fully connected layer fc2, and a final output layer fc3. The first fully connected layer fc1 takes the received state dimension as input and outputs 128 nodes. The second fully connected layer fc2 maintains 128 nodes, and the final output layer fc3 maps 128 nodes to the action space dimension.
[0077] Activation function: Use ReLU to activate the first two layers (the first fully connected layer fc1 and the second fully connected layer fc2), and use the Tanh function for the last layer to ensure that the output range is between [-1, 1], which is suitable for the standardized action space.
[0078] Definition of Critic network: A value function used to evaluate the value of taking a specific action in a given state, which helps to optimize the behavior of the Actor. The structure of the evaluation Critic network includes a first fully connected layer fc1, a second fully connected layer fc2, and a final output layer fc3. Among them, the first fully connected layer fc1 takes the vector obtained by concatenating the state and the action as input and outputs 128 nodes. The second fully connected layer fc2 maintains 128 nodes, and the final output layer fc3 outputs a single scalar value, representing the long-term cumulative reward Q value of the current state-action pair.
[0079] Activation function: Also use ReLU to activate the hidden layers (the first fully connected layer fc1 and the second fully connected layer fc2), and the output layer directly gives the Q value prediction.
[0080] The target network is for stabilizing the training process. Each agent is equipped with a set of target networks (actor_target and critic_target), which periodically copy parameters from the main network but with a lower update frequency, thus avoiding the instability caused by frequent changes.
[0081] Hyperparameters, including but not limited to learning rate (lr), discount factor (γ), soft update coefficient (τ), etc.
[0082] Step S2: Environment interaction: At each time step t, obtain the current state of each agent, and based on the current state of each agent, use the decision-making Actor network in the target network to generate actions and execute the actions to obtain a new state, reward, and a flag indicating whether the task is completed.
[0083] The state includes the agent's own state and the environment state. The agent's own state includes the real-time position, speed, and heading angle parameters of the agent, and the environment state includes the current position and movement trajectory of the target; the action representation generates discrete or continuous actions based on the environment state; the reward representation is the feedback when capturing the target.
[0084] Step S3: Collect the information of all agents including the current state, action, reward, new state, and the flag indicating whether the task is completed, and construct a trajectory with the information.
[0085] Step S4: Add the trajectory to the experience replay buffer. The target network randomly samples a batch of samples from the experience replay buffer at a set interval or when a predetermined condition is met, obtains the maximum expected return of the next state for each sample, obtains the loss with the maximum expected return and backpropagates to update the evaluation Critic network, guiding the decision-making Actor network to learn, and updates the parameters of the target network model until the termination condition is met.
[0086] Collect the information of all agent Agents to form a complete trajectory <s t ,a t ,r t ,s t+1 ,done>, add this trajectory to the experience replay buffer; Update model parameters: Start an update process at a certain interval or when a predetermined condition is met, randomly sample a batch of samples <s j ,a j ,r j ,s j+1 ,done j > from the buffer, calculate the maximum expected return in the next state for each sample, and update the model parameters.
[0087] For each sample j: Calculate the maximum expected return in the next state, and the specific expression is:
[0088] y j = r j + γQ'(s j+1 , μ'(s j+1 | θ' μ )| θ' Q );
[0089] where θ' μ and θ' Q are the parameters of the target network, y j is the target value, r j is the immediate reward, γ is the discount factor, Q′ is the Q function of the target Critic network, s j+1 is the state at the next time step, μ′ is the policy function of the target Actor network, θ ′μ and θ ′Q are the parameters of the target Actor network and the target Critic network respectively.
[0090] Use the mean squared error loss function to backpropagate and update the Critic network. The specific expression is:
[0091] L = (y j - Q(s j , a j | θ Q )) 2 ;
[0092] where L is the mean squared error loss function, Q is the Q function of the current Critic network, s j is the state at the current time step, a j is the action taken in state s j , and θ Q is the parameter of the current Critic network.
[0093] Use the gradient provided by the Critic to guide the learning of the Actor network. The specific expression is:
[0094] J(θ μ ) = E[Q(s j , μ(s j | θ μ )| θ Q )];
[0095] where J(θ μ ) is the objective function of the Actor network, E[] is the expected value, μ is the policy function of the current Actor network, and θ μ is the parameter of the current Actor network.
[0096] Perform a soft update to smoothly transition the parameters of the target network. The specific expression is:
[0097] θ' k = τ·θ k +(1 - τ)·θ' k ;
[0098] where k ∈ {μ, Q}, θ ′k is the updated target network parameter, θ k is the parameter of the current network, and τ is the soft update coefficient.
[0099] Check whether the termination condition is met (e.g., reaching the maximum number of iterations or the performance metric is met). If so, end the training; otherwise, return to step 2 to continue the loop.
[0100] Step S5: Obtain the current state of each agent through the target network model with updated parameters, pass the current state of each agent to other agents through the main network, and obtain the trajectory after processing the current state through the decision-making Actor network and the evaluation Critic network in the main network, and perform the cross-domain cluster cooperation and encirclement and capture of the unmanned vehicle - UAV in the urban lane environment.
[0101] When obtaining the trajectory by processing the current state through the decision-making Actor network and the evaluation Critic network in the main network, the method is the same as that of obtaining the trajectory by the target network.
[0102] To verify the effectiveness of ES-HADDPG, the present invention conducted a number of comparative experiments. The experimental platform was built in a simulated environment, including several movable robot agents, and the task was to cooperate to carry items to the specified location. ES-HADDPG was compared with other benchmark algorithms such as D3PG, and the performance of each group was recorded.
[0103] The performance evaluation metrics include: Average Cumulative Reward: Measuring the total return obtained by the Agent during the entire task, reflecting the quality of the policy. Convergence Speed: Observing the number of training rounds required for different algorithms to reach stable performance, the faster the better. Resource Utilization: Statistically analyzing the usage of hardware resources such as CPU / GPU occupancy rate and memory consumption, reflecting the lightweight characteristics of the algorithm. Communication Overhead: Calculating the amount of data transmitted between agents, evaluating the real-time response ability and scalability of the system.
[0104] After multiple repeated experiments, ES-HADDPG demonstrated obvious advantages over the existing technologies:
[0105] Within the same training time, ES-HADDPG can find a better solution faster, resulting in a significantly higher average cumulative reward than the control group. Due to the adoption of an efficient communication mechanism and a lightweight network architecture, ES-HADDPG not only reduces communication latency but also decreases the demand for computing resources, achieving a better energy efficiency ratio. More importantly, with the help of meta-learning and knowledge transfer techniques, ES-HADDPG demonstrates stronger generalization ability and can quickly adapt and achieve good results when encountering new environments.
[0106] The structure and working principle of the ES-HADDPG framework of the present invention:
[0107] Structure: Actor network: A neural network responsible for generating action policies, which receives the state as input and outputs continuous action values. Critic network: A value function used to evaluate the value of taking a specific action in a given state, helping to optimize the behavior of the Actor. Target network: To stabilize the training process, each Agent is equipped with a set of target networks (actor_target and critic_target), which regularly copy parameters from the main network but with a lower update frequency, thus avoiding the instability caused by frequent changes.
[0108] Attention mechanism module: Introduce a multi-head attention model to enable the agent to dynamically focus on other members crucial for its decision-making and achieve precise information sharing.
[0109] Experience replay buffer: Adopt a time decay factor and a diversity-based sampling strategy, combined with prioritized experience replay (PER), to ensure that different types of experiences can be fully learned.
[0110] Exploration strategy module: Integrate a variety of innovative exploration methods, such as noisy networks, random search, etc., to help the algorithm jump out of local optima and find better global solutions.
[0111] Reward adjustment module: Design a reasonable auxiliary reward function to guide the agent to learn the expected behavior faster and alleviate the challenges brought by sparse rewards.
[0112] Adaptive parameter adjustment module: Develop an adaptive learning rate adjustment rule and other hyperparameter adjustment schemes, enabling the algorithm to automatically adjust its own behavior pattern at different stages to cope with the changing task requirements.
[0113] Working principle: Efficient communication mechanism: By introducing the multi-head attention mechanism, unnecessary information redundancy is reduced, and data transmission efficiency is improved; combined with an event-triggered communication strategy, state updates are performed only when necessary, further optimizing the communication process. Enhanced generalization ability: Using the idea of meta-learning, a basic model that can quickly adapt to new environments is constructed. By sharing the weights of some neural network layers, knowledge transfer between different agents is promoted, enhancing the generality and adaptability of the system. Accelerated convergence and ensured stability: An adaptive step-size adjustment strategy is adopted to dynamically adjust the learning rate according to the task difficulty; combined with the prioritized experience replay and soft update mechanisms, a smooth transition during the policy iteration process is ensured, improving training stability; at the same time, multiple exploration strategies prevent premature convergence. Optimized experience replay buffer: Through the time decay factor and diversity sampling strategy, it is ensured that different types of experiences can be fully learned, improving learning efficiency. Adjusted reward function: A reasonable auxiliary reward function is designed to guide the agent to learn the expected behavior faster, promoting a more effective learning process. Reduced resource costs: By optimizing and simplifying the network architecture, the computational amount and memory occupation are significantly reduced without affecting performance; model compression technologies such as pruning and quantization are implemented, reducing the storage space requirement, enabling the ES-HADDPG framework to run on resource-constrained embedded devices. As Figure 3 shown, events are generated by an event trigger. After being encoded by the input encoder and decoded by the decoder, they are input to agent A. The filtered information is input to agent A through information filtering. Agent A and agent B interact in the external environment.
[0114] Beneficial effects that can be achieved:
[0115] 1. Improve communication efficiency, reduce latency and network load.
[0116] Beneficial effect: By designing an efficient encoding and decoding mechanism and an event-triggered communication strategy, unnecessary information exchange is significantly reduced, enhancing the real-time response ability and scalability of the system.
[0117] Economic benefits: The bandwidth requirement is reduced, communication costs are saved, especially in large-scale or rapidly changing application environments, which helps enterprises reduce operating costs.
[0118] 2. Enhance the model's generalization ability and improve the speed of adapting to new environments.
[0119] Beneficial effect: The pre-trained ES-HADDPG framework can converge to high-quality policies faster in new test environments, showing stronger generalization ability, reducing the time and resource consumption for retraining.
[0120] Economic benefits: For industries that need to frequently deal with unknown situations, such as autonomous driving and robot collaboration, this will greatly improve production efficiency and service quality, bringing significant competitive advantages.
[0121] 3. Accelerate the convergence process and ensure stable learning results.
[0122] Beneficial effects: The adaptive step size adjustment strategy and the soft update mechanism work together to ensure a smooth transition during the policy iteration process and improve the training stability; at the same time, priority experience replay and multiple exploration strategies enable the algorithm to achieve better performance on limited data sets.
[0123] Economic benefits: It shortens the R&D cycle, reduces the cost of trial and error, and enables new products and technologies to be brought to market faster, thus gaining an advantage for the company.
[0124] 4. Reduce computing resource consumption and lower cost threshold.
[0125] Beneficial effects: The application of lightweight network architecture and model compression technology enables the ES-HADDPG framework to run on resource-constrained embedded devices, expanding its scope of application.
[0126] Economic benefits: It reduces the cost of hardware procurement and the difficulty of technical implementation, which is conducive to small and medium-sized enterprises and individual developers to participate in the research and development of cutting-edge technologies and promotes the development of a technological innovation ecosystem.
[0127] 5. Improve exploration strategies and enhance exploration efficiency.
[0128] Beneficial effects: By introducing a variety of innovative exploration methods, such as noise networks and random searches, the algorithm is helped to escape from the local optimal solution and find a better global solution, thus improving the exploration efficiency.
[0129] Economic benefits: In complex task environments, more effective exploration strategies can speed up the process of finding the best solution, thereby improving work efficiency and service quality.
[0130] 6. Dynamically adjust the reward function to promote effective learning.
[0131] Beneficial effects: Designing a reasonable auxiliary reward function can guide the agent to learn the expected behavior faster, alleviate the challenges brought by sparse rewards, and promote a more effective learning process.
[0132] Economic benefits: Through better reward design, the learning process can be accelerated, the time and resources required to achieve ideal performance can be reduced, and costs can be saved for the enterprise.
[0133] Based on the above method, the present invention provides a collaborative hunting system for unmanned vehicle - unmanned aerial vehicle clusters in an urban alley environment, including: a construction module, an action module, a trajectory acquisition module, an update module, and a collaborative hunting module.
[0134] Among them, the construction module is used to construct a target network with the same initial weight as the main network, and regard the unmanned vehicle and the unmanned aerial vehicle as agents. Set a target network in each agent, and set an experience replay buffer in the main network; both the main network and the target network include a decision-making Actor network and an evaluation Critic network; the action module is used to obtain the current state of each agent at each time step t, and according to the current state of each agent, use the decision-making Actor network in the target network to generate actions and execute the actions to obtain a new state, a reward, and a task completion flag; the trajectory acquisition module is used to collect information of all agents including the current state, actions, rewards, new states, and task completion flags, and construct a trajectory through the information; the update module is used to add the trajectory to the experience replay buffer. The target network randomly extracts a batch of samples from the experience replay buffer at a set interval or when a predetermined condition is reached, obtains the maximum expected return of the next state of each sample, obtains the loss with the maximum expected return and backpropagates to update the evaluation Critic network, and guides the decision-making Actor network to learn until the termination condition is met and then updates the target network model parameters; the collaborative hunting module is used to obtain the current state of each agent through the target network model after parameter update, transmit the current state of each agent to other agents through the main network, and obtain a trajectory after processing the current state through the decision-making Actor network and the evaluation Critic network in the main network, so as to perform cross-domain cluster collaborative hunting of unmanned vehicles - unmanned aerial vehicles in an urban alley environment.
[0135] The present invention also provides a computer device, including a memory and a processor. A program is stored in the memory, and when the program is executed by the processor, the processor is caused to execute the steps of a method for collaborative hunting of unmanned vehicles - unmanned aerial vehicle clusters in an urban alley environment.
[0136] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.
[0137] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for collaborative hunting of unmanned vehicles - unmanned aerial vehicle clusters in an urban alley environment are implemented.
[0138] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0139] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the technical field of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A collaborative pursuit method for unmanned vehicle - unmanned aerial vehicle clusters in an urban alley environment, characterized in that, Including: Construct a target network with the same initial weights as the main network, and use the unmanned vehicle and the unmanned aerial vehicle as agents. Set a target network in each agent, and set an experience replay buffer in the main network; both the main network and the target network include a decision-making Actor network and an evaluation Critic network; At each time step t, obtain the current state of each agent, and according to the current state of each agent, use the decision-making Actor network in the target network to generate an action, and execute the action to obtain a new state, a reward, and a flag indicating whether the task is completed; Collect information of all agents including the current state, action, reward, new state, and flag indicating whether the task is completed, and construct a trajectory through the information; Add the trajectory to the experience replay buffer. The target network randomly extracts a batch of samples from the experience replay buffer at regular intervals or when a predetermined condition is reached, obtains the maximum expected return of the next state of each sample, obtains a loss based on the maximum expected return, and backpropagates to update the evaluation Critic network, guiding the decision-making Actor network to learn until the termination condition is met and then updating the target network model parameters; Obtain the current state of each agent through the target network model with updated parameters, transfer the current state of each agent to other agents through the main network, and obtain a trajectory after processing the current state through the decision-making Actor network and the evaluation Critic network in the main network for cross-domain cluster collaborative pursuit of unmanned vehicles and unmanned aerial vehicles in the urban alley environment.
2. The method for collaborative surrounding and capturing of an unmanned vehicle - unmanned aerial vehicle cluster in an urban alley environment according to claim 1, wherein, The obtaining of the maximum expected return of the next state of each sample, obtaining a loss based on the maximum expected return, and backpropagating to update the evaluation Critic network, guiding the decision-making Actor network to learn until the termination condition is met and then updating the target network model parameters is specifically as follows: Calculate the maximum expected return in the next state, and the specific expression is: y j = r j + γQ'(s j+1 , μ'(s j+1 |θ' μ )|θ' Q ); Among them, y j is the target value, r j is the immediate reward, γ is the discount factor, Q′ is the Q function for evaluating the Critic network, s j+1 is the state at the next time step, μ′ is the policy function of the decision-making Actor network, θ ′μ and θ′ Q are the parameters of the decision-making Actor network and the evaluation Critic network respectively; Use the mean squared error loss function to backpropagate and update the evaluation Critic network, and the specific expression is: L = (y j - Q(s j , a j |θ Q )) 2 ; where \(L\) is the mean squared error loss function, \(Q\) is the \(Q\)-function of the current evaluating Critic network, \(s\) j is the state at the current time step, \(a\) j is the action taken at state \(s\) j and \(\theta\) Q are the parameters of the current Critic network; Use the gradient provided by the evaluation Critic network to guide the learning of the decision-making Actor network, and the specific expression is: J(θ μ ) = E[Q(s j , μ(s j | θ μ )| θ Q )]; Among them, J(θ μ ) is the objective function of the decision-making Actor network, E[] is the expected value, μ is the policy function of the current decision-making Actor network, and θ μ are the parameters of the current decision-making Actor network; Use the soft update method to smoothly transition the parameters of the target network, and the specific expression is: θ' k = τ·θ k + (1 - τ)·θ' k ; where \(k\in\{\mu,Q\}\), \(\theta\) ′k is the updated target network parameter, \(\theta\) k is the parameter of the current target network, and \(\tau\) is the soft update coefficient.
3. The method for collaborative pursuit by an unmanned vehicle - unmanned aerial vehicle cluster in an urban alley environment according to claim 1, characterized in that, The specific expression for generating an action is: a t,i = μ(s t,i | θ μ ), where a t,i is the output action, μ is the policy function, s t,i is the input state, and θ μ is the parameter of the policy function μ.
4. The method for collaborative encirclement and capture of an unmanned vehicle - unmanned aerial vehicle cluster in an urban alley environment according to claim 1, wherein, The structure of the decision-making Actor network includes a first fully connected layer fc1, a second fully connected layer fc2, and a final output layer fc3. The first fully connected layer fc1 takes the state dimension as the input and outputs 128 nodes. The second fully connected layer fc2 maintains 128 nodes, and the final output layer fc3 maps the 128 nodes to the action space dimension.
5. A method for collaborative encirclement and capture of unmanned vehicle - unmanned aerial vehicle clusters in an urban alley environment according to claim 4, characterized in that, The decision-making Actor network uses the ReLU activation function to activate the first fully connected layer fc1 and the second fully connected layer fc2, and uses the Tanh function to ensure that the output range of the final output layer fc3 is between [-1, 1].
6. The method for collaborative encirclement and capture of an unmanned vehicle - unmanned aerial vehicle cluster in an urban alley environment according to claim 1, wherein, The structure of the evaluation Critic network includes a first fully connected layer fc1, a second fully connected layer fc2, and a final output layer fc3. Among them, the first fully connected layer fc1 receives the vector after concatenating the state and the action as input and outputs 128 nodes. The second fully connected layer fc2 maintains 128 nodes, and the final output layer fc3 outputs a single scalar value, representing the long-term cumulative reward Q value of the current state-action pair.
7. A collaborative surrounding and capturing method for unmanned vehicle - UAV clusters in an urban alley environment as claimed in claim 6, characterized in that, The evaluation Critic network uses ReLU to activate the first fully connected layer fc1 and the second fully connected layer fc2, and the final output layer directly gives the prediction of the long-term cumulative reward Q value.
8. The method for collaborative encirclement and capture of an unmanned vehicle - unmanned aerial vehicle cluster in an urban alley environment according to claim 1, wherein The state includes the self-state and the environmental state. The self-state includes the real-time position, speed, and heading angle parameters of the agent, and the environmental state includes the current position and movement trajectory of the target. The action representation generates discrete or continuous actions based on the environmental state. The reward representation is the feedback when the target is surrounded.
9. An unmanned vehicle - UAV cluster cooperative encirclement and capture system in an urban alley environment, characterized in that, It includes: A construction module for constructing a target network with the same initial weights as the main network, taking the unmanned vehicle and the unmanned aerial vehicle as agents, setting a target network in each agent, and setting an experience replay buffer in the main network. Both the main network and the target network include a decision-making Actor network and an evaluation Critic network. An action module for obtaining the current state of each agent at each time step t, and according to the current state of each agent, using the decision-making Actor network in the target network to generate an action and execute the action to obtain a new state, a reward, and a flag indicating whether the task is completed. A trajectory acquisition module for collecting information of all agents including the current state, action, reward, new state, and flag indicating whether the task is completed, and constructing a trajectory through the information. An update module for adding the trajectory to the experience replay buffer. The target network randomly extracts a batch of samples from the experience replay buffer at regular intervals or when a predetermined condition is met, obtains the maximum expected return of the next state of each sample, obtains the loss with the maximum expected return and backpropagates to update the evaluation Critic network, guiding the decision-making Actor network to learn until the termination condition is met and then updating the target network model parameters. A collaborative surrounding module for obtaining the current state of each agent through the target network model after parameter update, transmitting the current state of each agent to other agents through the main network, and obtaining a trajectory after processing the current state through the decision-making Actor network and the evaluation Critic network in the main network to perform cross-domain cluster collaborative surrounding of unmanned vehicles and unmanned aerial vehicles in the urban lane environment.
10. A computer device, characterized in that, It includes a memory and a processor. When the program stored in the memory is executed by the processor, the processor executes the steps of a method for cross-domain cluster collaborative surrounding of unmanned vehicles and unmanned aerial vehicles in the urban lane environment according to any one of claims 1 to 8.
Citation Information
Cited By
Dynamic unmanned aerial vehicle cluster multi-radar cooperative detection tracking method and system based on target resolution
CN121899801A