Intelligent group collaborative pursuit method and system under incomplete information condition
Through SimHash technology and multi-agent deep reinforcement learning algorithm, the problem of collaborative pursuit of multi-agent systems under incomplete information conditions is solved, efficient exploration and target pursuit are achieved, and the coordination efficiency and robustness of the system in complex environments is improved.
Patent Information
- Application Number
- CN202510738071.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-02
AI Technical Summary
In an environment with incomplete information and perception limitations, it is difficult for multi-agent systems to efficiently detect the environment, avoid obstacles, and achieve target pursuit under sparse rewards. The existing technology has problems of low synergy efficiency and difficulty in forming strategies.
The state-action space has been discretized using SimHash technology, combined with the multi-agent deep reinforcement learning algorithm, and optimize the agent strategy through the intrinsic incentive and reward mechanism to achieve efficient exploration and collaborative pursuit.
It improves the exploration efficiency and target pursuit capabilities of the agent in complex environments, ensures the efficient coordination and stability of the system under sparse reward conditions, and adapts to a dynamically changing environment.
Smart Images

Figure CN120578050A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-agent control optimization, and specifically relates to an intelligent group collaborative pursuit method and system under incomplete information conditions. Background Art
[0002] In recent years, multi-agent reinforcement learning (MARL) has shown significant research potential in complex tasks such as collaborative pursuit, group decision-making, and dynamic games. However, in practical applications, agents often face incomplete information environments, and their perception capabilities are constrained by factors such as sensor type, accuracy, and detection range. This makes it difficult to obtain timely and accurate global state information about the target or environment, which in turn affects the overall collaborative efficiency and task completion quality of the multi-agent system. Furthermore, the reward signals in the real-world environments of collaborative pursuit tasks are often very sparse, resulting in low exploration efficiency and difficulty in quickly forming effective pursuit strategies due to the limited local experience accumulated by agents. Furthermore, the increasing complexity of the environment, such as obstacles, not only increases the difficulty of executing pursuit tasks and avoiding obstacles, but also places higher demands on information sharing and collaboration between agents. Currently, multi-agent collaborative pursuit, under conditions of limited perception and complex environments, still faces numerous technical bottlenecks and challenges that need to be overcome. How to efficiently explore the environment and avoid unknown obstacles, improve exploration efficiency under sparse rewards, and achieve stable pursuit of the target, all require solutions. These challenges seriously restrict the application promotion and performance improvement of multi-agent systems in real scenarios. Summary of the Invention
[0003] To overcome these existing issues, this paper proposes a method and system for collaborative pursuit under incomplete information conditions. By incorporating SimHash-based state novelty evaluation and local perception modeling, this method effectively enhances the agent's ability to explore the location environment and adapt to complex environments, thereby improving the efficiency and robustness of collaborative pursuit tasks.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] An intelligent group collaborative pursuit method under incomplete information conditions includes the following steps:
[0006] S101. Establishing an incomplete information hunting environment;
[0007] S102, obtaining local perception information of each agent and constructing a joint state space and action space;
[0008] S103. Use the SimHash function to hash the joint state-action pairs, assign a hash count table, and update the corresponding access count in real time to achieve low-dimensional discrimination and novelty statistics in high-dimensional space; the more novel the state-action combination, the greater the intrinsic incentive reward the agent receives;
[0009] S104: Update the hash table and intrinsic incentive reward in real time for each interaction between the agent and the environment, and save the historical data of the interaction;
[0010] S105, integrating intrinsic motivation and environmental rewards, optimizing collaborative pursuit strategies through a multi-agent deep reinforcement learning algorithm with centralized training and distributed execution;
[0011] S106, using the trained policy network to autonomously generate pursuit action instructions based on the current environment state, driving each agent to collaboratively perform the pursuit task;
[0012] The present invention also provides a multi-agent collaborative pursuit system under incomplete information conditions, comprising:
[0013] The initialization module is configured to build a multi-agent collaborative pursuit simulation environment, set environmental parameters and constraints, including the number of agents, initial positions, escapee behavior patterns, obstacle distribution, etc., and complete the initialization of the state hash function and counting table, including initializing the SimHash function parameter matrix A and hash code length k, and establishing an empty hash counting table C for subsequent state novelty evaluation; at the same time, it initializes the multi-agent deep reinforcement learning network structure, including defining the hierarchical structure, activation function and initial weights of the Actor network and Critic network.
[0014] The learning and training module is configured to implement a multi-agent reinforcement learning process based on intrinsic incentives, including: collecting local observation information of each agent and constructing a joint state-action feature vector; achieving state novelty discrimination through hash discretization, and combining intrinsic incentives with environmental rewards to optimize the agent strategy using deep reinforcement learning methods. Specifically, the feature vector is mapped to a hash code through the SimHash function and the counting table is updated; the intrinsic incentive reward is calculated based on the hash count and combined with the environmental reward; the parameters are updated using the Actor-Critic architecture, where the Actor network outputs an action based on the current observation, and the Critic network evaluates the value of this action; during the training process, the module continuously optimizes the agent strategy so that it can not only efficiently explore the environment, adapt to the obstacle distribution, but also collaboratively pursue the target object; when the training reaches the preset conditions (such as reward convergence or the maximum number of iterations), the trained policy network model is saved.
[0015] The task execution module is configured to deploy and execute strategies based on the trained strategy model during actual application or testing. Specifically, it receives real-time environmental information and observations from each agent, outputs control commands for the agents, and enables multi-agent collaboration and target capture. It inputs each agent's local observation data into the trained actor network model. The actor network directly generates the optimal control actions (driving force) for each agent based on its current state. Each agent executes the corresponding actions according to the commands, achieving coordinated capture of the target while effectively avoiding obstacles. This module no longer requires calculating intrinsic incentive rewards or updating network parameters during execution, relying solely on the trained strategy model for real-time decision-making, ensuring the system's efficiency and stability in practical applications.
[0016] Compared with the prior art, the present invention has the following beneficial effects:
[0017] This invention scientifically divides the multi-agent collaborative pursuit task into two key phases: an intrinsically motivated environmental exploration phase and an efficient target pursuit phase. In the exploration phase, SimHash technology effectively compresses and classifies the state-action space, enabling the agents to proactively explore unknown or uncommon areas. In the pursuit phase, an optimized policy network is utilized to enable precise capture of the target and dynamic obstacle avoidance by the multi-agent system.
[0018] This paper employs a carefully designed hash code counting and intrinsic incentive mechanism for environments with incomplete information and limited perception, effectively addressing the problem of inefficient agent exploration in sparse reward environments. By converting state novelty into a numerical intrinsic incentive reward, the agent is encouraged to actively explore unvisited or less-visited areas of the environment, accelerating the policy learning process and significantly improving training efficiency and convergence speed.
[0019] This invention innovatively combines SimHash hash discretization technology with a multi-agent deep reinforcement learning algorithm. It not only solves the technical difficulty of directly counting high-dimensional continuous state-action spaces, but also effectively ensures the stability and convergence of algorithm training by updating the hash counting table in real time and dynamically calculating intrinsic incentive rewards.
[0020] This invention utilizes a centralized training and distributed execution architecture to achieve efficient coordinated pursuit while maintaining the autonomous decision-making capabilities of each agent. Even in complex scenarios with dense obstacles and limited information, the trained strategy demonstrates robustness and adaptability, enabling it to cope with dynamically changing environments and target behaviors.
[0021] Overall, the collaborative pursuit method of the present invention can enable intelligent agents to efficiently complete collaborative pursuit tasks in a perception-limited environment. This method can be flexibly applied in complex environments and provide more efficient collaboration and decision-making support for related applications in various fields.
[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which constitute part of this application, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0024] In the picture:
[0025] Figure 1 Flowchart of the multi-agent collaborative pursuit method based on intrinsic incentive reward mechanism.
[0026] Figure 2 Schematic diagram of the multi-agent collaborative pursuit simulation environment.
[0027] Figure 3 This is a schematic diagram of the working principle of the SimHash function.
[0028] Figure 4 Schematic diagram of the Actor-Critic network structure.
[0029] Figure 5 This is the reward convergence curve for multi-agent collaborative pursuit.
[0030] Figure 6 A sequence diagram demonstrating a multi-agent collaborative pursuit case. DETAILED DESCRIPTION
[0031] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate for the embodiments of the present invention described herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.
[0034] Figure 1 This is a flowchart of the multi-agent collaborative pursuit method based on the intrinsic incentive mechanism. It presents in detail the complete process from environment perception, state hash discretization, intrinsic incentive calculation to action generation, such as Figure 1 As shown, the method includes the following steps:
[0035] S101. Establish a hunting environment with incomplete information conditions.
[0036] Specifically, a kinematic model suitable for multi-agent coordinated pursuit is established, namely a simulation environment containing multiple homogeneous pursuers, a single escapee, and randomly distributed obstacles. Each agent is given a limited perception radius, allowing it to only obtain partial environmental information, creating a pursuit scenario under incomplete information conditions.
[0037] The state of each agent i is represented by its position coordinates (x i ,y i ) and velocity component (v x,i ,v y,i ) representation. The motion of the intelligent agent follows Newton’s second law, and its dynamic equation is:
[0038]
[0039] Among them, (F x,i ,F y,i ) is the control input (driving force) of agent i in the x and y directions. In this invention, the x direction is defined as the horizontal direction, the y direction is defined as the vertical direction, and m is the mass of the agent. are the speed and acceleration of agent i in the x and y directions. In the established simulation environment, the agent’s perception ability is limited and it can only obtain environmental information within a certain radius, resulting in a collaborative pursuit scenario under incomplete information conditions.
[0040] S102. Obtain the local perception information of each agent and construct a joint state space and action space.
[0041] Specifically, construct an individual observation space O for each hunting agent i i , which contains the agent's own state information (position coordinates and speed), the relative state information of other agents within the perception range (relative position and relative speed), the relative state information of escapees within the perception range (relative position and relative speed), and the relative position information of obstacles within the perception range. It is used to describe the basis for collaboration and decision-making among multiple agents; among them, the state is represented by position and speed.
[0042] The observation vector O of a single agent i It can be expressed as:
[0043] o i =[x i ,y i ,v x,i ,v y,i ,Δx i,j ,Δy i,j ,Δv x,i,j ,Δv y,i,j ,Δx i,e ,Δy i,e ,Δv x,i,e ,Δv y,i,e ,Δx i,obs ,Δy i,obs ]
[0044] Among them, (x i ,y i ,v x,i ,v y,i ) represents the position and velocity of agent i; (Δx i,j ,Δy i,j ,Δv x,i,j ,Δv y,i,j ) represents the relative position and velocity of agent i relative to agent j; (Δx i,e ,Δy i,e ,Δv x,i,e ,Δv y,i,e ) represents the relative position and velocity of agent i relative to the escaper e; (Δx i,obs ,Δy i,obs ) represents the relative position of agent i relative to the obstacles in the perception range. The action space of each agent is defined as a two-dimensional continuous driving force A i ={(F x,i ,F y,i )∣F x,i ,F y,i ∈[-F max ,F max ]},F x,i ,F y,i represents the driving force in the x and y directions, where Fmax is the maximum allowable driving force. The joint observation space of the multi-agent system is O = O1 × O2 × ... × O N , the joint action space is A=A1×A2×...×A N , N is the number of agents. During the training phase, although each agent makes independent decisions based on local observations, the system uses a centralized training framework that allows indirect information sharing between agents and promotes the formation of collaborative behaviors.
[0045] S103. Use SimHash function to hash and discretize the joint state-action pairs, allocate hash counting tables and update the corresponding access times in real time, so as to achieve low-dimensional discrimination and novelty statistics in high-dimensional space.
[0046] Specifically, since the state-action space of the multi-agent pursuit task is high-dimensional and continuous, the system introduces the SimHash function to perform hash discretization on the state-action pairs and establish a counting mechanism. The specific implementation process is as follows:
[0047] First, the joint observation of all agents O = O1 × O2 × ... × O N And the combined action A=A1×A2×...×A N The high-dimensional feature vector s||a is formed by concatenation. Then, the SimHash function φ is used to map the high-dimensional vector of the joint state-action pair into a low-dimensional hash code of fixed length:
[0048] φ(s||a)=sgn(M·(s||a))∈{-1,1} k
[0049] Where M is a k×D dimensional random matrix whose elements are independently sampled from a standard normal distribution, D is the dimension of the concatenated feature vector s||a, and k is the hash code length, which is usually set to 32 or 64. The sgn function projects the result into the {-1,1} space:
[0050]
[0051] The system establishes a counting table for each hash code and counts the access frequency of each state-action pair for the subsequent calculation and distribution of intrinsic incentive rewards. Specifically, a global hash counting table C is set up to record the frequency of occurrence of each state-action pair hash code φ(s||a). During the training process, when the agent completes an interaction with the environment and generates a state-action pair, the system calls the SimHash function to map it to a hash code and increases the corresponding count value in the hash table by 1: n(φ(s||a})=n(φ(s||a))+1. Subsequently, the intrinsic incentive reward is calculated based on the current count value of the hash code:
[0052]
[0053] Where n(φ(s t ||a t )) represents the current count value of the hash code, φ represents the SimHash function, s t ||a t It indicates that the current joint observation space and joint action space are concatenated to form a high-dimensional feature vector.
[0054] The core principle of SimHash is to map similar state-action pairs to identical or similar hash codes, effectively compressing high-dimensional continuous space into low-dimensional discrete space while preserving the similarity between state-action pairs. Combined with a counting mechanism, this approach enables the system to efficiently determine the novelty of state-action pairs, providing a computational basis for intrinsic incentives. The smaller the count value (i.e., the more novel the state-action combination), the greater the intrinsic incentive reward received by the agent, prompting the agent to actively explore less-visited regions of the state space in the environment, effectively addressing the problem of low exploration efficiency in sparse reward environments.
[0055] The hash code length k is a key parameter. Larger values of k provide finer-grained state differentiation but also increase computational complexity. Smaller values of k lead to excessive merging of different states, reducing the accuracy of novelty detection. By adjusting the k value and the coefficients in the reward calculation formula, we can balance the relationship between exploration and exploitation and optimize the agent's learning process.
[0056] S104. Update the hash count table and intrinsic incentive reward in real time for each interaction between the agent and the environment, and save the historical data of the interaction.
[0057] Specifically, during each environmental interaction, the hash counting table is updated in real time, and higher intrinsic incentive rewards are assigned to novel state-action pairs based on the counting results. At the same time, historical data including states, actions, and rewards generated by the interaction are stored to improve the experience pool.
[0058] In this invention, novel state-action pairs fall into two categories: those that appear for the first time and those with relatively low historical exploration frequency. This mechanism adheres to the principle of novelty reward: the reward value is negatively correlated with the exploration frequency of a state-action pair. That is, the lower the exploration frequency, the higher the novelty, and the correspondingly larger intrinsic incentive reward. This design encourages the agent to prioritize exploring unknown or rare state spaces, thereby improving exploration efficiency.
[0059] S105. Integrate intrinsic motivation and environmental rewards to optimize collaborative pursuit strategies through a multi-agent deep reinforcement learning algorithm with centralized training and distributed execution.
[0060] Specifically, the calculated intrinsic incentive reward is weightedly integrated with the external reward provided by the environment:
[0061] r t ′=r t +β·r t +
[0062] Among them, r t For environmental rewards, r t + is the intrinsic incentive reward, and β is the weight coefficient, which is used to adjust the relative importance of the two rewards. The combined reward r t + This data is fed into a multi-agent deep reinforcement learning framework based on the MADDPG algorithm. This framework leverages historical interaction data and employs a multi-agent deep reinforcement learning approach based on centralized training and distributed execution. The Actor network generates actions based on local observations, while the Critic network uses global information to evaluate the value of joint actions. This continuously optimizes the agents' collaborative pursuit strategies, achieving efficient collaboration and improved obstacle avoidance capabilities. Each time training data is sampled from the experience pool, the corresponding intrinsic incentive reward is dynamically calculated based on the latest hash count table.
[0063] Therefore, after each interaction with the environment (i.e., the agent performs a set of actions and obtains the next state and reward), the system immediately updates the count value of the corresponding state-action pair in the hash count table. It is important to note that the intrinsic incentive reward is not directly stored in the experience replay buffer. Instead, the corresponding intrinsic incentive reward is dynamically calculated based on the current latest hash count table each time training data is sampled from the experience pool. This real-time update and dynamic calculation mechanism avoids the problem of inconsistent rewards caused by different count values for the same state-action pair at different training stages, ensuring the timeliness of the incentive mechanism and the stable convergence of algorithm training.
[0064] S106. Utilize the trained policy network to autonomously generate pursuit action instructions based on the current environment state, and drive each intelligent agent to collaboratively perform the pursuit task.
[0065] Specifically, after training is completed, the trained Actor strategy network is deployed and applied. Each agent autonomously generates and executes corresponding pursuit actions based on the currently acquired environmental state information, achieving target capture and dynamic adaptation to the environment.
[0066] The present invention deploys the trained Actor policy network into actual application scenarios or simulation test environments. During the operation phase, each agent autonomously generates control actions based on the current local observation state through the policy network, without the need for additional calculation of intrinsic incentive rewards. The policy network can directly output appropriate action instructions based on the input environmental state, driving the group of agents to collaboratively perform pursuit tasks, achieving efficient capture and dynamic response to the target, while effectively avoiding obstacles in the environment. This strategy realizes adaptive collaborative control of multi-agent systems under conditions of incomplete information and limited perception, significantly improving the system's execution capability and task completion efficiency in complex environments.
[0067] Figure 2 This is a schematic diagram of the multi-agent collaborative pursuit simulation environment. Figure 2 As shown in the example, the map size is fixed at a 5m×5m square area, containing five pursuing agents (blue circles), one escaping target (red circle), and two randomly sized obstacles (black circles). Each pursuing agent has a perception radius of 0.75m, and can only perceive the position and velocity of other agents, targets, and obstacles within a 0.75m radius of itself. The escaping target is captured within a 0.5m radius: when at least two pursuing agents simultaneously enter a 0.5m range, the target is considered captured. The physical parameters of the pursuer (blue agent) are: radius 0.125m, mass 1.6kg, and maximum speed 1.0m / s. The physical parameters of the escaping target (red agent) are: radius 0.075m, mass 0.8kg, and maximum speed 1.0m / s. All entities are treated as ideal rigid bodies, subject to frictional damping during motion, with a damping coefficient set to 0.1. Collision elasticity coefficients between agents and between agents and obstacles are set to 0.01 to ensure sufficient rebound when collisions occur and to prevent overlap and penetration between entities. This environment simulates a multi-agent collaborative pursuit task under conditions of limited perception, multiple obstacles, and complex physical interactions.
[0068] Figure 3 The following is a diagram showing the working principle of the SimHash function. Figure 3As shown in the figure, it shows how to map a two-dimensional continuous state into a binary hash code of fixed length. Specifically, first, for each point in the two-dimensional state space, a k-bit binary code is obtained by calculating the inner product with k random projection vectors and taking the sign. The different colors or boundaries in the figure represent different value regions of the hash code. Similar continuous states are divided into the same or similar hash units, realizing clustering compression of the state space. It can be seen intuitively that the two-dimensional space is divided into several irregular but continuous regions by SimHash mapping, and each region corresponds to a unique hash code. This space division method based on SimHash can effectively retain the similarity relationship between the original states, that is, adjacent state points will be mapped to the same or similar hash code region with a high probability. When the dimension of the state space is small (2 dimensions), the division effect can be intuitively displayed through visualization. However, in actual multi-agent tasks, the space composed of state-action pairs is often high-dimensional and continuous. In this case, SimHash can still achieve discrete partitioning that preserves similarity in high-dimensional space. However, due to the high dimension and complex spatial structure, it is difficult to intuitively present it through a two-dimensional graph. Therefore, the partitioning effect in the high-dimensional case will not be elaborated here.
[0069] Figure 4 This is a schematic diagram of the Actor-Critic network structure. Figure 4 As shown in the figure, the MADDPG algorithm adopts a centralized training and distributed execution framework, with each agent consisting of a set of independent actor and critic networks. The actor network generates action decisions based on the local observations of the current agent and is responsible for independently controlling the behavior of the corresponding agent during the actual execution phase. During the training phase, the critic network receives the joint state and joint action information of all agents, performs a global value assessment of the current strategy, and provides gradient signals for the optimization of the actor network. During training, all critic networks have access to global information, enabling effective coordination and policy updates among multiple agents. During model deployment and actual execution, each actor network relies only on its own local observations. This structure not only enhances the collaborative capabilities of the multi-agent system, but also balances scalability and flexibility, making it suitable for collaborative pursuit tasks in environments with incomplete information.
[0070] Figure 5 This is the reward convergence curve of multi-agent cooperative pursuit. Figure 5As shown in the figure, the upper half shows the individual reward trends of each pursuing agent during training. Each curve corresponds to a pursuer, reflecting its cumulative reward over different training episodes. The lower half shows the average reward of all pursuing agents as training progresses. The horizontal axis represents the number of training episodes, and the vertical axis represents the corresponding reward value. As shown in the figure, as training progresses, the reward curves of each pursuer gradually stabilize, moving from early large fluctuations to a higher overall reward level, indicating that each agent has gradually learned an efficient pursuit strategy. In the early stages of training, due to the lack of convergence in the strategy, performance varies greatly between individuals, resulting in significant reward fluctuations. After a certain number of training episodes, the rewards of all pursuers increase significantly, eventually converging to a stable range. The average reward curve in the lower half further reflects the overall collaborative effect of the team. It can be observed that the average reward increases significantly with the number of training episodes and stabilizes in the later stages, indicating that the proposed multi-agent reinforcement learning method based on intrinsic incentives can effectively improve pursuit efficiency and achieve efficient collaboration among multiple agents. Overall, the reward converges quickly and is stable, further verifying the effectiveness and superiority of this method in complex collaborative pursuit tasks.
[0071] Figure 6 This is a sequence diagram demonstrating a multi-agent collaborative pursuit case. Figure 6 As shown, the blue curve represents the trajectory of each pursuing agent, and the red curve represents the movement path of the target escaper. Over time, multiple pursuers, starting from different initial positions, continuously adjust their movements along their respective trajectories, collaborating to gradually close the distance to the target. Ultimately, all pursuing agents successfully capture the target, and the escaper's trajectory ends in the encircled area, indicating the successful completion of the pursuit mission. This sequence diagram intuitively illustrates the entire process of multi-agent collaborative encirclement, dynamic tracking, and final capture of the target.
[0072] The intelligent swarm collaborative pursuit method and system under incomplete information conditions provided by this invention not only significantly improves exploration efficiency and collaborative capabilities under restricted perception, but also enables efficient target capture and autonomous obstacle avoidance in complex dynamic environments. Practical results demonstrate the invention's excellent generalization and practical application value, providing effective technical support for task execution and intelligent collaboration in multi-agent systems under incomplete information conditions.
[0073] The above specific implementation methods further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-agent collaborative pursuit method under incomplete information conditions, characterized by: The following steps are involved: S101. Establishing an incomplete information hunting environment; S102, obtaining local perception information of each agent and constructing a joint state space and action space; S103. Use the SimHash function to hash the joint state-action pairs, assign a hash count table, and update the corresponding access count in real time to achieve low-dimensional discrimination and novelty statistics in high-dimensional space; the more novel the state-action combination, the greater the intrinsic incentive reward the agent receives; S104: Update the hash table and intrinsic incentive reward in real time for each interaction between the agent and the environment, and save the historical data of the interaction; S105, integrating intrinsic motivation and environmental rewards, optimizing collaborative pursuit strategies through a multi-agent deep reinforcement learning algorithm with centralized training and distributed execution; S106. Utilize the trained policy network to autonomously generate pursuit action instructions based on the current environment state, and drive each intelligent agent to collaboratively perform the pursuit task.
2. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1, wherein S101 comprises: Construct a multi-agent collaborative pursuit simulation environment, which includes multiple homogeneous pursuers and escapees, as well as randomly distributed obstacles; A limited perception radius is set for each agent so that it can only obtain partial environmental information, forming a pursuit scenario under incomplete information conditions.
3. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1, characterized in that: The S102 includes: The state information of each agent, the relative state information of other agents within the perception range, and the relative position information of obstacles are collected to form a joint state space and action space, which are used to describe the collaboration and decision-making basis among multiple agents. The state is represented by position and velocity, and the action is defined as the driving force in the x and y directions. The x direction refers to the horizontal direction, and the y direction refers to the vertical direction.
4. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1 or 3, characterized in that: During the training phase, each agent makes independent decisions based on local observations, and a centralized training framework is used to allow indirect information sharing between agents to promote the formation of collaborative behavior.
5. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1 is characterized in that: The S103 includes: The SimHash function is used to convert the joint state-action pair into a low-dimensional hash code of fixed length, and a counting table is established for each hash code to count the access frequency of each state-action pair for the subsequent calculation and distribution of intrinsic incentive rewards.
6. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1 or 5, characterized in that: During training, when the agent completes an interaction with the environment and generates a state-action pair, the SimHash function is called to map it into a hash code, and the corresponding count value in the hash table is increased by 1. Subsequently, the intrinsic incentive reward is calculated based on the current count value of the hash code: Where n(φ(s t ||a t )) represents the current count value of the hash code, φ represents the SimHash function, s t ||a t It indicates that the current joint observation space and joint action space are concatenated to form a high-dimensional feature vector.
7. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1, characterized in that: The S104 includes: During each interaction with the environment, the hash counting table is updated in real time, and higher intrinsic incentive rewards are assigned to novel state-action pairs based on the counting results. At the same time, historical data including state, action, and reward generated by the interaction is stored to improve the experience pool.
8. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1 is characterized in that: The S105 includes: The system combines intrinsic incentive rewards with environmental rewards in a weighted manner, utilizes historical interaction data, and adopts a multi-agent deep reinforcement learning method based on centralized training and distributed execution structure. The Actor network generates actions based on local observations, and the Critic network uses global information to evaluate the value of joint actions. The collaborative pursuit strategy of the intelligent agents is continuously optimized to achieve efficient collaboration and improved obstacle avoidance capabilities. Each time training data is sampled from the experience pool, the corresponding intrinsic incentive reward is dynamically calculated based on the current latest hash count table.
9. The multi-agent collaborative pursuit method under incomplete information conditions according to claim 1, characterized in that: The S106 includes: Deploy and apply the trained Actor strategy network. Each agent autonomously generates and executes corresponding pursuit actions based on the currently acquired environmental state information, achieving target capture and dynamic environmental adaptation.
10. An intelligent group collaborative pursuit system under incomplete information conditions, characterized by: include: The initialization module is used to build a multi-agent collaborative pursuit simulation environment, set environmental parameters and constraints, and complete the initialization of the state hash function and counting table; The learning and training module is used to collect local observation information of each agent, construct joint state-action features, realize state novelty discrimination through hash discretization, and optimize the agent strategy using deep reinforcement learning methods by combining intrinsic motivation and environmental rewards; The task execution module is used to receive real-time environment and agent observation information based on the trained strategy model, output the agent's control instructions, and realize multi-agent collaboration and target pursuit.
Citation Information
Cited By
Multi-agent cooperative control method, device and equipment and storage medium
CN120848219A
Game decision-making system and method for cross-domain pursuit of dynamic target
CN121281271A