Training method and device for cluster coverage search model

By employing a cluster coverage search model training method and utilizing reinforcement learning models to optimize path planning for UAV clusters, the problems of coverage efficiency and security in collaborative multi-UAV cluster operations are solved, achieving efficient and safe path planning.

CN120952096APending Publication Date: 2025-11-14NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511237282.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In the path planning of multi-UAV swarm collaborative operations, how to effectively coordinate and cooperate the behavior of each intelligent agent to achieve efficient coverage of the target area, while avoiding path conflicts and resource waste, and ensuring safe flight in a dynamic environment, has become a key challenge.

Method used

A cluster coverage search model training method is adopted. By using the actor network and critic network in the reinforcement learning model, the movement direction of the agent is predicted. The path planning is optimized based on coverage, constraints and reward mechanism, including coverage constraints, time constraints, boundary constraints and collision constraints, to ensure the effectiveness and safety of path planning.

Benefits of technology

It achieves improved efficiency in collaborative coverage of multiple agents, avoids path conflicts and resource waste, and ensures safe flight in complex environments, thus meeting multi-objective optimization requirements for mission needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952096A_ABST
    Figure CN120952096A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster coverage search model training method and device, and the method comprises the steps: building a state matrix of a current time step based on the environment information and position information set of an intelligent agent cluster; inputting the state matrix into an initial reinforcement learning model, and predicting behavior decision information of a next time step through an actor network; controlling each agent to fly according to the moving direction, and determining the coverage rate of the agent cluster to the task space according to the second position information set; the critic network calculates a dominant value of the training according to the state matrixes of the current time step and the next time step and the dominant function; and calculating a loss value of the training based on the state matrix, the behavior decision information and the advantage value, and updating the model according to the loss value. According to the scheme, the intelligent agents can be ensured to reasonably allocate tasks, and path conflicts and resource waste are avoided, so that the overall coverage efficiency of the system is improved, and multi-target balance and optimization are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This solution relates to the field of reinforcement learning technology, and in particular to a training method and apparatus for a cluster coverage search model. Background Technology

[0002] In recent years, unmanned swarm systems composed of unmanned aerial vehicles (UAVs) have demonstrated significant technological advantages and application value in various fields such as military reconnaissance, disaster relief, environmental monitoring, and smart cities. The rapid development of these systems is mainly attributed to breakthroughs in three key technologies: first, advancements in micro-sensor technology, enabling UAVs to carry lighter and more precise detection equipment; second, optimization of autonomous navigation algorithms, including the application of SLAM (Simultaneous Localization and Mapping) technology and deep reinforcement learning path planning algorithms; and finally, innovations in communication protocols, particularly the introduction of 5G and ad hoc networking technologies, which have greatly enhanced swarm collaboration capabilities.

[0003] In terms of spatial coverage, modern drone swarms can overcome the field-of-view limitations of traditional ground equipment and achieve three-dimensional coverage, making them particularly suitable for monitoring tasks in complex terrain conditions.

[0004] In the field of Coverage Path Planning (CPP), the core objective is to design the optimal set of motion paths for a group of drones within a specified two-dimensional or three-dimensional spatial region. This optimization problem needs to satisfy two basic requirements: first, ensuring that all points within the target region are covered by at least one drone; and second, optimizing the overall system performance indicators, including key parameters such as coverage efficiency, energy consumption control, and task completion time. From a computational complexity perspective, the CPP problem is a typical combinatorial optimization problem, and its difficulty increases exponentially with the dimensionality of the search space and the number of agents.

[0005] For a single UAV system, the CPP problem can be transformed into an improved version of the Traveling Salesman Problem (TSP). The solution strategy mainly focuses on designing a continuous, non-intersecting flight path that completely covers the target area. Commonly used solutions include grid decomposition, spiral search algorithms, and computational geometry-based partitioning and covering algorithms. These methods discretize the target area into a regular grid or convex polygons, and then use heuristic search strategies to generate the optimal path.

[0006] However, when the system expands to multi-UAV collaborative operation, the complexity of the CPP problem undergoes a qualitative leap. This complexity is mainly reflected in three key dimensions: First, at the task allocation level, an efficient collaborative mechanism is needed to achieve a balanced distribution of workload, which involves optimizing the region partitioning algorithm and task scheduling strategy. Second, in terms of spatiotemporal constraints, it is essential to ensure that the flight paths of each UAV maintain a safe distance in both time and space to avoid potential collision risks. Finally, the system needs to possess dynamic adaptability, capable of adjusting path planning in real time to cope with unexpected situations, such as weather changes, obstacle appearances, or communication interruptions.

[0007] The current challenges in covering path planning tasks are as follows:

[0008] Most current methods for solving path planning problems are designed for single agents, but as the size of the cluster increases and the characteristics of the cluster increase, the path planning problem will become more complex and uncertain.

[0009] In the multi-agent path planning problem, effectively coordinating and cooperating the behaviors of each agent to achieve efficient coverage of the target area while avoiding path conflicts and resource waste between agents is an important challenge.

[0010] In the process of path planning, how to minimize the total path length, maximize coverage efficiency, reduce energy consumption, and meet multiple objectives such as task deadline is a major challenge in the design of path planning algorithms.

[0011] In dynamic and complex environments, ensuring that agents do not collide with each other and can safely avoid obstacles and other agents in the environment during task execution is a key challenge for path planning algorithms in practical applications. Summary of the Invention

[0012] This solution aims to at least address the technical problems existing in the prior art. To this end, the first aspect of this invention proposes a training method for a cluster coverage search model, the method comprising:

[0013] Obtain the first set of location information of the agent cluster in the task space at the current time step, and construct the first state matrix of the current time step based on the environmental information of the agent cluster and the first set of location information; the task space is the flight space in which the agent cluster executes cluster-covered tasks.

[0014] The first state matrix is ​​input into the initial reinforcement learning model, and the actor network is used to predict the behavioral decision information for the next time step. The behavioral decision information includes the movement direction of each agent. The initial reinforcement learning model includes an actor network and a critic network.

[0015] Each of the intelligent agents is controlled to fly in the direction of movement, and at the end of the next time step, a second set of location information of the intelligent agent cluster is obtained, and the coverage of the intelligent agent cluster to the task space is determined based on the second set of location information.

[0016] The reward score for the agent cluster to perform this cluster coverage task is determined based on the second location information set and the coverage rate.

[0017] Based on the second location information set, construct the state matrix for the next time step, and store the reward score, the behavior decision information, the first state matrix and the second state matrix as a set of data in the buffer. Control the agent to continue to execute the coverage task until the amount of data in the buffer is equal to the preset batch processing data amount, and then train the model.

[0018] During each training session, the critic network calculates the advantage value for this training session based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function.

[0019] The loss value for this training is calculated based on the state matrix at the current time step, the behavioral decision information, and the advantage value, and the actor network and the critic network are updated according to the loss value.

[0020] The updated actor and critic networks are used to continue training until the preset termination condition is met, at which point the training ends and the cluster coverage search model is obtained.

[0021] Optionally, before obtaining the first set of location information of the agent cluster in the task space at the current time step, the method further includes:

[0022] The mission space for the intelligent agent swarm flight is divided into multiple grid cells of the same size, and the mission space is determined by the effective detection range of the sensors carried by the intelligent agents;

[0023] The behavioral decision information for predicting the next time step through the actor network includes:

[0024] The actor network predicts the probability of each agent's action in each movement direction in the action space at the next time step, based on the principle of maximizing the instantaneous coverage of the grid cell by the agent cluster, combined with preset target constraints and the first policy parameters of the actor network.

[0025] The direction of target movement with the highest probability of the action is used as the behavioral decision information for the next time step;

[0026] The action space includes seven movement directions: forward, backward, left, right, stationary, up, and down; the target constraints include coverage constraints, time constraints, boundary constraints, obstacle avoidance constraints, and collision constraints.

[0027] Optionally, the coverage constraint means that when the distance between the agent and the grid cell is not less than a preset distance threshold, the grid cell is considered to have been covered;

[0028] The time constraint represents the smaller of the maximum flight time supported by the agent's battery capacity and the target time; the target time refers to the difference between the maximum possible search time for the agent cluster to complete the task and the task's start time.

[0029] The boundary constraint indicates that the movement path of the intelligent agent is not allowed to exceed the preset boundary range.

[0030] The obstacle avoidance constraint indicates that the distance between the agent and the obstacle is not less than a preset first distance;

[0031] The collision constraint indicates that the distance between the agents is not less than a preset second distance.

[0032] Optionally, determining the reward score for the agent cluster performing this coverage task based on the second location information set and the coverage rate includes:

[0033] The compliance status of the intelligent agent cluster with the boundary constraints, obstacle avoidance constraints, and collision constraints is determined based on the second location information set, and the constraint reward is determined based on the compliance status.

[0034] Based on the second location information set, determine whether an agent has moved to a previously uncovered grid cell. If so, give a positive movement reward; otherwise, give a negative movement reward.

[0035] When the flight time of the agent reaches the time corresponding to the time constraint, it is determined whether the coverage rate of the agent cluster on the grid cell exceeds the target coverage rate according to the second location information set. If yes, a positive coverage reward is given; if no, a negative coverage reward is given.

[0036] The constraint reward, the movement reward, and the coverage reward are weighted and summed to obtain the reward score for the agent cluster to perform this area coverage task.

[0037] Optionally, the step of calculating the loss value for this training based on the state matrix at the current time step, the behavioral decision information, and the advantage value, and updating the actor network and the critic network according to the loss value, includes:

[0038] The actor network calculates the action probability ratio based on the state matrix of the current time step and the behavior decision information, and calculates the first loss value of the actor network in the current training round based on the action probability ratio and the advantage value.

[0039] The Critic network calculates the state value of the current time step and the next time step according to the state value estimation function, and calculates the second loss value of the Critic network in the current training round based on the state value;

[0040] The parameters of the actor network are updated based on the first loss value, and the parameters of the critic network are updated based on the second loss value.

[0041] Optionally, the action probability of the target movement direction with the highest action probability is the first probability; the actor network calculates the action probability ratio based on the state matrix of the current time step and the behavior decision information, and calculates the first loss value of the actor network in the current training round based on the action probability ratio and the advantage value, including:

[0042] The actor network calculates the second probability of executing the target movement direction under the state matrix at the current time step based on the second policy parameters at the current time step.

[0043] Calculate the ratio between the second probability and the first probability to obtain the action probability ratio;

[0044] The action probability ratio is clipped according to a preset clipping function and the dominance value to obtain a clipping ratio;

[0045] The pruning ratio is used as the first loss value of the actor network in the current training round.

[0046] Optionally, the Critic network calculates the state value of the current time step and the next time step according to the state value estimation function, and calculates the second loss value of the Critic network in the current training round based on the state value, including:

[0047] The Critic network calculates the state value of the current time step and the state value of the next time step according to a preset state value estimation function;

[0048] The time difference error is calculated based on the state value of the current time step, the state value of the next time step, and a preset time difference error function.

[0049] The second loss value of the Critic network in the current training round is determined based on the time difference error.

[0050] A second aspect of the present invention provides a training apparatus for a cluster coverage search model, the apparatus comprising:

[0051] The state matrix construction module is used to obtain the first set of location information of the agent cluster in the task space at the current time step, and construct the first state matrix of the current time step based on the environmental information of the agent cluster and the first set of location information; the task space is the flight space in which the agent cluster executes cluster-covered tasks.

[0052] The behavior decision module is used to input the first state matrix into the initial reinforcement learning model and predict the behavior decision information for the next time step through the actor network. The behavior decision information includes the movement direction of each agent. The initial reinforcement learning model includes an actor network and a critic network.

[0053] An execution module is used to control each of the intelligent agents to fly in the direction of movement, and to acquire a second set of location information of the intelligent agent cluster at the end of the next time step, and to determine the coverage of the intelligent agent cluster of the task space based on the second set of location information.

[0054] The reward module is used to determine the reward score for the agent cluster to perform this cluster coverage task based on the second location information set and the coverage rate;

[0055] The training module is used to construct the state matrix for the next time step based on the second location information set, and store the reward score, the behavior decision information, the first state matrix and the second state matrix as a set of data into a buffer, and control the agent to continue to execute the coverage task until the amount of data in the buffer is equal to the preset batch processing data amount, and then train the model.

[0056] The advantage value calculation module is used to calculate the advantage value of the current training session based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function during each training session.

[0057] The network update module is used to calculate the loss value of this training based on the state matrix of the current time step, the behavior decision information, and the advantage value, and update the actor network and the critic network according to the loss value;

[0058] The termination module is used to continue training with the updated actor network and critic network until the preset termination condition is met, at which point the training ends and the cluster coverage search model is obtained.

[0059] A third aspect of the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the training method of the cluster coverage search model as described in the first aspect.

[0060] A fourth aspect of the present invention provides a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a training method for a cluster coverage search model as described in the first aspect.

[0061] The embodiments of the present invention have the following beneficial effects:

[0062] The present invention provides a training method for a cluster coverage search model, the method comprising: acquiring a first set of location information of an agent cluster in a task space at the current time step, and constructing a first state matrix for the current time step based on the environmental information of the agent cluster and the first set of location information; the task space being the flight space in which the agent cluster performs a cluster coverage task; inputting the first state matrix into an initial reinforcement learning model, predicting behavioral decision information for the next time step through an actor network, the behavioral decision information including the movement direction of each agent; the initial reinforcement learning model including an actor network and a critic network; controlling each agent to fly according to the movement direction, and acquiring a second set of location information of the agent cluster at the end of the next time step, and determining the coverage rate of the agent cluster over the task space based on the second set of location information; and based on the second set of location information and the coverage rate... The reward score for the agent cluster performing the current cluster coverage task is determined. Based on the second location information set, a state matrix for the next time step is constructed. The reward score, the behavioral decision information, the first state matrix, and the second state matrix are stored as a set of data in a buffer. The agents continue to perform the coverage task until the amount of data in the buffer equals a preset batch processing data size, at which point model training begins. During each training iteration, the critic network calculates the advantage value for the current training iteration based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function. The loss value for the current training iteration is calculated based on the state matrix of the current time step, the behavioral decision information, and the advantage value. The actor network and the critic network are updated based on the loss value. Training continues using the updated actor network and critic network until a preset termination condition is met, resulting in a cluster coverage search model. This scheme includes four stages: state construction, environment interaction, experience sharing, and batch updating. It effectively coordinates the collaboration of multiple agents in coverage path planning, ensuring that each agent can reasonably allocate tasks, avoiding path conflicts and resource waste, thereby improving the overall coverage efficiency of the system.

[0063] Furthermore, the path planning simultaneously considers multiple objectives, including minimizing the total path length, maximizing coverage efficiency, and reducing energy consumption. By optimizing the agent's action strategy, the algorithm can achieve a balance and optimization of multiple objectives while meeting task requirements. In addition, it enhances the agent's safety and collision avoidance capabilities: the path planning algorithm in this paper can operate effectively in complex and dynamic environments. By adjusting the agent's path in real time, it ensures that no collisions occur during task execution and avoids obstacles and other agents, guaranteeing the safe execution of the task. Attached Figure Description

[0064] Figure 1 This is a flowchart illustrating the steps of a training method for a cluster coverage search model provided in an embodiment of the present invention.

[0065] Figure 2 This is a schematic diagram illustrating the flight direction of an intelligent agent in three-dimensional space, provided as an embodiment of the present invention.

[0066] Figure 3 This is a structural block diagram of a training device for a cluster coverage search model provided in an embodiment of the present invention. Detailed Implementation

[0067] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present solution, and not all embodiments. Based on the embodiments of the present solution, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present solution.

[0068] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, the use of "based on" or "according to" implies openness and inclusiveness, because processes, steps, calculations, or other actions "based on" or "according to" one or more of the stated conditions or values ​​may in practice be based on additional conditions or beyond the stated values.

[0069] During their research, the inventors discovered that modern UAV swarms possess the following significant advantages: In terms of spatial coverage, UAV systems can overcome the field-of-view limitations of traditional ground equipment, achieving three-dimensional coverage, making them particularly suitable for monitoring tasks in complex terrain conditions. Regarding mission adaptability, thanks to their modular design, UAV platforms can quickly replace different types of payload equipment, such as high-resolution optical cameras, infrared thermal imagers, and synthetic aperture radar, thereby meeting diverse mission requirements. In terms of economic efficiency, compared to traditional manned systems, UAV systems not only have lower procurement costs but are also simpler to maintain and more flexible in deployment, significantly reducing mission execution costs. More importantly, in terms of safety, UAV systems can replace human labor in high-risk tasks, such as monitoring nuclear radiation areas and reconnaissance of forest fires, effectively ensuring the safety of personnel.

[0070] Research shows that these advantages of drone swarm systems give them unique value in multiple fields. In environmental monitoring, drone swarms can achieve large-scale, high-frequency ecological monitoring; in disaster response, they can establish emergency communication networks and conduct search and rescue operations immediately after a disaster; in urban management, they can be used for tasks such as traffic monitoring and illegal construction inspections. With the development of artificial intelligence and edge computing technologies, future drone swarm systems will have stronger autonomous decision-making capabilities and collaborative operation efficiency.

[0071] In recent years, unmanned swarm systems composed of unmanned aerial vehicles (UAVs) have demonstrated significant technological advantages and application value in various fields such as military reconnaissance, disaster relief, environmental monitoring, and smart cities. The rapid development of these systems is mainly attributed to breakthroughs in three key technologies: first, advancements in micro-sensor technology, enabling UAVs to carry lighter and more precise detection equipment; second, optimization of autonomous navigation algorithms, including the application of SLAM (Simultaneous Localization and Mapping) technology and deep reinforcement learning path planning algorithms; and finally, innovations in communication protocols, particularly the introduction of 5G and ad hoc networking technologies, which have greatly enhanced swarm collaboration capabilities.

[0072] From a system characteristics perspective, modern UAV swarms offer the following significant advantages: In terms of spatial coverage, UAV systems can overcome the field-of-view limitations of traditional ground equipment, achieving three-dimensional coverage, making them particularly suitable for monitoring tasks in complex terrain conditions. Regarding mission adaptability, thanks to their modular design, UAV platforms can quickly replace different types of payload equipment, such as high-resolution optical cameras, infrared thermal imagers, and synthetic aperture radar, thereby meeting diverse mission requirements. In terms of economic efficiency, compared to traditional manned systems, UAV systems not only have lower procurement costs but are also simpler to maintain and more flexible in deployment, significantly reducing mission execution costs. More importantly, in terms of safety, UAV systems can replace human labor in high-risk tasks, such as monitoring nuclear radiation areas and forest fire reconnaissance, effectively protecting personnel safety.

[0073] Research shows that these advantages of drone swarm systems give them unique value in multiple fields. In environmental monitoring, drone swarms can achieve large-scale, high-frequency ecological monitoring; in disaster response, they can establish emergency communication networks and conduct search and rescue operations immediately after a disaster; in urban management, they can be used for tasks such as traffic monitoring and illegal construction inspections. With the development of artificial intelligence and edge computing technologies, future drone swarm systems will have stronger autonomous decision-making capabilities and collaborative operation efficiency.

[0074] In the field of Coverage Path Planning (CPP), the core objective is to design the optimal set of motion paths for a group of drones within a specified two-dimensional or three-dimensional spatial region. This optimization problem needs to satisfy two basic requirements: first, ensuring that all points within the target region are covered by at least one drone; and second, optimizing the overall system performance indicators, including key parameters such as coverage efficiency, energy consumption control, and task completion time. From a computational complexity perspective, the CPP problem is a typical combinatorial optimization problem, and its difficulty increases exponentially with the dimensionality of the search space and the number of agents.

[0075] For a single UAV system, the CPP problem can be transformed into an improved version of the Traveling Salesman Problem (TSP). The solution strategy mainly focuses on designing a continuous, non-intersecting flight path that completely covers the target area. Commonly used solutions include grid decomposition, spiral search algorithms, and computational geometry-based partitioning and covering algorithms. These methods discretize the target area into a regular grid or convex polygons, and then use heuristic search strategies to generate the optimal path.

[0076] However, when the system expands to multi-UAV collaborative operation, the complexity of the CPP problem undergoes a qualitative leap. This complexity is mainly reflected in three key dimensions: First, at the task allocation level, an efficient collaborative mechanism is needed to achieve a balanced distribution of workload, which involves optimizing the region partitioning algorithm and task scheduling strategy. Second, in terms of spatiotemporal constraints, it is essential to ensure that the flight paths of each UAV maintain a safe distance in both time and space to avoid potential collision risks. Finally, the system needs to possess dynamic adaptability, capable of adjusting path planning in real time to cope with unexpected situations, such as weather changes, obstacle appearances, or communication interruptions.

[0077] The main challenges facing current research include: how to balance global optimization with local autonomous decision-making, how to handle real-time path planning under incomplete information conditions, and how to evaluate the trade-offs between different optimization objectives. These challenges make multi-UAV coverage path planning not only a complex engineering optimization problem, but also a theoretical research challenge involving multiple disciplines. Existing solutions usually require trade-offs between computational complexity, solution accuracy, and real-time requirements, which also represents an important direction for future research in this field.

[0078] The current challenges in covering path planning tasks are as follows:

[0079] Most current methods for solving path planning problems are designed for single agents, but as the size of the cluster increases and the characteristics of the cluster increase, the path planning problem will become more complex and uncertain.

[0080] In the multi-agent path planning problem, effectively coordinating and cooperating the behaviors of each agent to achieve efficient coverage of the target area while avoiding path conflicts and resource waste between agents is an important challenge.

[0081] In the process of path planning, how to minimize the total path length, maximize coverage efficiency, reduce energy consumption, and meet multiple objectives such as task deadline is a major challenge in the design of path planning algorithms.

[0082] In dynamic and complex environments, ensuring that agents do not collide with each other and can safely avoid obstacles and other agents in the environment during task execution is a key challenge for path planning algorithms in practical applications.

[0083] This paper proposes a novel method for UAV swarm area coverage search based on an independent proximal policy optimization algorithm, particularly for coverage path planning tasks. The main objective is to improve the efficiency and effectiveness of area coverage by optimizing the cooperative behavior within the UAV swarm.

[0084] Figure 1 This is a flowchart illustrating the steps of a training method for a cluster coverage search model provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0085] Step 101: Obtain the first location information set of the agent cluster in the task space at the current time step, and construct the first state matrix of the current time step based on the environmental information of the agent cluster and the first location information set; the task space is the flight space in which the agent cluster executes cluster coverage tasks.

[0086] An agent is an entity with a certain degree of autonomy, capable of perceiving its environment, making decisions, and executing actions. A multi-agent system (MAS) is a collection of multiple interacting agents that collaborate, compete, or coordinate to accomplish complex tasks. In this paper, the agent cluster can be a drone swarm.

[0087] Task space refers to the set of all possible actions, target states, and their relationships that an intelligent agent needs to consider when solving a problem or performing a task. In this embodiment of the invention, task space refers to the flight space of an intelligent agent cluster executing cluster-covered tasks.

[0088] State space refers to all the possible states of an environment and a cluster of intelligent agents, such as the location coordinates of a drone swarm.

[0089] The state space is defined as follows:

[0090]

[0091] in:

[0092] G represents the set of coordinates of all points in a 3D grid map.

[0093] M T This represents a set of target regions, defining the area that the agent needs to cover.

[0094] M N This refers to a set of no-movement zones, such as areas with fixed obstacles, which do not change over time during the mission.

[0095] M S This represents a set of starting regions used to define the initial position of the agent at the start of the task. These regions generally remain unchanged during the task.

[0096] M P This represents the set of agent positions, signifying the current positional state of all agents. As time progresses and the task advances, the position of each agent will change.

[0097] In the state space S, the agent's position (x, y, z) and position set are dynamically changing, reflecting the real-time state during the task; while the environmental state, such as the 3D mesh map, target area, restricted area, and starting area, is usually static, providing the constraints of the task. Through this combination of dynamic and static elements, the state space can effectively describe the changes and constraints of the agent in the region search task.

[0098] The first state matrix represents the joint state space of all agents at the current moment. It is a multi-dimensional matrix, with each row corresponding to the state space of one agent.

[0099] As an optional embodiment, the method further includes the following step before step 101:

[0100] The mission space for the flight of the intelligent agent swarm is divided into multiple grid cells of the same size, and the mission space is determined by the effective detection range of the sensors carried by the intelligent agents.

[0101] In the field of UAV swarm path planning, environmental modeling, as a key preprocessing step, has the core task of transforming the real three-dimensional space into a discrete mathematical model that can be processed by a computer. The 3D grid-based representation method has become one of the most mainstream modeling methods due to its simplicity and computational efficiency. This method uses a cubic grid to uniformly divide the task space. The determination of the grid size C requires consideration of two key factors: firstly, it must match the effective detection range of the sensors carried by the UAV (such as RGB cameras, LiDAR, etc.), typically taking C ≤ 0.5R (where R is the effective detection radius of the sensor); secondly, computational complexity must be considered, as an excessively small grid size will lead to a cubic growth of the state space.

[0102] In practical implementation, given the physical dimensions in three-dimensional space (K×K×H meters) and the grid side length C meters, a k×k×h grid matrix can be obtained through discretization, where... This represents the number of discrete elements in the length / width direction. This represents the number of discrete elements in the height direction. This is a rounding-up function. This modeling approach has three significant advantages: First, grid attributes (such as obstacle markers and coverage status) can be directly stored using a 3D matrix, facilitating rapid querying and updating; second, the grid-based neighborhood relationship definition simplifies the implementation of collision detection algorithms; and finally, it has natural compatibility with mainstream graph search algorithms (such as A* and Dijkstra). When deploying N drones for collaborative coverage, a distributed grid map management system can be established, allowing each drone to share local environmental information in real time.

[0103] Suppose N drones are collaboratively performing a coverage reconnaissance mission in a three-dimensional space. This space has a length of K meters, a width of K meters, and a height of H meters, with each grid cell having a side length of C meters. Under this assumption, the entire region is discretized into a grid region of size k×k×h, where k is the upper bound of K / C, and h is the upper bound of H / C. Through the collaborative work of multiple drones, more efficient three-dimensional coverage and reconnaissance can be achieved.

[0104] Step 102: Input the first state matrix into the initial reinforcement learning model, and predict the behavioral decision information for the next time step through the actor network. The behavioral decision information includes the movement direction of each agent. The initial reinforcement learning model includes an actor network and a critic network.

[0105] An initial reinforcement learning model refers to a rudimentary model that has not been fully trained and consists of two core components: an actor network and a critic network. The actor network is responsible for generating actions based on the current state, such as the direction of movement. The critic network evaluates the value of the actions generated by the actors and provides feedback to improve the policy. The behavioral decision information is the output of the actor network, typically the action instructions for each agent.

[0106] The initial reinforcement learning model is a decentralized partially observable Markov decision process (Dec-POMDP). The Dec-POMDP consists of tuples (I, S, A, Ω, T). a ,O,R) represents, where:

[0107] I: Set of intelligent agents, representing multiple intelligent agents in an unmanned system;

[0108] S: State space, representing the set of all possible states that the agent can be in;

[0109] A: Action space, {A i} i∈I This represents the set of actions that each agent i can perform;

[0110] Ω: Observation space, {Ω i} i∈l This represents the set of information that each agent i can observe;

[0111] T a : State transition function, T a (s,a,s′) represents the probability of transitioning from state s to state s′ under action a;

[0112] O: Observation function, O(o,s′,a) represents the probability of obtaining observation o after performing action a in state s′;

[0113] R: Reward function, R(s,a,s′) represents the immediate reward obtained when transitioning from state s to state s′ after performing action a in state s.

[0114] As an optional embodiment, step 102 includes:

[0115] Step 1021: Based on the principle of maximizing the instantaneous coverage of the grid cells by the agent cluster, and combined with preset target constraints and the first policy parameters of the actor network, the actor network predicts the probability of each agent's actions in each movement direction in the action space at the next time step; the action space includes seven movement directions, namely forward, backward, left, right, stationary, upward, and downward; the target constraints include coverage constraints, time constraints, boundary constraints, obstacle avoidance constraints, and collision constraints.

[0116] The Actor network selects its movement direction based on the current state of the agent cluster and the instantaneous coverage of the network units, so as to maximize the grid coverage in the next time step, allowing the agent cluster to explore as many uncovered areas as possible.

[0117] The first policy parameter is an internal parameter of the Actor network, typically the weights of the neural network. The first policy parameter learns how to balance coverage and constraints using historical training data.

[0118] The action space defines the operations that an agent can perform during a task, guiding it to select the center point of a target grid as its moving target. In a 3D space search task, by traversing all grid center points, the agent can achieve full coverage of a specified area, thus efficiently completing the search.

[0119] The goal of multi-agent coverage path planning is to design optimal strategies that enable agents to autonomously choose the best direction, thereby improving area coverage efficiency. The action space includes seven directions: east, west, south, north, ascending, descending, and stationary. The joint action space size is 7. N Where N represents the number of agents. The trained algorithm model assigns an action to each agent, and finally, the agents cooperate to complete full coverage in three-dimensional space. The action space is represented as follows:

[0120]

[0121] in, This represents the action of the i-th agent. Under this setting, by optimizing the agent's action selection, efficient and complete coverage of the target area can be ensured.

[0122] Each agent has seven possible movement directions, and the Actor network prioritizes the direction that covers the most new grid cells. While maximizing coverage, it must satisfy coverage constraints, time constraints, boundary constraints, obstacle avoidance constraints, and collision constraints.

[0123] The Actor network outputs an action probability distribution, for example: forward: 0.6, backward: 0.1, left: 0.1, right: 0.1, stationary: 0.05, rising: 0.03, falling: 0.02.

[0124] In UAV swarm coverage path planning, a 3D mesh-based environmental coverage assessment model is a key method for quantifying task completion. This model discretizes the sensor field of view of each UAV into 3D mesh cells (voxels) matching its resolution. The basic principle can be stated as follows: when a UAV's trajectory crosses a specific mesh cell, that cell is marked as covered (state value set to 1), while unvisited cells remain uncovered (state value kept at 0). This binary state representation has the following important characteristics:

[0125] (1) Irreversibility of coverage state: The state transition of each grid cell follows a one-way transition rule, that is, the transition from 0 to 1 is irreversible. This modeling method reflects the reconnaissance requirement of "first coverage is effective" in practical applications. No matter how many times the UAV repeatedly visits the cell, its state value always remains 1.

[0126] (2) Coverage calculation model:

[0127] At any evaluation time t, the instantaneous coverage η(t) of the system can be expressed as:

[0128]

[0129] Where Nc(t) is the number of grid cells with state 1 at time t, and M = k × k represents the total number of grid cells that need to be covered in the environment model. This metric provides a standardized measure for evaluating system performance.

[0130] (3) Algorithm optimization basis: Based on this coverage model, an objective function J = 1 - η(T) can be constructed, where T is the task deadline. By minimizing J, the path planning algorithm can be driven to seek the optimal coverage strategy. Experimental data show that, under typical parameter configurations (k = 20, h = 4), this model can keep the computational complexity of coverage evaluation at the O(M) level, meeting the real-time requirements.

[0131] Figure 2 This is a schematic diagram illustrating the flight direction of an intelligent agent in three-dimensional space, as provided in an embodiment of the present invention.

[0132] like Figure 2 As shown, in a three-dimensional space at height h, the flight direction of each UAV is discretized into seven fixed directions: forward, backward, left, right, stationary, ascending, and descending, labeled 1 to 7 respectively. Considering battery capacity limitations, the maximum number of flight steps for each UAV is set to 200 steps. Therefore, within this three-dimensional space, the UAV can not only cover areas on the horizontal plane but also cover areas at different heights through vertical movement, thus achieving more comprehensive spatial reconnaissance.

[0133] As an optional embodiment, the coverage constraint means that when the distance between the agent and the grid cell is not less than a preset distance threshold, the grid cell is considered to be covered; the time constraint means the smaller of the maximum flight time supported by the agent's battery capacity and the target time; the target time refers to the difference between the maximum possible search time for the agent cluster to complete the task and the start time of the task;

[0134] The boundary constraint indicates that the movement path of the intelligent agent is not allowed to exceed the preset boundary range.

[0135] The obstacle avoidance constraint indicates that the distance between the agent and the obstacle is not less than a preset first distance;

[0136] The collision constraint indicates that the distance between the agents is not less than a preset second distance.

[0137] This model introduces multiple constraints, including coverage constraints, time constraints, boundary constraints, and collision constraints, to simulate the limitations of real-world environments, thereby ensuring the practicality and effectiveness of the algorithm. Coverage and time constraints ensure that the drone swarm completes its task within a specified time, boundary constraints prevent drones from exceeding boundaries or crashing to the ground, and collision constraints prevent drones from colliding with obstacles or within the swarm itself.

[0138] The following provides a detailed explanation of each constraint:

[0139] (1) Time constraints

[0140] The time constraint defines the method for calculating the search time of the drone swarm. As part of the time constraint, the drone's battery life is converted into a finite number of flight steps for each drone, i.e., the maximum flight time for each drone. This method ensures effective control of the drone's flight time while taking battery life into account.

[0141] t sea =max{t 1,sea ,…,t i,sea ,…,t N,sea}

[0142] t in the formula sea This represents the maximum possible search time for the drone swarm to complete its mission. This time is the time it takes for the last of the N drones to finish its mission. This is because the drone swarm uses a region-based coverage method when executing its mission; the mission is considered complete only when all target areas are completely covered.

[0143] t end =min{t sea -t start ,t battery}

[0144] In the formula, t end This represents the actual search time of the drone swarm. Here, t sea t represents the maximum possible search time for the drone swarm to complete the task. start Indicates the start time of the task, t battery This indicates the maximum flight time supported by the drone's battery capacity. The formula explains that the mission end time depends on the shorter of the time from mission start to maximum search time and battery endurance.

[0145] (2) Boundary constraints

[0146] For illegal actions, such as going beyond boundaries, colliding with obstacles or other drones, dangerous maneuvers will be predicted within the warning area. If an illegal next action of the drone is predicted, the drone will be forced to remain in its current position until its subsequent actions are deemed legal.

[0147] Boundary constraints define that, in area coverage missions, drones are not allowed to exceed predetermined boundary ranges, ensuring that they always perform missions within the mission area. The formula is defined as follows:

[0148]

[0149] The left side of the inequality represents the position of the drone in three-dimensional space (x). i,Uav ,y i,Uav ,z i,Uav (x) must be located within the predetermined area. min ,y min ,z min ) and (x max ,y max ,z max These are the minimum and maximum boundary values ​​of the region, respectively. This constraint ensures that the agent does not exceed the boundary of the target region when performing a region coverage task, thereby improving the safety and effectiveness of task execution.

[0150] (3) Obstacle avoidance constraints:

[0151] Obstacle avoidance constraints define that in area coverage tasks, the distance between the drone and obstacles must not be less than a certain distance threshold to ensure that the drone will not collide with the obstacles. The formula is defined as follows:

[0152]

[0153] The left side of the inequality represents the position of the drone in three-dimensional space (x). Uav ,y Uav ,z Uav ) and a point (x) in the obstacle regionobs ,y obs ,z obs The Euclidean distance between ) and r obs It is a predefined distance threshold used to determine the minimum safe distance between the drone and obstacles. This ensures that the drone avoids collisions with obstacles when performing area coverage tasks, thereby improving the safety and reliability of the mission.

[0154] (4) Collision constraints:

[0155] Collision constraints define that, in area coverage tasks, the distance between drones and other drones must not be less than a certain distance threshold to ensure that drones do not collide. The formula is defined as follows:

[0156]

[0157] The left side of the inequality represents the position (x) of a drone in three-dimensional space. i,Uav ,y i,Uav ,z i,Uav ) and the position of another drone in three-dimensional space (x j,Uav ,y j,Uav ,z j,Uav The Euclidean distance between ) and r collide It is a predefined distance threshold used to determine the minimum safe distance between drones and other drones. This ensures that drones can avoid collisions when performing area coverage missions, thereby improving mission coordination and safety.

[0158] Step 1022: Use the target movement direction with the highest probability of the action as the behavior decision information for the next time step.

[0159] From the action probability distribution output by the Actor network, the direction with the highest probability is selected as the final decision. For example, in the example action probability distribution in step 1021, "forward" with the highest action probability is selected as the target movement direction.

[0160] Step 103: Control each of the intelligent agents to fly in the direction of movement, and at the end of the next time step, obtain the second location information set of the intelligent agent cluster, and determine the coverage of the intelligent agent cluster on the task space based on the second location information set.

[0161] After determining the behavioral decision information for the next time step, all agents execute their selected actions in parallel and proceed to the next time step.

[0162] At the next time step, the agent's position is updated, and the coverage of the agent cluster over the task space is updated based on the agent's second position information.

[0163] Step 104: Determine the reward score for the intelligent agent cluster to perform this cluster coverage task based on the second location information set and the coverage rate.

[0164] In reinforcement learning, reward scores are key feedback signals used to evaluate the quality of task completion by a cluster of agents after a specific action.

[0165] In the collaborative search mission of a drone swarm, the agent receives a corresponding reward value from the environment at each time step to help it efficiently complete the region search task in three-dimensional space.

[0166] As an optional embodiment, step 104 includes:

[0167] Step 1041: Determine the compliance status of the intelligent agent cluster with the boundary constraints, obstacle avoidance constraints, and collision constraints based on the second location information set, and determine the constraint reward based on the compliance status.

[0168] Constraint rewards are a key mechanism used to guide a swarm of intelligent agents to adhere to preset rules during task execution. The system uses the agent swarm's second set of positional information—its latest positional state after an action—to detect violations of boundary constraints, obstacle avoidance constraints, and collision constraints, and calculates the corresponding constraint rewards accordingly. Constraint rewards are typically negative, i.e., penalties.

[0169] R penalty This represents a constraint-based reward, specifically defined as follows:

[0170]

[0171] Constraint rewards are penalties imposed when an agent violates constraints. For example, if an agent collides with a boundary, obstacle, or other agent, or attempts to move outside a restricted area, these actions are considered illegal and will be rewarded with a negative reward r. penalty Actions that do not trigger these situations will not be punished.

[0172] Step 1042: Determine whether an agent has moved to a previously uncovered grid cell based on the second location information set. If so, give a positive movement reward; if not, give a negative movement reward.

[0173] R move This refers to the mobile reward, specifically defined as follows:

[0174]

[0175] Movement rewards are given based on the agent's exploration progress at each step. For example, if the agent moves to a previously uncovered area, it receives a positive movement reward r. uncovered This encourages exploration of uncovered areas; if the behavior results in repeatedly covering already explored areas, a negative movement reward r is given. covered .

[0176] Step 1043: When the flight time of the agent reaches the time corresponding to the time constraint, determine whether the coverage rate of the agent cluster on the grid cell exceeds the target coverage rate according to the second location information set. If yes, a positive coverage reward is given; if no, a negative coverage reward is given.

[0177] R coverage The coverage reward is defined as follows:

[0178]

[0179] When the coverage of the agent cluster exceeds the target coverage, the task is considered complete, and a large one-time positive reward is given. For example, if the global coverage exceeds 95% when the specified time corresponding to the time constraint is reached, the agent cluster receives a high reward r. complete Conversely, if the target coverage is not achieved within the specified time, the search task is terminated, and a corresponding negative reward / penalty is imposed. uncomplete .

[0180] Step 1044: Perform a weighted summation of the constraint reward, the movement reward, and the coverage reward to obtain the reward score for the agent cluster performing this area coverage task.

[0181] The reward score obtained during the training of the agent for this task can be represented by the following formula:

[0182]

[0183] Where T represents the total number of time steps of the task, R coverage Indicates coverage reward, R penalty R represents the constraint reward. move This represents the movement reward, where α, β, and γ are the corresponding weighting parameters used to balance the influence of different rewards.

[0184] Step 105: Construct the state matrix for the next time step based on the second location information set, and store the reward score, the behavior decision information, the first state matrix and the second state matrix as a set of data in the buffer, and control the agent to continue to execute the coverage task until the amount of data in the buffer is equal to the preset batch processing data amount, and then train the model.

[0185] The data from each task execution is stored in a buffer. Model training only begins when the amount of data in the buffer reaches the preset batch processing data size.

[0186] The system utilizes a centralized replay mechanism to manage experience. Each agent filters its acquired experience samples based on criteria such as importance and novelty, uploading only high-value data to a shared buffer. This mechanism enables knowledge transfer and sample diversity accumulation across agents, while avoiding the transmission of low-value redundant samples, thus improving communication efficiency.

[0187] Step 106: During each training session, the critic network calculates the advantage value for this training session based on the state matrix of the current time step, the state matrix of the next time step, and the preset advantage function.

[0188] In reinforcement learning, the advantage function measures the degree of advantage of taking a specific action relative to the average policy in a given state. Calculating the advantage function helps in better understanding and adjusting policy selection, and is commonly used in Actor-Critic algorithms to improve policy efficiency.

[0189] The advantage function A(s,a) represents the relative advantage of taking action a relative to the average policy, given state s. Mathematically, the advantage function is defined as:

[0190] A(s,a)=Q π (s,a)-V π (s)

[0191] Among them, Q π (s,a) is the action-value function, representing the expected reward an agent can obtain by taking a specific action a in state s and then following policy π. π (s) is the state-value function, representing the expected reward that an agent can obtain when starting from a specific state s and following a specific policy π.

[0192] Step 107: Calculate the loss value for this training based on the state matrix of the current time step, the behavior decision information, and the advantage value, and update the actor network and the critic network according to the loss value.

[0193] The Independent Proximal Policy Optimization (IPPO) algorithm used in this invention is an innovative multi-agent reinforcement learning method. While maintaining the advantages of the PPO algorithm, it solves the multi-agent cooperation problem through a unique architectural design. The core innovation of this algorithm lies in its hybrid architecture combining independent individual training with global information sharing, which retains the flexibility of distributed decision-making while achieving effective collaboration at the group level.

[0194] The Actor network's task is to optimize the policy π, and its loss function is based on the policy gradient and the advantage value. The Critic network's task is to accurately predict the state value V(s), and its loss function typically uses the mean squared error.

[0195] As an optional embodiment, step 107 includes:

[0196] Step 1071: The actor network calculates the action probability ratio based on the state matrix of the current time step and the behavior decision information, and calculates the first loss value of the actor network in the current training round based on the action probability ratio and the advantage value.

[0197] The Actor network is updated based on optimizing a specific objective function. This objective function aims to adjust the parameters of the Actor network to improve the expected return of the policy, while limiting the magnitude of changes brought about by policy updates to ensure the stability of the learning process.

[0198] The action probability ratio rt(θ) represents the action a to be selected under the current policy parameters. t The probability of selecting the desired action is the ratio of the probability of selecting the same action under the parameters of the old policy. This is used to adjust the stride of the policy update, preventing excessive shifts during the update process and ensuring training stability.

[0199] As an optional embodiment, the probability of the target movement direction with the highest probability of action is the first probability; step 1071 includes:

[0200] Step 10711: The actor network calculates the second probability of executing the target movement direction under the state matrix of the current time step based on the second policy parameters of the current time step.

[0201] Step 10712: Calculate the ratio between the second probability and the first probability to obtain the action probability ratio;

[0202] Step 10713: Clip the action probability ratio according to the preset clipping function and the dominance value to obtain the clipping ratio;

[0203] Step 10714: Use the pruning ratio as the first loss value of the actor network in the current training round.

[0204] In steps 10711-10714, the formula for calculating the action probability ratio is as follows:

[0205]

[0206] Where, π old (a t |s t The probability of the old policy is given by the agent based on state s before the policy is updated. t Select action a t The probability. This value is usually recorded before the policy update and used for subsequent ratio calculations. In this embodiment of the invention, this value is the first probability corresponding to the target movement direction in the behavior decision information of the actor network in step 102.

[0207] π θ (a t |s t ) represents the current policy probability, which is the probability that the agent, under the current second policy parameter θ, determines based on the state s. t Select action a t The probability is given by the second policy parameter of the current actor network.

[0208] This article uses a special clipping function L CLIP This function helps avoid training instability caused by excessively large update steps. The pruning function considers the expected performance when the policy ratio is restricted to the range [1-∈, 1+∈]. By limiting the range of policy ratio variation, the volatility of policy updates is reduced, increasing the robustness of the algorithm.

[0209] The clipping function includes an advantage value as a parameter, which clips the action probability ratio to obtain the clipping ratio. The clipping function is as follows:

[0210]

[0211] Where A is the advantage value, E represents the expectation, and rt(θ) represents the action probability ratio. [1-∈, 1+∈] represents the range of the action probability ratio.

[0212] The pruning ratio is used as the first loss value for the actor network in the current training round.

[0213] Step 1072: The Critic network calculates the state value of the current time step and the next time step according to the state value estimation function, and calculates the second loss value of the Critic network in the current training round based on the state value.

[0214] The Critic network attempts to learn a value function that estimates the expected reward, i.e., the state value, obtained by following the current policy in a given state s. Based on the state value, the second loss value of the Critic network in the current training round is calculated.

[0215] As an optional embodiment, step 1072 includes:

[0216] Step 10721: The Critic network calculates the state value of the current time step and the state value of the next time step according to a preset state value estimation function;

[0217] Step 10722: Calculate the time difference error based on the state value of the current time step, the state value of the next time step, and the preset time difference error function;

[0218] Step 10723: Determine the second loss value of the Critic network in the current training round based on the time difference error.

[0219] In steps 10721-10723, the time difference error (TD-error) is a key quantity in the Critic network update, measuring the difference between the value function estimate and the actual return. The formula for calculating TD-error is as follows:

[0220] δ t =r t +γV(s t+1 )-V(s t )

[0221] Where γ is a discount factor used to adjust the current value of future rewards. t The instant reward obtained at time step t, V(s) t+1 ) and V(s t ) are the state value function estimates for time steps t+1 and t, respectively.

[0222] V(s t+1 ) and V(s t The value is calculated based on a preset state value estimation function.

[0223] The loss function in Critic is typically the square of the TD-error, used to quantify the accuracy of the value function prediction. The loss function L(φ) is defined as:

[0224]

[0225] Where, δ t L(φ) represents the time difference error, and L(φ) represents the second loss value.

[0226] Step 1073: Update the parameters of the actor network according to the first loss value, and update the parameters of the Critic network according to the second loss value.

[0227] Specifically, the gradient of the parameter θ is calculated to update the Actor network. More specifically, the policy parameters of the Actor network are updated using stochastic gradient ascent, and the update process is as follows:

[0228]

[0229] Where θ represents the policy parameters of the actor network, α is the learning rate of the actor network, and L... CLIP This is the clipping function.

[0230] For the Critic network, the gradient of parameter φ is calculated using the loss function L(φ):

[0231]

[0232] Update the parameters of the Critic network using gradient descent:

[0233]

[0234] Where β is the learning rate of the Critic network.

[0235] By continuously adjusting the Critic network to accurately predict the value function, bias in policy evaluation can be effectively reduced, improving overall learning performance. Accurate updates to the Critic network contribute to stabilizing the Actor network's learning, as the Actor's updates depend on the value estimate provided by the Critic.

[0236] Step 108: Continue training using the updated actor network and critic network until the preset termination condition is met, then end the training to obtain the cluster coverage search model.

[0237] By continuously updating the Actor network and Critic network, the agent cluster gradually optimizes its strategy until it meets the preset stopping criteria, ultimately resulting in a mature cluster coverage search model.

[0238] This paper employs a decentralized multi-agent reinforcement learning framework, where each agent independently optimizes its policy, constructing a complete local learning system including a policy network and a value evaluation network. These networks are trained and make decisions based solely on the agent's own local observations, ensuring the system's scalability and autonomy. Simultaneously, to balance information collaboration and communication overhead among multiple agents, the framework introduces a centralized experience-sharing mechanism. By establishing a global experience replay buffer, each agent can periodically upload its selected key state information and learn from the high-quality experiences of others. During information transmission, only the selected important state data is retained, avoiding the transmission of large amounts of redundant information, effectively improving training efficiency and controlling communication bandwidth usage.

[0239] The entire training process consists of four key phases, which cycle continuously until convergence. These four phases are state construction, environment interaction, experience sharing, and batch update. In the state construction phase, the agent collects current observation information through its local perception module and combines it with historical multi-frame observations to form a time-dependent state representation to capture dynamic changes in the environment. Subsequently, in the environment interaction phase, the agent inputs its current state into its policy network to generate decision actions. The actions of each agent work together on the environment, triggering environmental updates, which in turn return new observations and immediate rewards. These interaction results constitute empirical data, recording the agent's behavioral feedback in the current round.

[0240] During the experience-sharing phase, the system utilizes a centralized replay mechanism to manage experience. Each agent filters its acquired experience samples based on criteria such as importance and novelty, uploading only high-value data to the shared buffer. This mechanism enables knowledge transfer and sample diversity accumulation across agents, while avoiding the transmission of low-value redundant samples, thus improving communication efficiency. In the batch update phase, agents extract experience samples from the shared buffer to optimize and update the policy network and value network. The policy network stably optimizes decision policies through an objective function that reduces the magnitude of policy changes, while the value network continuously improves the accuracy of state evaluation by minimizing the error between prediction and actual reward.

[0241] This framework demonstrates significant advantages in dynamic and complex multi-agent systems. First, because each agent is trained independently, the system can flexibly scale to different numbers of agents, maintaining learning consistency even as the number of agents changes. Newly added agents can quickly benefit from shared experience, accelerating the learning process. Second, the experience filtering and uploading mechanism effectively controls communication bandwidth consumption. Experimental results show that with a linear increase in the number of agents, the communication load only increases by about 15–20%. Third, the decentralized policy decision-making structure enhances system robustness; even if some agents fail, the overall system can still maintain core functionality and continue to perform tasks.

[0242] In practical applications of drone swarms, the algorithm of this invention exhibits three prominent features: stable training process, linear growth in collaborative efficiency, and high adaptability to environmental changes. Therefore, this method is particularly suitable for scenarios with high real-time requirements, high collaborative demands, and strong robustness, such as large-area collaborative search and continuous tracking of dynamic targets.

[0243] The drone swarm coverage search technology of this invention has produced significant beneficial effects and economic benefits in multiple fields. In urban emergency response and public safety assurance, this technology effectively solves the "last mile" search and rescue problem in natural disasters and public safety emergencies by enabling rapid and comprehensive exploration of complex alleyways, significantly improving search and rescue efficiency (estimated to be over 60%). It provides key technical support for disaster assessment and counter-terrorism, and is of great value in protecting people's lives and property and social stability. In the field of smart city construction, this technology provides core support for applications such as automated inspection of urban infrastructure (e.g., bridges, tunnels, high-voltage lines), real-time environmental monitoring, violation investigation, and 3D modeling. It is expected to reduce inspection costs by 40%-50% while improving urban management efficiency by over 30%.

[0244] From an industrial development perspective, the breakthroughs in key technologies such as autonomous navigation, intelligent planning, and collaborative control of unmanned aerial vehicles (UAVs) in this invention will help break down foreign technological barriers, enhance my country's core competitiveness in this field, and is expected to drive the formation of emerging markets, including customized inspection services and urban low-altitude logistics networks, creating significant economic benefits. In terms of operational cost control, the application of UAV swarms can effectively replace manual labor in high-risk environments, expected to reduce labor costs by 60%-70% while reducing operational safety risks by more than 80%, achieving the strategic goal of "machine replacement of human labor." Comprehensive evaluation indicates that the promotion and application of this project will generate significant socio-economic benefits in improving urban governance capabilities, cultivating emerging industries, and reducing operational risks.

[0245] In summary, the training method for the cluster coverage search model provided in this embodiment of the invention includes: acquiring a first set of location information of an agent cluster in the task space at the current time step, and constructing a first state matrix for the current time step based on the environmental information of the agent cluster and the first set of location information; the task space is the flight space in which the agent cluster performs the cluster coverage task; inputting the first state matrix into an initial reinforcement learning model, predicting behavioral decision information for the next time step through an actor network, the behavioral decision information including the movement direction of each agent; the initial reinforcement learning model includes an actor network and a critic network; controlling each agent to fly according to the movement direction, and acquiring a second set of location information of the agent cluster at the end of the next time step, and determining the coverage of the agent cluster on the task space based on the second set of location information; based on the second set of location information and the coverage... The reward score for the agent cluster performing the current cluster coverage task is determined by the rate. Based on the second location information set, a state matrix for the next time step is constructed. The reward score, the behavioral decision information, the first state matrix, and the second state matrix are stored as a set of data in a buffer. The agents continue to perform the coverage task until the amount of data in the buffer equals the preset batch processing data amount, at which point model training begins. During each training iteration, the critic network calculates the advantage value for this training iteration based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function. The loss value for this training iteration is calculated based on the state matrix of the current time step, the behavioral decision information, and the advantage value. The actor network and the critic network are updated based on the loss value. Training continues using the updated actor network and critic network until a preset termination condition is met, resulting in a cluster coverage search model. This scheme includes four stages: state construction, environment interaction, experience sharing, and batch updating. It effectively coordinates the collaboration of multiple agents in coverage path planning, ensuring that each agent can reasonably allocate tasks, avoiding path conflicts and resource waste, thereby improving the overall coverage efficiency of the system.

[0246] Furthermore, the path planning simultaneously considers multiple objectives, including minimizing the total path length, maximizing coverage efficiency, and reducing energy consumption. By optimizing the agent's action strategy, the algorithm can achieve a balance and optimization of multiple objectives while meeting task requirements. In addition, it enhances the agent's safety and collision avoidance capabilities: the path planning algorithm in this paper can operate effectively in complex and dynamic environments. By adjusting the agent's path in real time, it ensures that no collisions occur during task execution and avoids obstacles and other agents, guaranteeing the safe execution of the task.

[0247] Figure 3 This is a structural block diagram of a training device for a cluster coverage search model provided in an embodiment of the present invention. Figure 3 As shown, the device 200 includes:

[0248] The state matrix construction module 201 is used to obtain the first set of location information of the agent cluster in the task space at the current time step, and construct the first state matrix of the current time step based on the environmental information of the agent cluster and the first set of location information; the task space is the flight space in which the agent cluster performs cluster-covered tasks.

[0249] The behavior decision module 202 is used to input the first state matrix into the initial reinforcement learning model and predict the behavior decision information for the next time step through the actor network. The behavior decision information includes the movement direction of each of the agents. The initial reinforcement learning model includes an actor network and a critic network.

[0250] The execution module 203 is used to control each of the intelligent agents to fly in the direction of movement, and to obtain the second position information set of the intelligent agent cluster at the end of the next time step, and to determine the coverage of the intelligent agent cluster of the task space based on the second position information set;

[0251] Reward module 204 is used to determine the reward score of the agent cluster for performing this cluster coverage task based on the second location information set and the coverage rate;

[0252] The training module 205 is used to construct the state matrix of the next time step based on the second location information set, and store the reward score, the behavior decision information, the first state matrix and the second state matrix as a set of data into a buffer, and control the agent to continue to execute the coverage task until the amount of data in the buffer is equal to the preset batch processing data amount, and then train the model.

[0253] The advantage value calculation module 206 is used to calculate the advantage value of the current training based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function during each training session.

[0254] The network update module 207 is used to calculate the loss value of this training based on the state matrix of the current time step, the behavior decision information, and the advantage value, and update the actor network and the critic network according to the loss value;

[0255] The termination module 208 is used to continue training with the updated actor network and critic network until the preset termination condition is met, thus ending the training and obtaining the cluster coverage search model.

[0256] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0257] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0258] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that comply with the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0259] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A training method for a cluster coverage search model, characterized in that, The method includes: Obtain the first set of location information of the agent cluster in the task space at the current time step, and construct the first state matrix of the current time step based on the environmental information of the agent cluster and the first set of location information; the task space is the flight space in which the agent cluster executes cluster-covered tasks. The first state matrix is ​​input into the initial reinforcement learning model, and the actor network is used to predict the behavioral decision information for the next time step. The behavioral decision information includes the movement direction of each agent. The initial reinforcement learning model includes an actor network and a critic network. Each of the intelligent agents is controlled to fly in the direction of movement, and at the end of the next time step, a second set of location information of the intelligent agent cluster is obtained, and the coverage of the intelligent agent cluster to the task space is determined based on the second set of location information. The reward score for the agent cluster to perform this cluster coverage task is determined based on the second location information set and the coverage rate. Based on the second location information set, construct the state matrix for the next time step, and store the reward score, the behavior decision information, the first state matrix and the second state matrix as a set of data in the buffer. Control the agent to continue to execute the coverage task until the amount of data in the buffer is equal to the preset batch processing data amount, and then train the model. During each training session, the critic network calculates the advantage value for this training session based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function. The loss value for this training is calculated based on the state matrix at the current time step, the behavioral decision information, and the advantage value, and the actor network and the critic network are updated according to the loss value. The updated actor and critic networks are used to continue training until the preset termination condition is met, at which point the training ends and the cluster coverage search model is obtained.

2. The method according to claim 1, characterized in that, Before obtaining the first set of location information of the agent cluster in the task space at the current time step, the following steps are also included: The mission space for the intelligent agent swarm flight is divided into multiple grid cells of the same size, and the mission space is determined by the effective detection range of the sensors carried by the intelligent agents; The behavioral decision information for predicting the next time step through the actor network includes: The actor network predicts the probability of each agent's action in each movement direction in the action space at the next time step, based on the principle of maximizing the instantaneous coverage of the grid cell by the agent cluster, combined with preset target constraints and the first policy parameters of the actor network. The direction of target movement with the highest probability of the action is used as the behavioral decision information for the next time step; The action space includes seven movement directions: forward, backward, left, right, stationary, up, and down; the target constraints include coverage constraints, time constraints, boundary constraints, obstacle avoidance constraints, and collision constraints.

3. The method according to claim 2, characterized in that: The coverage constraint means that when the distance between the agent and the grid cell is not less than a preset distance threshold, the grid cell is considered to have been covered. The time constraint represents the smaller of the maximum flight time supported by the agent's battery capacity and the target time; the target time refers to the difference between the maximum possible search time for the agent cluster to complete the task and the task's start time. The boundary constraint indicates that the movement path of the intelligent agent is not allowed to exceed the preset boundary range. The obstacle avoidance constraint indicates that the distance between the agent and the obstacle is not less than a preset first distance; The collision constraint indicates that the distance between the agents is not less than a preset second distance.

4. The method according to claim 2, characterized in that, The step of determining the reward score for the agent cluster to perform this coverage task based on the second location information set and the coverage rate includes: The compliance status of the intelligent agent cluster with the boundary constraints, obstacle avoidance constraints, and collision constraints is determined based on the second location information set, and the constraint reward is determined based on the compliance status. Based on the second location information set, determine whether an agent has moved to a previously uncovered grid cell. If so, give a positive movement reward; otherwise, give a negative movement reward. When the flight time of the agent reaches the time corresponding to the time constraint, it is determined whether the coverage rate of the agent cluster on the grid cell exceeds the target coverage rate according to the second location information set. If yes, a positive coverage reward is given; if no, a negative coverage reward is given. The constraint reward, the movement reward, and the coverage reward are weighted and summed to obtain the reward score for the agent cluster to perform this area coverage task.

5. The method according to claim 2, characterized in that, The step of calculating the loss value for this training based on the state matrix at the current time step, the behavioral decision information, and the advantage value, and updating the actor network and the critic network according to the loss value, includes: The actor network calculates the action probability ratio based on the state matrix of the current time step and the behavior decision information, and calculates the first loss value of the actor network in the current training round based on the action probability ratio and the advantage value. The Critic network calculates the state value of the current time step and the next time step according to the state value estimation function, and calculates the second loss value of the Critic network in the current training round based on the state value; The parameters of the actor network are updated based on the first loss value, and the parameters of the critic network are updated based on the second loss value.

6. The method according to claim 5, characterized in that, The action probability of the target movement direction with the highest action probability is the first probability; the actor network calculates the action probability ratio based on the state matrix of the current time step and the behavior decision information, and calculates the first loss value of the actor network in the current training round based on the action probability ratio and the advantage value, including: The actor network calculates the second probability of executing the target movement direction under the state matrix at the current time step based on the second policy parameters at the current time step. Calculate the ratio between the second probability and the first probability to obtain the action probability ratio; The action probability ratio is clipped according to a preset clipping function and the dominance value to obtain a clipping ratio; The pruning ratio is used as the first loss value of the actor network in the current training round.

7. The method according to claim 6, characterized in that, The Critic network calculates the state value of the current time step and the next time step based on the state value estimation function, and calculates the second loss value of the Critic network in the current training round based on the state value, including: The Critic network calculates the state value of the current time step and the state value of the next time step according to a preset state value estimation function; The time difference error is calculated based on the state value of the current time step, the state value of the next time step, and a preset time difference error function. The second loss value of the Critic network in the current training round is determined based on the time difference error.

8. A training device for a cluster coverage search model, characterized in that, The device includes: The state matrix construction module is used to obtain the first set of location information of the agent cluster in the task space at the current time step, and construct the first state matrix of the current time step based on the environmental information of the agent cluster and the first set of location information; the task space is the flight space in which the agent cluster executes cluster-covered tasks. The behavior decision module is used to input the first state matrix into the initial reinforcement learning model and predict the behavior decision information for the next time step through the actor network. The behavior decision information includes the movement direction of each agent. The initial reinforcement learning model includes an actor network and a critic network. An execution module is used to control each of the intelligent agents to fly in the direction of movement, and to acquire a second set of location information of the intelligent agent cluster at the end of the next time step, and to determine the coverage of the intelligent agent cluster of the task space based on the second set of location information. The reward module is used to determine the reward score for the agent cluster to perform this cluster coverage task based on the second location information set and the coverage rate; The training module is used to construct the state matrix for the next time step based on the second location information set, and store the reward score, the behavior decision information, the first state matrix and the second state matrix as a set of data into a buffer, and control the agent to continue to execute the coverage task until the amount of data in the buffer is equal to the preset batch processing data amount, and then train the model. The advantage value calculation module is used to calculate the advantage value of the current training session based on the state matrix of the current time step, the state matrix of the next time step, and a preset advantage function during each training session. The network update module is used to calculate the loss value of this training based on the state matrix of the current time step, the behavior decision information, and the advantage value, and update the actor network and the critic network according to the loss value; The termination module is used to continue training with the updated actor network and critic network until the preset termination condition is met, at which point the training ends and the cluster coverage search model is obtained.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the training method of the cluster coverage search model as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or instruction set is loaded and executed by a processor to implement the training method of the cluster coverage search model as described in any one of claims 1-7.

Citation Information

Cited By

  • Unmanned aerial vehicle cluster collaborative search method and system based on multi-agent reinforcement learning

    CN121742524A