Dynamic optimization and real-time decision method and device for multi-agent cooperative target search

Through the collaborative optimization of the global optimization layer and the local reinforcement learning layer, combined with the environmental perception mechanism, the coordination of global strategy and local execution in the multi-agent collaborative target search task is achieved, the efficiency and completion of task execution are improved, and the problem of inefficiency in path planning and task allocation in complex dynamic environments is solved.

CN119828460BActive Publication Date: 2025-10-24XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411892930.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-10-24
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

In complex dynamic environments, in multi-agent collaborative target search tasks, existing technologies have difficulty achieving global information integration, rapid decision-making and coordination, resulting in inefficient path planning and task allocation, and unable to effectively respond to real-time changes in target positions and obstacles.

Method used

The global optimization level is used to generate a global model, and the lightweight local reinforcement learning level is combined with reinforcement learning to make real-time decisions. The planning and resource allocation are dynamically updated through environmental perception data to ensure the implementation of path planning and task allocation plans. The decision-making is combined with the dynamic environmental perception data to realize the implementation of path planning and task allocation plans. Combined with real-time decision-making, the dynamic planning and execution strategy of path planning and task allocation is realized.

Benefits of technology

By generating a global model at the global optimization level and combining it with environmental perception data to dynamically update the planning and resource allocation plan, the implementation of path planning and task allocation plans is realized. The implementation of path planning and task allocation plans is realized, and the implementation of environmental dynamic planning and resource allocation plans is combined to ensure the coordination of global strategy and local execution, avoiding global failure caused by local decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119828460B_ABST
    Figure CN119828460B_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic optimization and real-time decision method and device of multi-agent collaborative target search, comprising: global optimization layer constructs global model, obtains the initial path planning of multiple intelligent agents and initial task allocation scheme, sends to each intelligent agent;While, global optimization layer updates global model with first preset frequency;Each intelligent agent adjusts initial path planning according to initial path planning and initial task allocation scheme, obtains execution result, and sends execution result to global optimization layer;While, each intelligent agent updates execution result with second preset frequency;Global optimization layer judges whether global model needs to be updated according to the execution result of each intelligent agent, if needs, then update global model, send to each intelligent agent, if not, save the execution result of each intelligent agent;Real-time acquisition path planning and task allocation scheme.The application can improve task execution efficiency and completion degree.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent optimization and automatic decision-making, and particularly relates to a dynamic optimization and real-time decision-making method and device for multi-agent cooperative target search. BACKGROUND

[0002] In a target search scenario, a controller needs to control multiple unmanned aerial vehicles (UAVs) to perform a target search task. The target search scenario is full of unknown terrains, obstacles and dynamically changing environments. The main task of the controller is to search for and collect hidden targets in a challenging space through the cooperation of the UAVs. These targets can be distributed in a wide and complex environment, and the environment itself changes over time, the positions of the targets change over time, or the appearance of new obstacles hinders the search path of the UAVs, increasing the difficulty of the task and greatly affecting the search strategy and execution efficiency of the UAVs. Therefore, the UAVs need to constantly adjust their search strategies to adapt to the dynamically changing environment.

[0003] In a highly dynamic task environment, the target search of the UAVs is not only a simple path planning problem, but also requires each UAV to make decisions quickly in a complex environment. How to reasonably allocate resources, select the optimal path, and respond to the changing target positions and external disturbances is the key to successfully completing the task. In addition, the controller not only needs to control the individual actions of the UAVs, but also needs to ensure the coordination and cooperation between the UAVs to avoid repeated search or missed targets, so as to complete the search task within a limited time. At the same time, as the task progresses, the controller will gradually face more and more complex situations, which requires the controller not only to have a global strategic vision, but also to quickly adjust the search strategy according to the local environment. Under such a background, every decision, every path planning and every search action of the controller directly affects the completion efficiency and final result of the task.

[0004] Therefore, it is urgent to provide a dynamic optimization and real-time decision-making method for multi-agent cooperative target search to improve the search efficiency of targets in a dynamic search environment. SUMMARY

[0005] In order to solve the above problems existing in the prior art, the application provides a dynamic optimization and real-time decision-making method and device for multi-agent cooperative target search. The technical problem to be solved by the application is solved by the following technical scheme:

[0006] In a first aspect, the application provides a dynamic optimization and real-time decision-making method for multi-agent cooperative target search, comprising:

[0007] The global optimization layer constructs a global model according to the target position and obstacle distribution of the search scene, obtains initial path planning and initial task allocation scheme of the multiple agents, and sends the initial path planning and the initial task allocation scheme to each agent; meanwhile, the global optimization layer updates the global model at a first preset frequency, and updates the initial path planning and the initial task allocation scheme of the multiple agents;

[0008] Each agent adjusts the initial path planning to execute a task according to the corresponding initial path planning and initial task allocation scheme, combines real-time perception data, obtains an execution result, and sends the execution result to the global optimization layer; meanwhile, each agent adjusts the initial path planning to execute a task at a second preset frequency, and updates the execution result;

[0009] The global optimization layer judges whether the initial path planning and the initial task allocation scheme of the multiple agents need to be updated according to the execution result of each agent, if yes, updates the global model, updates the initial path planning and the initial task allocation scheme of the multiple agents, and sends the initial path planning and the initial task allocation scheme to each agent, and if no, saves the execution result of each agent;

[0010] The global optimization layer and each agent are kept in cooperative updating, and the path planning and the task allocation scheme are obtained in real time.

[0011] In a second aspect, the present application further provides a dynamic optimization and real-time decision device for multi-agent cooperative target search, which is used to realize the multi-agent cooperative target search dynamic optimization and real-time decision method provided above, and includes:

[0012] A global optimization module is configured to construct a global model according to the target position and obstacle distribution of the search scene, obtain initial path planning and initial task allocation scheme of the multiple agents, and send the initial path planning and the initial task allocation scheme to each agent; meanwhile, the global optimization layer updates the global model at a first preset frequency, and updates the initial path planning and the initial task allocation scheme of the multiple agents;

[0013] A lightweight local reinforcement learning module is configured to adjust the initial path planning to execute a task according to the corresponding initial path planning and initial task allocation scheme, combine real-time perception data, obtain an execution result, and send the execution result to the global optimization layer; meanwhile, each agent adjusts the initial path planning to execute a task at a second preset frequency, and updates the execution result;

[0014] An information feedback module is configured to judge whether the initial path planning and the initial task allocation scheme of the multiple agents need to be updated according to the execution result of each agent, if yes, update the global model, update the initial path planning and the initial task allocation scheme of the multiple agents, and send the initial path planning and the initial task allocation scheme to each agent, and if no, save the execution result of each agent;

[0015] A multi-layer time scale optimization module is used to keep the global optimization layer and the individual agent cooperative update, and obtain the path planning and task allocation scheme in real time.

[0016] The beneficial effects of the present application are:

[0017] The dynamic optimization and real-time decision method and device for multi-agent cooperative target search provided by the present application ensure the coordination of global strategy and local search target execution by performing overall task planning and resource allocation at the global optimization level and using reinforcement learning for real-time decision making at the local reinforcement learning level, avoiding global failure that may be caused by local decision making. The division of labor and cooperation between the global optimization layer and the local reinforcement learning layer enables the system to flexibly adjust task guidance while responding to changes in local execution, thereby improving task execution efficiency and completion degree.

[0018] The present application will be further described in detail below in conjunction with the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is a flowchart of the dynamic optimization and real-time decision method for multi-agent cooperative target search provided by the present application;

[0020] Figure 2 is a schematic diagram of the dynamic optimization and real-time decision method for multi-agent cooperative target search provided by the present application;

[0021] Figure 3 is a schematic diagram of the dynamic optimization and real-time decision device for multi-agent cooperative target search provided by the present application. DETAILED DESCRIPTION

[0022] The present application will be further described in detail below in conjunction with the accompanying drawings and examples, but the embodiments of the present application are not limited thereto.

[0023] In the prior art, the multi-agent cooperative target search problem involves multiple agents (such as unmanned aerial vehicles, robots, etc.) working cooperatively to efficiently search for targets or cover areas, mainly relying on task allocation, path planning, collaboration mechanisms and dynamic environment adaptation methods to optimize efficiency and accuracy; existing target search methods include:

[0024] Swarm Intelligence and distributed cooperative algorithms are classical methods for solving multi-agent target search problems. Common algorithms include Ant Colony Optimization (ACO), which simulates the foraging process of ants and achieves target search through local information exchange. In multi-agent cooperative search tasks, ACO can optimize path planning and task allocation by simulating individual local search behavior and information sharing. Particle Swarm Optimization (PSO) also includes simulating the flight behavior of a particle swarm in the search space to find the optimal solution. PSO can combine local updates with global information to perform path planning in a multi-agent system that adapts to complex dynamic environments. Bee Colony Algorithm (BCA) is similar to ACO and simulates the way bees find food. Through collaboration and information sharing among multiple agents, BCA improves search efficiency.

[0025] Reinforcement Learning (RL) is widely used in multi-agent systems, especially in path planning and task allocation in dynamic environments. Common algorithms include Q-learning, which is based on Q-value updates. Agents learn the optimal strategy in the environment through exploration and trial-and-error. In multi-agent cooperative tasks, Q-learning can achieve autonomous decision-making for individual agents and adjust behavior through a reward and punishment mechanism. Deep Q Network (DQN) combines deep learning and Q-learning. DQN enables reinforcement learning to handle high-dimensional state spaces. In multi-agent systems, DQN can support more complex state and action spaces, allowing agents to adjust their behavior in real-time in dynamic environments. Multi-Agent Reinforcement Learning (MARL) further extends the idea of single-agent RL to handle multiple agents working together. In target search, MARL learns local strategies for each agent and combines global information to optimize overall search efficiency.

[0026] Distributed task allocation and coordination control techniques propose some optimization solutions for task allocation and coordination in multi-agent systems. Common methods include auction algorithms, in which tasks are auctioned as items, and agents select tasks through bidding. Through the bidding process, agents can reasonably allocate tasks and ensure the rapid execution of tasks. Distributed optimal control is another method that allocates tasks and coordinates behavior among multiple agents by optimizing a global objective function. Each agent independently executes local tasks while considering collaboration with other agents to optimize the overall goal.

[0027] In complex dynamic environments, the positions of targets and obstacles are constantly changing. Real-time path planning and task allocation in such environments is a significant challenge. Common techniques include adaptive path planning, which dynamically adjusts paths based on real-time environmental changes such as target movement and obstacle appearance. Common algorithms include A* and D* algorithms, which can dynamically adjust paths through environment map updates. Game theory provides a theoretical framework for multi-agent collaboration by simulating decision-making games between agents to achieve efficient task allocation and collaboration. In target search, game theory can help agents coordinate actions in complex dynamic environments to maximize overall efficiency.

[0028] In complex dynamic environments, the positions of targets and obstacles are constantly changing. Real-time path planning and task allocation in such environments is a significant challenge. Common techniques include adaptive path planning, which dynamically adjusts paths based on real-time environmental changes such as target movement and obstacle appearance. Common algorithms include A* and D* algorithms, which can dynamically adjust paths through environment map updates. Game theory provides a theoretical framework for multi-agent collaboration by simulating decision-making games between agents to achieve efficient task allocation and collaboration. In target search, game theory can help agents coordinate actions in complex dynamic environments to maximize overall efficiency.

[0029] In summary, the existing methods have the following defects:

[0030] Existing swarm intelligence and distributed collaboration algorithms (such as ant colony algorithm, particle swarm optimization algorithm, etc.) mainly rely on local information and adaptive mechanisms for path planning and task allocation during target search. Although these algorithms can optimize search paths to some extent, they often lack effective global information integration mechanisms, leading to local optimal solutions in complex dynamic environments and failing to guarantee optimal search results for global targets.

[0031] When reinforcement learning (such as Q-learning, deep Q-network, etc.) is applied to target search tasks, although it can achieve autonomous decision-making and path selection of agents, its learning process is usually slow and difficult to meet real-time requirements. Especially in multi-agent cooperative search tasks, reinforcement learning lacks sufficient rapid adaptation ability and is difficult to respond to environmental dynamics such as target changes and obstacle appearance in real time, which limits the decision-making efficiency and search accuracy of agents.

[0032] In multi-agent target search tasks, distributed task allocation and cooperative control techniques (such as auction algorithms, distributed optimal control, etc.) are usually used for task allocation among agents. However, these methods lack efficient cooperation and coordination capabilities, especially when the task demand is complex or the number of agents is large, it is difficult to efficiently handle task allocation and cooperation. The flexibility of traditional algorithms in dynamic environments is poor, which leads to task allocation often unable to adjust in time according to environmental changes, thereby affecting search efficiency and target discovery speed.

[0033] Although existing environmental perception and dynamic planning techniques (such as adaptive path planning, game theory, etc.) can respond to certain environmental changes, in target search tasks, these methods often lack sufficient real-time response capabilities. Especially in the case of frequent changes in target location and obstacles, the path planning and task allocation effect of these algorithms may become unstable, unable to quickly adjust the search strategy, leading to reduced search efficiency and increased risk of missing targets.

[0034] Therefore, the present application provides a dynamic optimization and real-time decision-making method for multi-agent cooperative target search, which combines global optimization and local reinforcement learning to improve the accuracy and efficiency of task completion.

[0035] In the present application, the idea of combining global and local is used to improve the search efficiency of the task.

[0036] 1. The global optimization layer generates global path planning and task allocation scheme, and responds to environmental dynamic changes to update the global path planning and task allocation scheme.

[0037] The global optimization layer is responsible for integrating all information in the task environment, including target location, obstacle distribution, etc., to generate initial path planning and initial task allocation scheme. The global optimization layer is equivalent to the perspective of the "command department" and has the ability to control global information. Its main task is to update the path planning and task allocation scheme at fixed time periods or when major environmental dynamics occur, ensuring the rationality of task allocation and the efficiency of path planning, and providing strategic guidance for local execution.

[0038] 2. The local reinforcement learning layer dynamically adjusts the behavior of agents executing tasks to generate actual execution results of agents.

[0039] The local reinforcement learning layer adjusts the path selection and action strategy of each agent according to the planning guidance provided by the global optimization layer combined with real-time perception data. The local reinforcement learning layer is equivalent to the perspective of a "pilot", and is characterized by lightweight design, which can only obtain environmental information within a limited range, and make rapid responses within the local range through reinforcement learning algorithms. Through the reward mechanism, the system can promote the agent to efficiently search for the target and avoid obstacles in the dynamic environment, and enhance the flexibility and adaptability of execution.

[0040] It should be noted that the agent of the present application is a UAV.

[0041] 3. Bidirectional information feedback mechanism to coordinate global planning and local execution.

[0042] The bidirectional information feedback mechanism is the coordination link between the global optimization layer and the local reinforcement learning layer. The data generated during local execution (such as task progress, path deviation, etc.) will be fed back to the global optimization layer in real time, which will adjust the global planning scheme accordingly, and issue new task guidance to the local execution layer, thereby realizing dynamic collaborative optimization of planning and execution.

[0043] 4. Multi-layer time scale optimization to distinguish between global optimization and local optimization.

[0044] The multi-layer time scale optimization module ensures that the global optimization layer and the local reinforcement learning layer work in parallel on different time scales. The global optimization layer updates the planning scheme in a longer time interval, i.e. updates according to the first preset frequency, to ensure the coordination and resource optimization of the overall strategy; the local reinforcement learning layer responds to local environmental changes at a second frequency, i.e. updates according to the second preset frequency, to ensure the flexibility and execution efficiency of tactical decision-making, thereby maintaining the coordination between global and local.

[0045] 5. Environmental dynamic perception mechanism to trigger global planning update and optimize local execution.

[0046] The environmental dynamic perception mechanism is responsible for real-time monitoring of changes in target position and obstacles. When the system perceives significant environmental changes, dynamic perception data will be passed to the local reinforcement learning layer to trigger individual behavior adjustment, ensuring that global and local planning and execution always maintain the best collaborative state.

[0047] In the present application, it is assumed that the search environment is a two-dimensional planar region In this airspace, there are multiple targets and obstacles, and the position and size of the obstacles will affect the path selection of the UAV, and the state of the targets and obstacles is dynamically changing, which will be adjusted or added over time.

[0048] Please refer to Figure 1 and Figure 2 ,Figure 1 This is a flow chart of a dynamic optimization and real-time decision-making method for multi-agent collaborative target search provided by an embodiment of the present invention. Figure 2 : This is a schematic diagram of a dynamic optimization and real-time decision-making method for multi-agent collaborative target search provided by an embodiment of the present invention. The dynamic optimization and real-time decision-making method for multi-agent collaborative target search provided by the present invention includes:

[0049] S101. Based on the target position and obstacle distribution of the search scene, the global optimization layer constructs a global model, obtains the initial path planning and initial task allocation plan of the multi-agent, and sends it to each agent; at the same time, the global optimization layer updates the global model at a first preset frequency, and updates the initial path planning and initial task allocation plan of the multi-agent.

[0050] Specifically, in this embodiment, based on the target position and obstacle distribution of the search scene, the global optimization layer constructs a global model, including:

[0051] Obtain decision variables based on the target location and obstacle distribution of the search scene; the decision variables include path sequence and task allocation matrix;

[0052] The path sequence refers to the path sequence of each agent, including the target point, starting point and end point that the agent passes through. The path sequence P i The expression is:

[0053]

[0054] The task assignment matrix shows whether each goal is assigned to a certain agent. The expression of the task assignment matrix A is:

[0055] A={A ij};

[0056]

[0057] Among them, p ik represents the kth path point of the i-th agent, K i represents the total number of path points of the i-th agent, i = 1, 2, ..., N represents the index of the agent, N represents the total number of agents, j = 1, 2, ..., M represents the target index, M represents the total number of targets;

[0058] Construct a global model, that is, the objective function J global , the global optimization goal is to minimize the path length, maximize the target coverage, and reduce the threat impact, and its expression is:

[0059] J global =α·L path +β·T threat- g-coverage;

[0060] wherein L path represents the total length of the multi-agent path, a represents a first weight, T threat represents the total threat intensity of the multi-agent, b represents a second weight, coverage represents the coverage efficiency of the task, and g represents a third weight;

[0061] Obtain the path constraint of each agent, each agent starts from a starting path point, passes through a series of path points, and reaches a final path point, and the expression is:

[0062]

[0063] wherein each path point p ik is a coordinate point in a plane R 2 , S i represents the starting path point, and E i represents the final path point.

[0064] Obtain the task assignment constraint, each target is searched by an agent, and each agent can be responsible for searching multiple targets, and the expression is:

[0065]

[0066] Obtain the obstacle constraint, the path of the agent must avoid all obstacles, and the distance between each path point and any obstacle is greater than the radius of the obstacle, and the expression is:

[0067]

[0068] wherein the area of the obstacle region is circular, the center is O k , and the radius is r k .

[0069] In this embodiment, for each agent i, the expression of the path length L i is:

[0070]

[0071] wherein d(p ik , p i(k+1) ) represents the distance between the path point p ik and the path point p i(k+1) , and the expression of the total length L path of the multi-agent path is:

[0072]

[0073] The total threat intensity T threatThe expression of coverage is:

[0074]

[0075] Wherein, r represents the index of the radar, R represents the total number of radars, p i represents the position of the i-th agent, R max represents the maximum detection range of the radar, f threat (p i ,r) represents the threat intensity between the agent i and the radar r, d(p i ,r) represents the distance between the agent i and the radar r; the multi-agent threat comes from the radar;

[0076] The expression of coverage of the task is:

[0077]

[0078] Wherein, cover(j) represents whether the target j is searched, if the target j is searched by at least one agent, cover(j) = 1, otherwise 0, and the coverage of the task coverage represents the degree of the task being searched.

[0079] In this embodiment, the particle swarm optimization algorithm is used to solve the above constructed objective function, to obtain the initial path planning of the multi-agent and the initial task allocation scheme, including:

[0080] Initialize the particle swarm, and initialize the position and speed of the particle to a random value; wherein each particle represents a solution, including an agent path sequence and a task classification matrix, and the dimension of each particle includes the path planning of the multi-agent and the task allocation scheme;

[0081] The fitness value of each particle is calculated using the objective function; wherein the smaller the fitness value of the particle, the better the solution.

[0082] According to the experience of the particle itself and the experience of the group, the speed and position of the particle are updated, and the expression is:

[0083]

[0084] x u (t+1)=x u (t)+v u (t+1);

[0085] Wherein, v u (t) represents the speed of the particle, x u (t) represents the position of the ion, ω represents the inertia weight, c1 and c2 represent the learning factor, rand1 and rand2 represent random numbers, pbest urepresents the historical best position of particle u, gbest represents the global optimal position, and t represents the number of iterations;

[0086] According to the particle fitness value, update the historical optimal position pbest of each particle u and the global optimal position gbest;

[0087] When the preset number of iterations is reached or the fitness value of the particle is less than the threshold, the optimization process is stopped and the initial path planning and initial task allocation plan of the multi-agent are obtained.

[0088] It should be noted that the final global optimization result is the optimal path sequence and task allocation plan. These results are output by the global optimization layer and passed to the local reinforcement learning layer. The local reinforcement learning layer adjusts the paths of individual drones based on these plans and executes the tasks based on real-time perception data.

[0089] S102. Each intelligent agent adjusts the initial path planning to perform the task based on the corresponding initial path planning and initial task allocation plan, combined with real-time perception data, obtains the execution result, and sends the execution result to the global optimization layer; at the same time, each intelligent agent adjusts the initial path planning to perform the task at a second preset frequency and updates the execution result.

[0090] Specifically, in this embodiment, each agent adjusts the initial path plan to execute the task based on the corresponding initial path plan and initial task allocation plan, combined with real-time perception data, to obtain the execution results, including:

[0091] The i-th agent is initialized according to the received corresponding initial path planning and initial task allocation scheme, and sets the reward function; wherein, the position of the agent and the task allocation are used as the initial state of reinforcement learning, which includes the current position of the agent, the position and allocation scheme of the target, and the search state of the task, such as the search progress; the agent generates the initial action according to the corresponding initial path planning, and the action can be based on the adjustment of the path points, such as path points p1, p2, ..., p n It can be used as part of the initial action space for the agent to adjust in subsequent training;

[0092] The i-th agent uses the initial state and initial action as the starting path point and executes the task according to the initial path plan;

[0093] The i-th agent combines the real-time perception data and the k-1th path point p in the initial path planning ik-1 The Q value for the kth path point p ik , adjust the current state s of the agent i, select a current action a matching the current state i , as the kth path point, and obtain a reward; meanwhile, obtain the Q value of the kth path point p ik , expressed as:

[0094] Q(s i ,a i ) = Q(s i ,a i ) + a[r i + gmax a′ Q(s i+1 ,a') - Q(s i ,a i )];

[0095] wherein Q(s i ,a i ) represents the current Q value of the current action a i under the current state s i , r i represents the reward after executing the current action a i , g represents the discount factor, a represents the learning rate, max a′ Q(s i+1 ,a') represents the maximum Q value of all possible actions under the state s i+1 , s i represents the current state of the ith intelligent agent, and s i+1 represents the current state of the (i+1)th intelligent agent.

[0096] When the initial task allocation scheme is executed by the ith intelligent agent, the initial path planning is updated to obtain an execution result; wherein the execution result includes a path actual execution result, threat avoidance information, and environmental dynamic changes.

[0097] It should be noted that the path actual execution result includes the fine-tuning of the reinforcement learning intelligent agent to the path during the execution process (for example, bypassing obstacles, adjusting flight height, etc.).

[0098] The threat avoidance information includes the avoidance strategy of the reinforcement learning intelligent agent to the threat source during the execution process.

[0099] The environmental dynamic changes include that new obstacles, targets or threats may appear during the execution process, and the reinforcement learning layer needs to update its behavior through real-time perception. These environmental changes are passed to the global optimization layer to help the global optimization layer dynamically adjust the planning.

[0100] It can be understood that the reinforcement learning intelligent agent will select a behavior based on the current state and action space. Reinforcement learning does not select an action "from scratch", but will use the preliminary decision provided by the global optimization layer.

[0101] The specific process is as follows:

[0102] 1. Using global planning information, the reinforcement learning layer will first select the action (path adjustment, target selection, etc.) that best matches the current state based on the output of the global optimization layer. For example, if the global optimization layer suggests that the agent travel along a certain path, the reinforcement learning agent will choose this path as the initial action plan.

[0103] 2. Exploration and optimization, the reinforcement learning agent adjusts its behavior through the exploration process, gradually optimizing the path. Although the global optimization layer provides preliminary path planning, the agent will adjust the path according to changes in the environment (such as the appearance of obstacles or threats) to achieve the optimal task execution effect.

[0104] In this embodiment, the Q-learning algorithm is used for solving, Q-Learning is a value iteration-based reinforcement learning algorithm used to learn the optimal policy in a given state-action space. The core of reinforcement learning is to update the Q value of each state-action pair to estimate the value of the agent taking a certain action in a certain state. The reinforcement learning agent will perceive changes in the environment in real time and adjust its behavior when performing tasks. As the task progresses, new obstacles or threats may appear on the path, and the agent needs to perceive these changes in time and adjust the path. Reinforcement learning also requires the agent to judge whether to continue executing the current target or switch to the next target based on the task progress (e.g., target coverage).

[0105] In this embodiment, the reward function is a key in reinforcement learning, which defines the reward value after performing a certain action, the purpose is to guide the agent to learn the optimal policy, the reward function includes:

[0106] Task completion reward, positive reward is given when the agent completes the search task;

[0107] Threat avoidance reward, positive reward is given when the agent successfully avoids the threat source;

[0108] Path adjustment reward, positive reward is given if the agent's path length is less than the preset path length or avoids obstacles; it can be understood that the smaller the agent's path length, the better;

[0109] Collision penalty, negative reward is given if the agent collides with obstacles or enters a high-threat area.

[0110] In this embodiment, the process of acquiring real-time perception data includes:

[0111] Real-time perception data is acquired through the environmental perception module set on the agent.

[0112] In this embodiment, the environment perception module plays a crucial role in the dynamic triggering process, continuously monitoring the entire task environment, and real-time perception and judgment of significant changes. These changes may come from the dynamic changes of target position, obstacles, threat sources and other factors, or may be caused by some sudden situations during task execution. When these changes are perceived, the system will trigger the corresponding response mechanism according to the type and influence range of the changes.

[0113] During the operation of the whole process, the dynamic triggering mechanism runs through the information flow and coordination between the global optimization layer and the local reinforcement learning layer. When the environment changes, the local reinforcement learning layer needs to respond quickly and adjust the strategy, while feeding back the updated execution results to the global optimization layer to provide the latest information for subsequent task planning. This mechanism ensures the continuous coordination of global strategy and local tactics, enabling the system to flexibly respond to complex and changing task environments, maintaining the continuity and adaptability of the task.

[0114] S103, the global optimization layer determines whether to update the initial path planning and initial task allocation scheme of the multi-agent according to the execution results of each agent, if necessary, updates the global model, updates the initial path planning and initial task allocation scheme of the multi-agent, and sends to each agent, if not, saves the execution results of each agent.

[0115] Specifically, in this embodiment, it specifically includes:

[0116] The global optimization layer analyzes the received execution results of each agent, compares the initial path planning of each agent with the actual path execution results, and evaluates the effectiveness of the actual path execution results; evaluates the distribution of threat sources in the threat avoidance information of the agent; evaluates the environmental changes in the environmental dynamic transformation of each agent; obtains the evaluation results of the execution results of each agent;

[0117] It can be understood that the evaluation of the path execution includes comparing the path of the global optimization planning with the actual path after the execution of the reinforcement learning layer to evaluate the effectiveness of the execution path. For example, if some paths deviate or collide with threat sources, it indicates that some path designs in the global planning are unreasonable, and it may be necessary to re-optimize. The evaluation of the distribution of threat sources includes evaluating the distribution of threat sources according to the threat avoidance information provided by the reinforcement learning layer. For example, if the threat in a certain area is large, it may be necessary to adjust the path planning or change the task allocation strategy; the global optimization layer will comprehensively evaluate the threat source information and other task constraints to update the threat avoidance strategy. The evaluation of the environmental changes in the environmental dynamic changes includes that the dynamic changes in the environment (such as the appearance of new obstacles or targets) will affect the global task planning; the global optimization layer will recompute and adjust the task allocation according to these changes. For example, if the reinforcement learning layer reports that the position of a target changes or a new obstacle is added, the global optimization layer may re-optimize the path to ensure that the task continues to be effectively performed.

[0118] According to the evaluation results, it is judged whether the initial path planning and the initial task allocation scheme of the multi-agent need to be updated.

[0119] If needed, the global optimization layer updates the global model in combination with the actual execution results of the paths of each agent, threat information and environmental dynamic changes, and updates the initial path planning and the initial task allocation scheme of the multi-agent.

[0120] It can be understood that updating the initial path planning includes that the global optimization layer will re-evaluate the flight path of the agent. By combining the threat avoidance information, the task progress and the environmental changes, the flight path planning of each agent is updated. Specifically, if there is an obvious obstacle or threat in a certain path, the obstacle avoidance path is recalculated. If the task progress is slow, the target is re-allocated, the path is adjusted, or the task participation of the agent is increased.

[0121] Updating the initial task allocation scheme includes that if the reinforcement learning layer finds that the search efficiency of some targets is low or the task completion progress in some areas is fast, the global optimization layer may re-allocate the tasks. For example, if the target of an agent has been completely covered, it can be re-allocated to a new target.

[0122] If not needed, the global optimization layer saves the execution results of each agent.

[0123] S104, the global optimization layer and each agent are kept updated to obtain the path planning and the task allocation scheme in real time.

[0124] Specifically, in the embodiment, the first preset frequency is less than the second preset frequency.

[0125] In this embodiment, the global optimization layer and the local reinforcement learning layer work in parallel and optimize decisions at their respective time scales to achieve high coordination of global goal search and local execution. This process involves division of labor and cooperation between global and local, ensuring that the system can maintain the unity of global strategy while flexibly responding to immediate changes in the local environment when facing complex dynamic environments.

[0126] Specifically, the cooperative updating of the global optimization layer and the local optimization layer includes the following beneficial effects:

[0127] 1. Time scale difference;

[0128] The global optimization layer is responsible for long-term task planning and strategic decision-making on a large scale, usually involving a long time span, such as overall task planning, long-distance flight route of unmanned aerial vehicles, target allocation, etc. The optimization decisions at this level are relatively stable, and the main goal is to ensure the optimization of the global goal of the task. Due to the large amount of calculation of global optimization, it will be updated periodically. For example, the first preset frequency can be to adjust the task once every certain time (such as every 10 minutes or every hour) for global optimization to ensure that the task can adapt to the long-term changing environment (such as target migration, task priority change, etc.).

[0129] The local reinforcement learning layer makes decisions and optimizations on a shorter time scale, focusing on real-time target search decisions, such as immediate path selection for individual agents, obstacle avoidance, threat response, etc. Each local decision is usually based on the current state and environmental perception to make an optimized choice, and the cycle is usually shorter, such as every second or every minute. The reinforcement learning layer continuously optimizes behavior strategies through interaction with the environment to ensure that the agent can make flexible responses in a dynamic environment and adjust path planning and task execution in real time.

[0130] 2. Optimization goals and synergistic effects;

[0131] The cooperation of the global optimization layer and the reinforcement learning layer also needs to be unified in terms of optimization goals, ensuring that the goals of the two are consistent and complementary. The global optimization goal is path optimization, target coverage maximization, and threat avoidance. The goal of the reinforcement learning layer is to quickly make the best tactical decision based on the current local environment and task requirements, such as selecting the appropriate search area or path.

[0132] 3. Real-time adjustment and execution;

[0133] The reinforcement learning layer makes a decision every certain period of time, adjusts the actions of the agent, and feeds back the execution to the global optimization layer in real time. The global optimization layer will update the task planning based on this information and form new task guidelines. The whole process forms a closed loop, ensuring the close cooperation of global strategy and local execution.

[0134] In summary, the present application provides a dynamic optimization and real-time decision-making method for multi-agent cooperative target search, which has the following beneficial effects:

[0135] 1. Global and local cooperative optimization. The system plans the overall task and allocates resources at the global optimization level, while using reinforcement learning at the local reinforcement learning level for real-time decision-making, ensuring the coordination of global strategy and local search target execution, and avoiding global failure caused by local decision-making. The division of labor between the global optimization layer and the local reinforcement learning layer allows the system to flexibly adjust task guidance while responding to changes in local execution, thereby improving task execution efficiency and completion.

[0136] 2. Real-time dynamic adaptability. By introducing an environment perception module and combining it with a dynamic triggering mechanism, the system can perceive environmental changes in real time and automatically adjust planning and execution strategies. This dynamic adaptability effectively deals with uncertainties and sudden changes in the task environment, ensuring that tasks can continue in complex and dynamic scenarios, improving the reliability and adaptability of task execution.

[0137] 3. Improve task execution efficiency. The local reinforcement learning layer can adjust the behavior of individual agents in real time based on real-time feedback, optimize local task execution strategies, and reduce invalid paths and time waste. At the same time, the sequential decision-making mechanism of reinforcement learning can optimize the behavior of agents in complex environments, making task execution more efficient.

[0138] 4. Enhance the flexibility and robustness of the system. The system can handle various environmental and task requirements, ensuring that even in the face of drastic environmental changes, agents can still complete tasks efficiently based on global planning and local strategies. In addition, the parallel optimization feature of the system allows different levels of decision-making to be executed in parallel at different time scales, better handling multiple objectives and various constraints in tasks.

[0139] 5. Promote efficient coordination between agents. Through the combination of global planning and local reinforcement learning, agents can not only optimize their own task execution, but also coordinate and cooperate with other agents, fully utilizing the advantages of multi-agent systems to achieve complex goals. This collaboration mechanism enhances the overall ability of the system, especially in the face of complex and variable tasks, improving the execution of overall strategies.

[0140] Based on the same inventive concept, please see Figure 3 , Figure 3is a schematic view of the dynamic optimization and real-time decision device for multi-agent cooperative target search provided by the embodiment of the present application, and the present application further provides a dynamic optimization and real-time decision device for multi-agent cooperative target search, which is used to realize the dynamic optimization and real-time decision method for multi-agent cooperative target search provided by the above-mentioned embodiment of the present application, and the embodiment of the method is described above, and will not be described here again; the device comprises:

[0141] a global optimization module, configured to construct a global model according to the target position and obstacle distribution of the search scene, obtain an initial path planning and an initial task allocation scheme of the multi-agent, and send them to each agent; meanwhile, the global optimization layer updates the global model at a first preset frequency, and updates the initial path planning and the initial task allocation scheme of the multi-agent;

[0142] a lightweight local reinforcement learning module, configured to adjust the initial path planning to execute the task according to the corresponding initial path planning and initial task allocation scheme, combine real-time perception data, obtain an execution result, and send the execution result to the global optimization layer; meanwhile, each agent adjusts the initial path planning to execute the task at a second preset frequency, and updates the execution result;

[0143] an information feedback module, configured to judge whether the initial path planning and the initial task allocation scheme of the multi-agent need to be updated according to the execution result of each agent, if yes, update the global model, update the initial path planning and the initial task allocation scheme of the multi-agent, and send them to each agent, and if not, save the execution result of each agent;

[0144] a multi-layer time scale optimization module, configured to keep the global optimization layer and each agent cooperative updating, and obtain the path planning and the task allocation scheme in real time.

[0145] It is to be understood that the terminology "first", "second", and the like used herein is merely intended to differentiate one element from another element, and does not imply or suggest any actual relationship or sequence between the elements. Also, the terms "comprise", "comprising", or any other variant thereof are intended to cover non-exclusive inclusions, such that an item or apparatus that comprises a list of elements does not exclude other elements not expressly listed. An element defined by the phrase "comprising a... " does not exclude the presence of additional identical elements in the item or apparatus comprising the element. The terms "connected", "coupled", and the like are not limited to direct connections or physical connections, but can include indirect connections or indirect physical connections between devices or elements such as electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are used only for the purpose of facilitating the description of the present application and simplifying the description, and thus cannot be construed as indicating or implying that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and thus cannot be construed as limiting the present application.

[0146] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific feature or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. The illustrative expressions of the above terms in the present specification do not necessarily refer to the same embodiment or example. Also, the specific feature or characteristic described can be combined in any suitable manner in one or more embodiments or examples. In addition, a person skilled in the art can combine and combine different embodiments or examples described in the present specification.

[0147] The above is a further detailed description of the present application in conjunction with specific preferred embodiments, and cannot be considered as limiting the specific implementation of the present application to these descriptions. For those of ordinary skill in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or substitutions can be made, which should be considered as falling within the scope of protection of the present application.

Claims

1. A dynamic optimization and real-time decision method for multi-agent cooperative target search, characterized in that, Comprise: According to the target position and obstacle distribution of the search scene, the global optimization layer constructs a global model, obtains the initial path planning and initial task allocation scheme of the multi-agent, and sends them to each agent; at the same time, the global optimization layer updates the global model at a first preset frequency, and updates the initial path planning and initial task allocation scheme of the multi-agent; According to the corresponding initial path planning and initial task allocation scheme, each agent adjusts the initial path planning to execute the task in combination with real-time perception data, obtains an execution result, and sends the execution result to the global optimization layer; at the same time, each agent adjusts the initial path planning to execute the task at a second preset frequency, and updates the execution result; According to the execution result of each agent, the global optimization layer judges whether the initial path planning and initial task allocation scheme of the multi-agent need to be updated, if yes, updates the global model, updates the initial path planning and initial task allocation scheme of the multi-agent, and sends them to each agent, and if not, saves the execution result of each agent; The global optimization layer and each agent are kept in cooperative updating, and the path planning and task allocation scheme are obtained in real time.

2. The method of claim 1, wherein, According to the target position and obstacle distribution of the search scene, the global optimization layer constructs a global model, comprising: According to the target position and obstacle distribution of the search scene, a decision variable is obtained; the decision variable comprises a path sequence and a task allocation matrix; The path sequence P i The expression is: The expression of the task allocation matrix A is: A = {A ij}; where p ik represents the kth waypoint of the ith agent, K i represents the total number of waypoints of the ith agent, i = 1, 2, …, N represents the index of the agent, N represents the total number of agents, j = 1, 2, …, M represents the index of the target, and M represents the total number of targets; Constructing the global model, i.e. the objective function J global whose expression is: J global = a · L path + b · T threat - g · coverage; wherein, L path represents the total length of the multi-agent path, a represents a first weight, T threat represents the total threat intensity of the multi-agent, β represents a second weight, coverage represents the coverage efficiency of the task, and γ represents a third weight; The path constraint of each agent is obtained, each agent starts from a starting path point, passes through a series of path points, and reaches a final path point, and its expression is: wherein each path point p ik is a coordinate point of the plane R 2 , S i denotes the starting path point, and E i denotes the final path point; The task allocation constraint is obtained, each target is searched by one agent, and each agent can be responsible for the search of multiple targets, and its expression is: The obstacle constraint is obtained, the path of the agent must avoid all obstacles, and the distance between each path point and any obstacle is greater than the radius of the obstacle, and its expression is: Wherein, the area of the obstacle region is circular, the center is O k , and the radius is r k .

3. The method of claim 2, wherein, For each agent i, the path length L i is expressed as: where d(p ik ,p i(k+1) ) denotes the distance between path point p ik and path point p i(k+1) , and the total length L path of the multi-agent path is expressed as: Threat strength T in multi-agent threat The expression for T is: wherein r represents the index of the radar, R represents the total number of radars, p i represents the position of the i-th agent, R max represents the maximum detection range of the radar, f threat (p i ,r) represents the threat intensity between the agent i and the radar r, d(p i ,r) represents the distance between the agent i and the radar r; The expression of the coverage efficiency of the task is: Wherein, cover(j) represents whether target j is searched, if target j is searched by at least one agent, cover(j)=1, otherwise 0.

4. The method of claim 2, wherein, The initial path planning and initial task allocation scheme of the multi-agent are obtained, comprising: Initialize the particle swarm, and initialize the position and speed of the particle to a random value; wherein each particle represents a solution, including an agent path sequence and a task classification matrix, and the dimension of each particle includes the path planning and task allocation scheme of the multi-agent; The fitness value of each particle is calculated using the objective function; According to the experience of the particle itself and the experience of the group, the speed and position of the particle are updated, and the expression is: x u (t+1) = x u (t) + v u (t+1); where v u (t) represents the velocity of the particle, x u (t) represents the position of the particle, ω represents the inertia weight, c1 and c2 represent the learning factors, rand1 and rand2 represent random numbers, pbest u u represents the historical best position of the particle u, gbest represents the global optimal position, and t represents the iteration number; updating the historical best position pbest of each particle according to the particle fitness value u and the global best position gbest; When the preset iteration number is reached or the fitness value of the particle is less than the threshold value, the initial path planning and initial task allocation scheme of the multi-agent are obtained.

5. The method of claim 1, wherein, According to the corresponding initial path planning and initial task allocation scheme, each agent adjusts the initial path planning to execute the task in combination with real-time perception data, obtains an execution result, comprising: The ith agent is initialized according to the received corresponding initial path planning and initial task allocation scheme, and a reward function is set; wherein the position of the agent and the task allocation are taken as the initial state of reinforcement learning, and the initial state includes the current position of the agent, the position of the target and the allocation scheme, and the search state of the task; the agent generates an initial action according to the corresponding initial path planning; The ith agent takes the initial state and the initial action as the behavior of the starting path point, and executes the task according to the initial path planning; The i-th agent combines the real-time perception data and the k-1th path point p in the initial path planning ik-1 The Q value for the kth path point p ik , adjust the current state s of the agent i , select the current action a that matches the current state i , as the behavior of the kth path point, and obtain the reward; at the same time, obtain the kth path point p ik The Q value is expressed as: Q(s i ,a i ) = Q(s i ,a i ) + a[r i + y max a′ Q(s i+1 ,a' ) - Q(s i ,a i )] ; where Q(s i ,a i ) represents the current Q value of the current action a i under the current state s i , r i represents the reward after executing the current action a i , γ represents a discount factor, α represents a learning rate, and max a′ Q(s i+1 ,a′) represents the maximum Q value of all possible actions under the state s i+1 . When the ith agent executes the initial task allocation scheme, the initial path planning is updated to obtain an execution result; wherein the execution result includes a path actual execution result, threat avoidance information and environmental dynamic changes.

6. The method of claim 5, wherein, The reward function includes: Task completion reward: when the agent completes the search task, a positive reward is given; Threat avoidance reward: when the agent successfully avoids the threat source, a positive reward is given; Path adjustment reward: if the path length of the agent is less than the preset path length, or the agent avoids obstacles, a positive reward is given; Collision penalty: if the agent collides with obstacles or enters a high-threat area, a negative reward is given.

7. The method of claim 5, wherein, The acquisition process of the real-time perception data includes: Real-time perception data is acquired through an environmental perception module arranged on the agent.

8. The method of claim 1, wherein, The global optimization layer judges whether the initial path planning and the initial task allocation scheme of the multi-agent need to be updated according to the execution results of the agents, if yes, updates the global model, updates the initial path planning and the initial task allocation scheme of the multi-agent, and sends them to the agents, and if not, saves the execution results of the agents, including: The global optimization layer analyzes the received execution results of the agents, compares the initial path planning and the path actual execution result of each agent, and evaluates the effectiveness of the path actual execution result; evaluates the distribution of the threat source in the threat avoidance information of each agent; evaluates the environmental changes in the environmental dynamic changes of each agent; obtains an evaluation result of the execution results of each agent; According to the evaluation result, it is judged whether the initial path planning and the initial task allocation scheme of the multi-agent need to be updated; If yes, the global optimization layer updates the global model in combination with the path actual execution result, the threat information and the environmental dynamic changes of each agent, updates the initial path planning and the initial task allocation scheme of the multi-agent; If not, the global optimization layer saves the execution results of the agents.

9. The method of claim 1, wherein, The first preset frequency is less than the second preset frequency.

10. A device for dynamic optimization and real-time decision of multi-agent cooperative target search, for implementing the method of dynamic optimization and real-time decision of multi-agent cooperative target search according to any one of claims 1-9, characterized in that, It includes: A global optimization module is configured to construct a global model according to the target position and the obstacle distribution of the search scene, to obtain the initial path planning and the initial task allocation scheme of the multi-agent, and to send them to the agents; at the same time, the global optimization layer updates the global model at a first preset frequency, and updates the initial path planning and the initial task allocation scheme of the multi-agent; A lightweight local reinforcement learning module is configured to adjust the initial path planning of each intelligent agent to perform a task according to the initial path planning and the initial task allocation scheme and real-time perception data, obtain an execution result, and send the execution result to the global optimization layer; meanwhile, the intelligent agent adjusts the initial path planning to perform a task at a second preset frequency, and updates the execution result; An information feedback module is configured to determine, by the global optimization layer, whether the initial path planning and the initial task allocation scheme of the multi-agent need to be updated according to the execution result of each intelligent agent, update the global model, update the initial path planning and the initial task allocation scheme of the multi-agent, and send the initial path planning and the initial task allocation scheme to each intelligent agent if the initial path planning and the initial task allocation scheme of the multi-agent need to be updated, or save the execution result of each intelligent agent if the initial path planning and the initial task allocation scheme of the multi-agent do not need to be updated; A multi-layer time scale optimization module is configured to keep the global optimization layer and each intelligent agent updated in cooperation, and obtain the path planning and the task allocation scheme in real time.

Citation Information

Patent Citations

  • Multi-agent coverage search and task allocation path optimization method, equipment and medium

    CN117762148A

  • Intelligent agent collaborative exploration method

    CN118707973A