An Energy-Efficient Multi-Agent Exploration System Based on Deep Reinforcement Learning

By introducing deep reinforcement learning, efficiency-oriented observation and waiting mechanisms into the multi-agent exploration system, the problem of insufficient collaboration capabilities in multi-agent exploration is solved, and efficient exploration and energy-saving effects are achieved.

CN119761220BActive Publication Date: 2025-05-27DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510258711.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-27
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing multi-agent exploration methods are insufficiently designed in unknown environments, resulting in increased redundant exploration and energy consumption, making it difficult to achieve efficient collaboration.

Method used

The energy-saving multi-agent exploration system based on deep reinforcement learning is adopted, and combined with efficiency-oriented observation, inter-agent policy network and waiting mechanism, the collaboration and exploration behavior of the agent are optimized.

Benefits of technology

By accurately modeling the environment and strengthening collaboration among agents, avoiding redundant exploration, significantly improving the efficiency of collaborative exploration of multiple agents in unknown environments and saving energy costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761220B_ABST
    Figure CN119761220B_ABST
Patent Text Reader

Abstract

The present invention belongs to the fields of artificial intelligence, multi-agent systems, deep reinforcement learning, and autonomous exploration, and discloses an energy-saving multi-agent exploration system based on deep reinforcement learning, including scenario modeling, an efficiency-oriented observation system, an inter-agent policy network, a sequential decision-making mechanism module, and a waiting mechanism. The efficiency-oriented observation system proposed by the present invention constructs a dual structure of a connection graph and an interaction graph, and attaches multi-dimensional cooperation-oriented features such as exploration values, distance features, and complexity information to the nodes, enabling the system to accurately grasp the environmental features and the interaction relationships between agents, and significantly improving the cooperation efficiency of the multi-agent system. The sequential decision-making mechanism module designed by the present invention combines a sequential decision-making mechanism and a waiting mechanism. By allowing agents to refer to the decisions of previous agents and choose to wait at appropriate times, it avoids the energy waste caused by blind cooperation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence, multi-agent systems, deep reinforcement learning and autonomous exploration, and relates to an energy-saving multi-agent exploration system based on deep reinforcement learning, specifically an energy-saving multi-agent exploration system, which combines efficiency-oriented observation, inter-agent strategy network and waiting mechanism to improve the collaborative exploration efficiency of multi-agents in unknown environments. Background Art

[0002] With the development of artificial intelligence and robotics, multi-agent systems have shown great application potential in the exploration of unknown environments and have been widely used in target transportation, security monitoring, disaster detection, and environmental management. Multi-agent exploration tasks require multiple agents to collaboratively explore in partially observable unknown environments, build environmental maps, and complete designated tasks. This technology has important practical value in many scenarios such as smart cities, emergency rescue, and deep-sea exploration.

[0003] At present, multi-agent exploration methods are mainly divided into three categories: boundary-based methods, path planning-based methods, and deep reinforcement learning-based methods. Boundary-based methods guide the exploration direction by identifying the boundaries between known and unknown areas, and have the advantage of being simple to implement. Path planning-based methods (such as RRT, PRM, etc.) focus on achieving regional coverage while ensuring path feasibility. In recent years, with the breakthrough of deep learning technology, methods based on deep reinforcement learning (such as MAPPO, QMIX, etc.) have begun to emerge in multi-agent exploration tasks, achieving agent strategy optimization through end-to-end training.

[0004] Although deep reinforcement learning-based methods can use learned strategies to achieve intelligent exploration, existing methods are mostly limited to the perception of local information and the optimization of short-term goals. There are still deficiencies in the design of intelligent agent collaboration mechanisms, making it difficult to achieve global and efficient collaboration. In addition, in complex environments, due to the limitations of distributed decision-making, existing methods are prone to repeated exploration and redundant movement, further increasing the total path length and energy cost of exploration.

[0005] Therefore, how to improve exploration efficiency while reducing redundant exploration and energy consumption, and design a more efficient and robust agent collaboration mechanism, has become a core issue that needs to be solved in the field of multi-agent exploration. The solution to these challenges will directly affect the promotion and deployment of multi-agent exploration technology in practical applications. Summary of the invention

[0006] The present invention aims to provide a multi-agent energy-saving autonomous exploration system based on deep reinforcement learning, aiming to solve problems such as path redundancy caused by insufficient collaboration in existing exploration algorithms. By proposing efficiency-oriented observations to accurately model the environment and extract interaction information, the collaboration between agents is strengthened and redundant exploration is avoided. Combined with the inter-agent strategy network, the ability of the entire system to understand the environment is deepened. In addition, the novel waiting mechanism further avoids additional exploration behavior. The present invention has made good improvements in reducing exploration overlap and saving energy costs for multi-agent exploration problems.

[0007] The technical solution of the present invention:

[0008] An energy-saving multi-agent exploration system based on deep reinforcement learning, the energy-saving multi-agent exploration system includes scenario modeling, efficiency-oriented observation system, inter-agent strategy network, sequential decision mechanism module and waiting mechanism;

[0009] (1) Scenario Modeling

[0010] The scene to be explored is defined as Two-dimensional occupancy grid map of , which contains the explored area and unexplored areas ,satisfy ; The explored area is further divided into free areas and occupied area ,satisfy ; Each agent is equipped with a detection range of Laser scanner to update explored areas ; At the beginning of the exploration, ; The optimization goal of the energy-saving multi-agent exploration system is to find the shortest agent trajectory To complete the exploration process of the entire area while ensuring energy efficiency and collaborative performance during the exploration process;

[0011] (2) Constructing an efficiency-oriented observation system

[0012] Enhance overall situational awareness and emphasize the interaction between agents by first building a connection graph , used to characterize the connectivity of the environment; in the explored area Uniform sampling in the middle, establish a node set , and based on the collision-free path for each node and its The nearest neighbor nodes establish an edge set ; Connection diagram Over time Dynamic update, the agent is updated according to the connection graph The feature selection of the middle neighbor nodes for the next action; the main function of the connection graph is to construct the topological structure of the environment to ensure that the agent can efficiently cover the unknown area during the exploration process; to strengthen the collaborative relationship between agents, an interaction graph is further constructed , whose node set is the same as that of the connection graph , and the edge set is to dynamically connect the current position node of the agent and the target node with a non-zero exploration value (the number of observable exploration boundaries at this node), forming a logical undirected graph; the interaction graph emphasizes the collaboration between agents. By dynamically connecting high-utility nodes and the positions of other agents, it promotes the effectiveness of collaborative exploration; the edge set is constructed to dynamically capture the collaborative needs between agents at time ;

[0013] Among the features representing the exploration value of each node, complexity information is a key feature used to quantify the exploration difficulty of the area and the potential collaborative needs; complexity information is obtained by clustering the boundary points corresponding to the nodes through the density clustering algorithm (DBSCAN), reflecting the bifurcation structure and exploration complexity within the area; complexity information is comprehensively defined by these characteristics. A higher complexity value indicates that the area may contain more branches, higher exploration difficulty, and agents need stronger collaboration to complete the task. For example, in a single-path area, the value of complexity information is usually low, while in a bifurcated or cross-path area, the value of complexity information will increase significantly, indicating that agents need to collaborate to allocate tasks. The calculation method of complexity information is as follows: First, extract the boundary points from the perception range of each node. The boundary points are the boundaries between the explored area and the unknown area; through the density clustering algorithm, the distribution of boundary points is divided into different clusters to reflect the geometric characteristics of the area where the node is located, including the number of bifurcations, the spacing between bifurcation points, and the angle of bifurcation; finally, directly use the number of clusters output by the density clustering algorithm as the value of complexity information ;

[0014] The introduction of complexity information makes up for the limitations brought by relying solely on exploration values. The exploration value mainly represents the number of unknown areas within the area, but it cannot effectively identify the potential collaborative needs in complex bifurcation scenarios. Through complexity information, the energy-saving multi-agent exploration system can identify key areas with high geometric complexity and reasonably allocate agent resources, thus avoiding redundant paths and resource waste caused by simply pursuing high exploration values. Complexity information can also dynamically adjust the behavior strategy of agents, enabling them to complete exploration tasks more accurately in complex scenarios and significantly improving the overall exploration efficiency.

[0015] (3) Construct an inter-agent policy network to enhance the cognitive and decision-making abilities of agents through an attention mechanism, which includes two encoders and a decoder ;

[0016] The two encoders extract node information from the connection graph and the interaction graph in parallel; Both are composed of 6 layers of multi-head self-attention layers, where the output of the upper multi-head self-attention layer becomes the input of the lower multi-head self-attention layer; the input of each multi-head self-attention layer includes a query vector , a key vector and a value vector , which are generated by learning matrices , , ; , , are iterable parameters initialized randomly; the weight is calculated by the dot product of the query vector and the key vector and normalized by a scaling factor , and then processed by the softmax function; to ensure that there is no information transfer between nodes without edges, a mask matrix is introduced. When there is an edge between node and node , otherwise ; after the connection graph and the interaction graph pass through the two encoders in parallel, the respectively output and are concatenated and projected into the -dimensional feature space to generate enhanced node features ; ;

[0017] The decoder is composed of one attention layer and one pointer layer, and generates an action probability distribution through multi-stage processing; first, the current node feature and the neighbor feature are extracted from the enhanced node feature , and the previous node feature is used as the query, and the enhanced node feature is used as the key and value input to the attention layer; the output generated by the attention layer is concatenated with the previous node feature and projected into the current enhanced feature ; finally, an action probability distribution is generated through the pointer layer, where the current enhanced feature As a query, as keys and values; the pointer layer dynamically adjusts the action space and outputs the policy distribution , representing the time The probability of selecting the next node, where represents the time the corresponding environmental observation, respectively represent the decision points at times , the time the node where it is located and the time all neighboring nodes of the node where it is located;

[0018] (4) Sequence decision mechanism module, which decomposes the actions of multiple agents into a series of ordered decision-making processes;

[0019] To solve the unstable behavior and inefficient cooperation problems in multi-agent exploration, the sequence decision mechanism module proposes a sequence decision mechanism, which innovatively decomposes the actions of multiple agents into a series of ordered decision-making processes, thus significantly improving the cooperation efficiency. Although the decision-making of the energy-saving multi-agent exploration system is executed in a synchronous framework, by introducing a sequential decision-making method, each agent can fully refer to the decision-making results of the previous agents when formulating its own actions, thus forming a more collaborative and holistic exploration behavior globally.

[0020] In the input of the pointer layer of the decoder , the current enhanced feature of agent is further enhanced by the previous action, where for agent , the energy-saving multi-agent exploration system combines the enhanced target node feature sequence selected by its previous agent at time with the current enhanced node feature output by the decoder . Here, is the th agent's , , ; the concatenated features are processed by a two-dimensional convolutional layer with a dimension of to reduce their dimensions back to a unified dimension, where represents the total number of agents in the energy-saving multi-agent exploration system, represents the feature dimension; finally, the output of the convolutional layer is used as the new input to the pointer layer to integrate the previous action information; the design of the convolutional layer can efficiently fuse the action sequence features between agents while retaining the independence and cooperation between individuals, thus enhancing the integrity and robustness of decision-making.

[0021] (5) Waiting mechanism, which realizes the reasonable allocation of exploration resources and the efficient cooperation of multiple agents by dynamically adjusting the behavior states of the agents.

[0022] In the multi-agent exploration task in a complex scenario, due to the high degree of non-uniformity of the environment and task distribution, there are often significant differences in the exploration progress of different agents. Some agents may complete the exploration tasks in their responsible areas earlier than other agents. If these agents blindly continue to cooperate or attempt to explore other areas after completing the tasks, it usually leads to unnecessary movement and energy waste. To solve this problem, the present invention proposes a waiting mechanism, which realizes the reasonable allocation of exploration resources and the efficient cooperation of multiple agents by dynamically adjusting the behavior states of the agents.

[0023] When a certain agent completes the exploration task in its responsible area, the energy-saving multi-agent exploration system makes a real-time assessment of its current state and the exploration task states of other agents; by adding a "wait" action to the action space of the agent to allow it to judge whether it needs to continue the exploration behavior according to the current observation; the energy-saving multi-agent exploration system dynamically judges whether the agent needs to continue to participate in cooperation or maintain a stationary state by sensing the task progress, regional complexity, and target node distribution of other agents; if the exploration tasks of other agents in their respective areas have not reached the level that requires external cooperation, the energy-saving multi-agent exploration system keeps the agent that has completed the task in a stationary state to avoid its meaningless node selection or regional movement; on the contrary, when other agents encounter exploration bottlenecks due to task difficulty or regional complexity, the agent that has completed the task is dynamically activated and continues to participate in the cooperation.

[0024] The implementation of the waiting mechanism depends on the real-time perception and dynamic assessment of the multi-agent states by the energy-saving multi-agent exploration system. The agent that has completed the task does not directly enter a completely stationary state, but is in a standby mode that can be activated at any time. When the system senses new task requirements or other agents encounter exploration problems, these standby agents can quickly resume their active states and re-participate in the exploration tasks. At the same time, the system can dynamically adjust the behavior of the agents by continuously monitoring the task states, so that they always match the current task requirements.

[0025] Advantages of the present invention:

[0026] (1) The efficiency-oriented observation system proposed by the present invention constructs a dual structure of a connection graph and an interaction graph, and attaches multi-dimensional cooperation-oriented features such as exploration values, distance features, and complexity information to the nodes, enabling the system to accurately grasp the environmental features and the interaction relationships between agents, and significantly improving the cooperation efficiency of the multi-agent system.

[0027] (2) The sequence decision-making mechanism module designed in the present invention innovatively combines the sequence decision-making mechanism and the waiting mechanism. By allowing the agent to refer to the decisions of previous agents and choose to wait at the appropriate time, it avoids the energy waste caused by blind cooperation. Tests show that, while maintaining the same exploration coverage rate, this system significantly improves the exploration efficiency of the system in large-scale complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a pipeline flow chart of the multi-agent exploration method.

[0029] Figure 2 It is a node connectivity graph.

[0030] Figure 3 It is an interaction graph between nodes.

[0031] Figure 4 It is an exploration case before introducing the interaction graph.

[0032] Figure 5 It is an exploration case after introducing the interaction graph. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following further illustrates the specific embodiments of the present invention in conjunction with the drawings and technical solutions.

[0034] Embodiment 1

[0035] A multi-agent energy-saving autonomous exploration system based on deep reinforcement learning (the overall process is as Figure 1 shown), the steps are as follows:

[0036] (1) First, the efficiency-oriented observation system accurately represents the observed features of the environment.

[0037] For a specific complex real-world scenario, the efficiency-oriented observation system uniformly samples 1024 points in the explored free area as the vertex set, and establishes collision-free connections for each vertex and its 20 nearest neighbors, thereby constructing a connection graph, as Figure 2 shown. To enhance the interaction perception between agents, the present invention attaches exploration value, distance feature, complexity information, and access feature to each node.

[0038] Among them, the exploration value and distance feature are obtained by calculating the number of observable boundaries and the Euclidean distance at the node respectively. For the calculation of complexity information, first, the set of observable boundary points of the current node is detected. Nodes with more than zero observable boundary points are defined as points with exploration potential. Subsequently, the density-based clustering algorithm DBSCAN is used to perform clustering analysis on the boundary points. Specifically, for each boundary point, the number of points within its neighborhood is calculated with a radius of ε. When the number of points exceeds the threshold, it is marked as a core point. By iteratively classifying all density-reachable boundary points into one class, k clusters are finally obtained. The complexity information is determined by the number k of clusters. The access feature is binarized according to whether the node has been visited.

[0039] When constructing the interaction graph, the system maintains the same node set and node features as the connection graph, but uses a manually defined edge set to highlight the collaborative relationship between agents. This edge set connects all nodes where agents are located to points with exploration potential to ensure the smooth flow of exploration information. Its schematic diagram is as Figure 3 shown. Finally, the efficiency-oriented observation module integrates the features of the connection graph and the interaction graph and outputs them, providing a structured environmental representation for the subsequent policy network. This graph-based multi-level feature extraction mechanism provides the necessary environmental perception information for the collaborative decision-making of agents. Figure 4 and Figure 5 shows the impact on the exploration performance before and after adding the interaction graph. From left to right are the exploration cases without the interaction graph and the exploration cases using the interaction graph. In the left figure, the exploration routes of multiple agents overlap with each other, resulting in a large amount of redundant exploration. In the right figure, the exploration routes of each agent complete the exploration of the entire map with a low overlap rate.

[0040] (2) The inter-agent policy network enhances the cognitive and decision-making abilities of agents through the inter-agent attention mechanism.

[0041] The proxy-inter-agent policy network enhances the agents' cognitive and decision-making abilities through the inter-agent attention mechanism. In the feature encoding stage, the system first starts two encoders in parallel to process the connection graph and the interaction graph respectively. Each encoder consists of 6 multi-head self-attention layers stacked sequentially. The attention layer calculates the similarity between the query vector and the key vector, and applies the obtained attention weights to the value vector, thereby adaptively extracting the important structural information in the graph. In the feature fusion stage, the outputs of the two encoders are integrated to form enhanced node features. The system extracts the node features representing the current position information of the agent and the adjacent node features describing the surrounding environment from the enhanced node features, and these features are then input into the decoder for processing. The decoder also uses the attention mechanism to generate a context representation that fuses local and global information by modeling the correlation between the current position features and the adjacent features. In the action generation stage, the output of the decoder is processed by the pointer layer. The pointer layer first calculates the correlation scores between the context representation and the candidate actions, and then converts these scores into a probability distribution through the softmax function. The entire processing flow is trained end-to-end through backpropagation, enabling the system to learn efficient feature extraction and decision-making strategies. Among them, the features are encoded into 128 dimensions, and the multi-head attention layer is set to 8 heads for learning.

[0042] (3) Avoid unstable actions and redundant exploration through the sequential decision-making mechanism and the waiting mechanism.

[0043] The sequential decision-making mechanism module avoids unstable actions and redundant exploration through the sequential decision-making mechanism and the waiting mechanism. In the sequential decision-making processing stage, the system processes the decisions of each agent in a preset order (from 1 to N) in sequence. For agent , first collect the set of enhanced target node features selected by the previous agents from 1 to , and splice these features with the enhanced node features of the current agent to form an extended feature vector. This extended feature vector is then processed by a 2D convolutional layer. The output of the convolutional layer is processed by a non-linear activation function to generate a fusion feature reflecting the collaborative relationship between agents.

[0044] In the action space construction stage, the system constructs a mixed action space containing two types of actions, namely moving and waiting, for each agent. The set of moving actions contains the neighboring nodes of the current position of the agent, representing the possible moving target positions that the agent can choose. At the same time, the system selects the waiting action as a single action vector, enabling the agent to choose to stay still at the appropriate time. This design of the mixed action space enables the system to flexibly respond to different exploration scenarios.

[0045] During the prediction phase of the policy network, the fused features are input into the policy network composed of a multi-layer neural network. The policy network first transforms the features through a fully connected layer and then uses an attention mechanism to enhance the extraction of important information. Finally, it outputs the probability distribution over the mixed action space through a softmax layer and samples to determine the action to be executed. This end-to-end training method enables the system to adaptively learn when to move and when to wait, thus effectively avoiding redundant exploration and energy waste. The entire decision-making process is optimized through experience replay and the SAC method to maximize the long-term exploration efficiency.

[0046] (4) Use a collaborative reward mechanism to encourage exploration behavior

[0047] To further improve the exploration efficiency of the agent and encourage collaborative behavior, the present invention uses an exploration incentive method based on a collaborative reward mechanism. This mechanism guides the agent to adopt a more globally valuable exploration strategy through a carefully designed reward function, optimizing resource utilization while enhancing the exploration efficiency. The reward function consists of three parts: exploration reward, distance penalty, and completion reward.

[0048] During the exploration process of the agent, the exploration reward mainly measures the amount of environmental information newly discovered, that is, the unknown area newly explored by the agent in each step of action. This reward encourages the agent to move towards unvisited areas to maximize the overall environmental coverage. At the same time, considering resource utilization, the system introduces a distance penalty to constrain the movement cost of the agent from the current position to the next position. This penalty term ensures that the agent does not perform overly long or unnecessary movements during exploration, thus saving energy and improving the economy of actions. In addition, to further motivate the agent to complete the entire exploration task, the system introduces a completion reward. When all areas to be explored are completely covered, the agent will receive a fixed reward. This reward is only given when the task is completed to ensure the ultimate achievement of the exploration goal.

[0049] Finally, the entire reward function, through the weighted combination of the exploration reward, distance penalty, and completion reward, enables the agent to reduce redundant actions and optimize resource consumption while ensuring efficient exploration, thereby enhancing the overall task completion efficiency.

Claims

1. An energy-saving multi-agent exploration system based on deep reinforcement learning, characterized in that: The energy-saving multi-agent exploration system includes scenario modeling, efficiency-oriented observation system, inter-agent strategy network, sequential decision mechanism module and waiting mechanism; (1) Scenario Modeling The scene to be explored is defined as Two-dimensional occupancy grid map of , which contains the explored area and unexplored areas ,satisfy ; The explored area is further divided into free areas and occupied area ,satisfy ; Each agent is equipped with a detection range of Laser scanner to update explored areas ; At the beginning of the exploration, ; The optimization goal of the energy-saving multi-agent exploration system is to find the shortest agent trajectory To complete the exploration process of the entire area while ensuring energy efficiency and collaborative performance during the exploration process; (2) Constructing an efficiency-oriented observation system First build the connection graph , used to characterize the connectivity of the environment; in the explored area Uniform sampling in the middle, establish a node set , and based on the collision-free path for each node and its The nearest neighbor nodes establish an edge set ; Connection diagram Over time Dynamic update, the agent is updated according to the connection graph The features of neighbor nodes in the middle select the next action; to strengthen the collaborative relationship between agents, further build an interaction graph , its node set and connection graph Consistent, edge set To dynamically connect the agent's current position node and the target node with a non-zero exploration value to form a logical undirected graph; Among the features that characterize the exploration value of each node, complexity information It is a key feature used to quantify the exploration difficulty and potential collaboration requirements of the region; complexity information The density clustering algorithm is used to cluster the boundary points corresponding to the nodes, which reflects the bifurcation structure and exploration complexity in the region; (3) Construct an inter-agent policy network to improve the cognitive and decision-making capabilities of the agent through the attention mechanism, including two encoders and a decoder ; Two encoders In parallel from the connection graph and interaction diagrams Extract node information from Each layer consists of 6 layers of multi-head self-attention layers, where the output of the upper multi-head self-attention layer becomes the input of the lower multi-head self-attention layer; the input of each multi-head self-attention layer includes the query vector , key vector Sum value vector , by learning the matrix , , generate, , , Iterable parameters for random initialization; weight By query vector With key vector The dot product is calculated and scaled by After normalization, it is processed by the softmax function; to ensure that there is no information transmission between nodes without edge connections, a mask matrix is ​​introduced , when the node and nodes When there is an edge ,otherwise ; Connection diagram and interaction diagrams After passing through two encoders in parallel, the output and are stitched and projected onto dimensional feature space to generate enhanced node features ; Decoder It consists of an attention layer and a pointer layer, and generates action probability distribution through multi-stage processing; first, it enhances node features Extract the current node features and neighbor characteristics , the previous node feature As a query, enhance node features Input to the attention layer as key and value; Output generated by the attention layer and features of the previous node After stitching, the projection is the current enhanced feature ; Finally, the action probability distribution is generated through the pointer layer, where the current enhanced feature As a query, As keys and values; the pointer layer dynamically adjusts the action space and outputs the strategy distribution , indicating the time The probability of selecting the next node is Indicates time The corresponding environmental observations, Respectively indicate time Decision points and moments Node and time All neighboring nodes of the node; (4) Sequential decision mechanism module, which decomposes the actions of multiple agents into a series of orderly decision-making processes; In the decoder In the input of the pointer layer, the agent Current Enhancements is further enhanced by the preceding action, where For the intelligent agent , the energy-saving multi-agent exploration system At the moment Selected enhanced target node feature sequence With decoder Output current enhancement node features To splice, here, For the Intelligent , ; The concatenated features are passed through a dimension of The 2D convolutional layer reduces it back to a uniform dimension, where represents the total number of agents in the energy-saving multi-agent exploration system, represents the feature dimension; finally, the output of the convolutional layer is used as the new Input pointer layer to integrate the previous action information; (5) Waiting mechanism: By dynamically adjusting the behavior state of the intelligent agent, it can achieve the rational allocation of exploration resources and the efficient collaboration of multiple intelligent agents.

2. The energy-saving multi-agent exploration system based on deep reinforcement learning according to claim 1 is characterized in that: In step (2), the complexity information The calculation method of is as follows: first, the boundary points are extracted from the perception range of each node. The boundary points are the boundaries between the explored area and the unknown area. Through the density clustering algorithm, the distribution of the boundary points is divided into different clusters to reflect the geometric characteristics of the area where the node is located, including the number of bifurcations, the spacing between bifurcations, and the angle of bifurcations. Finally, the number of clusters output by the density clustering algorithm is directly used as the complexity information. The value of .

3. The energy-saving multi-agent exploration system based on deep reinforcement learning according to claim 1 is characterized in that: The specific implementation process of step (5) is as follows: When an agent completes the exploration task of the area it is responsible for, the energy-saving multi-agent exploration system evaluates its current state and the exploration task status of other agents in real time; by adding a "wait" action in the action space of the agent, it allows it to judge whether it needs to continue the exploration behavior based on the current observation; the energy-saving multi-agent exploration system dynamically judges whether the agent needs to continue to participate in the collaboration or remain stationary by sensing the task progress, regional complexity and target node distribution of other agents; if the exploration tasks of other agents in their respective areas have not yet reached the level that requires external collaboration, the energy-saving multi-agent exploration system will keep the agent that has completed the task stationary to avoid meaningless node selection or regional movement; on the contrary, when other agents encounter exploration bottlenecks due to task difficulty or regional complexity, the agent that has completed the task is dynamically activated to continue to participate in the collaboration.

Citation Information

Patent Citations

  • Dynamic path optimization problem solving method based on deep reinforcement learning

    CN112116129A

  • Unmanned aerial vehicle group scheduling method based on reinforcement learning and attention mechanism

    CN113625757A