Multi-Agent Environment Exploration Method, Device, System, Agent and Medium
By providing environmental exploration tasks and map information for multiple agents, using local feature extractors and action delay randomization technology to perform exploration actions asynchronously, the problem of low efficiency of synchronization algorithm framework in the existing technology is solved, and the efficiency of multi-robot exploration tasks is improved.
Patent Information
- Application Number
- CN202211080551.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-05
AI Technical Summary
In the existing multi-agent reinforcement learning method, the use of a synchronization algorithm framework causes all agents to operate simultaneously at each time step, which is inefficient and cannot effectively solve the multi-robot exploration task in the real world.
By providing each agent with environmental exploration tasks and map information for the target area, local feature extractors are used to extract local feature maps and perform exploration actions asynchronously, combining action delay randomization technology to simulate action delays in the real world.
It improves the efficiency of multi-robot exploration tasks, can better adapt to asynchronous action execution in the real world, and solves the problem of low efficiency of the synchronization algorithm framework.
Smart Images

Figure CN115357025B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine learning, and particularly relates to a method, device, system, agent and medium for multi-agent environment exploration. Background Art
[0002] Exploration is an important task in building intelligent robot systems and has been widely studied in many application fields such as rescue, autonomous driving, unmanned aerial vehicles, mobile robots, etc. It mainly focuses on the multi-robot cooperative exploration task, that is, multiple homogeneous robots explore an unknown spatial area in a cooperative manner. Due to the existence of multiple robots, learning the optimal cooperation strategy is particularly challenging. These robots must effectively allocate the exploration workload so that they can always navigate to different spatial areas to avoid trajectory conflicts, thereby achieving a significantly higher exploration efficiency than single-robot exploration.
[0003] Multi-agent reinforcement learning has become a trending method to solve this cooperative navigation challenge. The reinforcement learning-based method directly learns a neural policy end-to-end by interacting with a simulated environment. Compared with the planning-based solution, the reinforcement learning-based method has the powerful representation ability of complex policies, and once the policy is trained, the time overhead of forward inference can be ignored.
[0004] Classical multi-agent reinforcement learning algorithms usually adopt a synchronous algorithm framework, that is, all agents perform actions simultaneously, and all actions will be immediately executed at each time step. This process is mathematically represented as a distributed Markov decision process and is widely used in multi-agent reinforcement learning. Although such a mathematical framework is simple and elegant, for multi-robot exploration tasks in the real world, it may cause a series of problems. For a real robot system, each actual action is not synchronous and may take different times to complete. Due to unpredictable network communication or hardware failures, the resulting action delay may be more serious. Simply following the synchronous setting, that is, waiting for each robot to complete the previous action before making a new action plan, may be particularly inefficient in actual execution. Summary of the Invention
[0005] The present application provides a method, device, system, agent and medium for multi-agent environment exploration to solve the problems in the related art that a synchronous algorithm framework is adopted, all agents perform actions simultaneously, and all actions will be immediately executed at each time step, resulting in low efficiency and being unable to solve multi-agent exploration tasks in the real world.
[0006] The first aspect of the embodiments of the present application provides a method for multi-agent environment exploration, which is applied to an agent. Wherein, the method includes the following steps: obtaining an environment exploration task and map information of a target area; extracting a local feature map from the map information, and receiving local feature maps of other agents, determining an unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and using the unexplored area and the positional relationship to plan the next exploration action; controlling the agent to execute the exploration action and collecting exploration data when performing the environment exploration task, wherein the exploration data is used to generate an environment exploration result of the target area, and the exploration actions are executed asynchronously between different agents.
[0007] Optionally, in an embodiment of the present application, the extracting of the local feature map of the map information includes: using a local feature extractor of the agent to extract the local feature map of the map information, wherein the local feature extractor is composed of a three-layer convolutional neural network with shared weights.
[0008] Optionally, in an embodiment of the present application, the using of the unexplored area and the positional relationship to plan the next exploration action includes: controlling the agent to wait for a preset random simulation step; after waiting for the preset random simulation step, querying the next macro action of the agent according to the unexplored area and the positional relationship, and generating the next exploration action of the agent based on the next macro action.
[0009] Optionally, in an embodiment of the present application, after collecting the exploration data when performing the environment exploration task, it further includes: storing the exploration data in an internal cache of the agent, and pushing the cache data in the internal cache to a preset centralized data buffer at preset time intervals until the environment exploration task is completed.
[0010] The second aspect of the embodiments of the present application provides a multi-agent environment exploration device, which is applied to an agent. Wherein, the device includes: an obtaining module, configured to obtain an environment exploration task and map information of a target area; a planning module, configured to extract a local feature map from the map information, and receive local feature maps of other agents, determine an unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and use the unexplored area and the positional relationship to plan the next exploration action; a control module, configured to control the agent to execute the exploration action and collect exploration data when performing the environment exploration task, wherein the exploration data is used to generate an environment exploration result of the target area, and the exploration actions are executed asynchronously between different agents.
[0011] Optionally, in an embodiment of the present application, it is characterized in that the planning module is further configured to extract a local feature map of the map information by using a local feature extractor of the agent, wherein the local feature extractor consists of a three-layer convolutional neural network with shared weights;
[0012] The planning module is further configured to control the agent to wait for a preset random simulation step; after waiting for the preset random simulation step, query the next macro action of the agent according to the unexplored area and the position relationship, and generate the exploration action of the agent for the next step based on the next macro action.
[0013] Optionally, in an embodiment of the present application, it further includes: a cache module, configured to store the exploration data in an internal cache of the agent after collecting the exploration data for the exploration task of the execution environment, and push the cache data in the internal cache to a preset centralized data buffer at preset time intervals until the environment exploration task is completed.
[0014] An embodiment of the third aspect of the present application provides an agent, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the multi-agent environment exploration method as described in the above embodiment.
[0015] An embodiment of the fourth aspect of the present application provides a multi-agent environment exploration system, including: a plurality of agents; wherein each agent is configured to obtain an environment exploration task and map information of a target area; extract a local feature map from the map information, and receive local feature maps of other agents, determine an unexplored area and a position relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and plan the next exploration action by using the unexplored area and the position relationship; control the agent to execute the exploration action and collect exploration data for the exploration task of the execution environment, wherein the exploration data is used to generate an environment exploration result of the target area, and the exploration actions are executed asynchronously between different agents; a preset centralized data buffer, configured to receive the exploration data sent by each agent at preset time intervals; and a generation module, configured to generate an environment exploration result of the target area according to the exploration data of each agent.
[0016] An embodiment of the fifth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the multi-agent environment exploration method as described in the above embodiment.
[0017] Thus, the present application has at least the following beneficial effects:
[0018] By sending the environmental exploration tasks and map information of the target area to multiple agents, using the local feature extractors of each agent to extract local feature maps, extracting the local information features of each agent, and combining the local features of each agent to perform motion planning, controlling each agent to asynchronously execute the corresponding exploration actions, and using the method of action delay randomization to better simulate different action delays in the real world. Thus, the problems in the related art are solved, where a synchronous algorithm framework is adopted, all agents perform actions simultaneously, and all actions will be immediately executed at each time step, resulting in low efficiency and being unable to solve multi-robot exploration tasks in the real world, etc.
[0019] The additional aspects and advantages of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and / or additional aspects and advantages of the present application will become apparent and be easily understood from the following description of the embodiments in conjunction with the drawings, where:
[0021] Figure 1 is a flowchart of a method for environmental exploration of multiple agents according to an embodiment of the present application;
[0022] Figure 2 is a schematic diagram of MultiExploration and the real world according to an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of the workflow of a policy (MCP) based on multi-tower CNN according to an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of the pseudo code of asynchronous MAPPO according to an embodiment of the present application;
[0025] Figure 5 is a schematic block diagram of a device for environmental exploration of multiple agents according to an embodiment of the present application;
[0026] Figure 6 is a schematic block diagram of a system for environmental exploration of multiple agents according to an embodiment of the present application;
[0027] Figure 7 is a schematic diagram of the structure of an agent according to an embodiment of the present application.
[0028] Explanation of the reference numerals: acquisition module-100, processing module-200, exploration module-300, multiple agents-400, preset centralized data buffer-500, generation module-600, memory-701, processor-702, memory-703. DETAILED DESCRIPTION
[0029] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0030] The following describes the multi-agent environment exploration method, device, system, agent and medium of the embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a multi-agent environment exploration method, in which the environment exploration task and map information of the target area are sent to multiple agents, and the local feature extractor of each agent is used to extract the local feature map, and the action planning is carried out by extracting the local information features of each agent and combining the local features of each agent, and each agent is controlled to asynchronously execute the corresponding exploration action, and the action delay is randomized to better simulate the different action delays in the real world. Thus, the problem of using a synchronous algorithm framework in the related technology, in which all agents act at the same time, and all actions will be executed immediately at each time step, is low in efficiency, and cannot solve the problems of multi-robot exploration tasks in the real world.
[0031] Specifically, Figure 1 A flowchart of a multi-agent environment exploration method provided in an embodiment of the present application.
[0032] In order to bridge the gap between synchronous simulators and asynchronous action generation processes in real-world multi-agent exploration tasks, the present application embodiment proposes an asynchronous coordination explorer (ACE). ACE consists of three main parts: (1) a policy representation based on multi-tower CNN, (2) an asynchronous MAPPO for multi-agent reinforcement learning training, and (3) an action delay randomization mechanism for zero-shot generalization in the real world.
[0033] Specifically, in ACE, the policy of each agent is represented as a policy based on a multi-tower CNN. Among them, the CNN module is used to extract the local information features of each robot, and the fusion module combines the local features of each robot to perform motion planning. During the execution process, efficient information communication between robots can be achieved by directly exchanging the extracted low-dimensional CNN features. In addition, the embodiments of the present application can use an asynchronous variant of the multi-agent PPO algorithm to train the multi-agent policy, and also utilize the action delay randomization technique to better simulate the generalization of the real world.
[0034] As Figure 1 shown, the environment exploration method for this multi-agent includes the following steps:
[0035] In step S101, obtain the environment exploration task and map information of the target area.
[0036] The embodiments of the present application mainly consider a multi-agent collaborative exploration task, that is, a group of agents need to maximize the exploration area as quickly as possible within a limited time range. The embodiments of the present application can conduct simulations and experiments in a multi-agent exploration environment in a grid-based multi-room scenario, such as Figure 2 shown, implement the MultiExploration environment based on the GridWorld simulator. The original design of this simulator is for synchronous tasks. For the map information of the target area, the embodiments of the present application can consider two different map sizes, namely a 15*15 scenario containing 4-9 random rooms and a 25*25 scenario containing 4-25 random rooms. All agents are evenly distributed in the scenario at the beginning, and the available underlying actions are moving forward, turning left, and turning right. The details of the simulation environment of the embodiments of the present application can include three parts: the observation space, the action space, and the reward function, which are detailed as follows:
[0037] 1) Observation space
[0038] The input of ACE is an S*S image with 7 channels, where S represents the size of the map to be explored, including an obstacle channel, an explored map channel, a location information channel, a historical trajectory information channel, and three 7*7 agent local perspective channels. It should be noted that each agent only needs to maintain its local information, which has high memory efficiency and high communication efficiency in actual deployment.
[0039] 2) Action space
[0040] ACE executes in a hierarchical manner, generating a global target through a macro operation and several underlying operations for this target. The role of ACE is to generate a global target (u x , u y ), representing the grid selected in the map.
[0041] 3) Reward function
[0042] The team-based reward function is the sum of coverage reward, successful exploration reward, and overlapping exploration penalty. Let Ratio t represent the total coverage rate at time t, Exp a t be the exploration area of robot a, and Exp t represent the combined exploration area of all robots. Exp t and Exp t a are both sets of explored grids.
[0043] 3.1 Coverage reward: It is proportional to the size of the newly discovered area of the team Exp t \Exp t-1 .
[0044] 3.2 Successful exploration reward: Agent a will receive a success reward of 1*Ratio t when it reaches 98% coverage.
[0045] 3.3 Overlapping exploration penalty: The overlapping penalty r Overlap $ is designed to punish repeated exploration and encourage cooperation with other agents. Its definition is:
[0046]
[0047] where A overlap is the increment of the overlapping exploration area between robot a and other robots. The overlapping area between agent a and robot w is:
[0048]
[0049] In step S102, extract the local feature map from the map information, receive the local feature maps of other agents, determine the unexplored area and the positional relationship between the agent and other agents based on the local feature map of the agent and the local feature maps of other agents, and use the unexplored area and the positional relationship to plan the next exploration action.
[0050] During actual execution, the embodiment of the present application can use the local feature extractor of the agent to extract the local feature map of the map information, where the local feature extractor consists of a three-layer convolutional neural network with weight sharing.
[0051] It is understandable that the local feature extractor in the embodiments of the present application is a three - layer CNN with shared weights, which can extract a G*G*4 feature map from the s*s*7 local information of each agent. The robot transmits the extracted feature map instead of the original local information, which will greatly reduce the communication traffic. In the embodiments of the present application, G = 5 is adopted. Therefore, the communication traffic is reduced by ~97% in the 25*25 mapping and by ~93% in the 15*15 mapping.
[0052] Specifically, in the embodiments of the present application, in the scenario of a randomly sized 15*15 room in the MultiExploration environment, after training a global target planner, it can be directly deployed to a real agent system, such as Figure 2 As shown, the embodiments of the present application set up a real - world grid map of 15*15 that is the same as the simulator, and each grid is 0.31m long. The agent is equipped with Mecanum steering and an NVIDIA Jetson Nano processor. The position and orientation of the agent are tracked by an OptiTrack camera and motion capture software. In the embodiments of the present application, by sending environment exploration tasks and map information to multiple agents, local feature maps are extracted using the local feature extractor of each agent, and the feature maps extracted from different agents are aggregated using a relational encoder to better capture the internal interaction mechanism of the agents. Thus, according to the exploration area and the positional relationship of each agent, the agent is controlled to plan corresponding exploration actions.
[0053] Furthermore, the embodiments of the present application can generate global targets based on a multi - tower CNN - based policy (MCP). As Figure 3 shown, the MCP consists of three parts, namely a CNN - based local feature extractor, an attention - based relational encoder, and an action decoder.
[0054] The relational encoder mainly aggregates the feature maps extracted from different agents to better capture the internal interaction mechanism of the agents. In team - based exploration, the agent should not only discover undiscovered areas but also discover the movements between teammates for better scheduling among the agents.
[0055] The embodiments of the present application can adopt a simplified Transformer module as a team - size - invariant relational encoder. Different from the previous applications of transformers to one - dimensional text in NLP and two - dimensional images in vision, three - dimensional information (the feature map is two - dimensional and the team is one - dimensional) can be fused in multi - agent exploration. Inspired by the vision transformation model, the embodiments of the present application apply multi - head cross - attention to derive a single team - size - invariant representation of size G*G*4, as Figure 3As shown in the figure. Finally, the action decoder predicts the agent's strategy based on the aggregated representation of the multivariate categorical distribution to select a global goal from the plane.
[0056] In step S103, the control agent executes exploration actions and collects exploration data during the execution of the environment exploration task. The exploration data is used to generate the environment exploration result of the target area, and the exploration actions are executed asynchronously among different agents.
[0057] It can be understood that the ideal multi-agent reinforcement learning framework used in the real world should be asynchronous, that is, whenever an agent completes the current action, it should immediately generate the next action. Therefore, the embodiments of the present application can control each real agent to execute the corresponding exploration action in a distributed and asynchronous manner and collect the exploration data during the execution of the environment exploration task. The local information of each agent is sent to the global planner obtained by reinforcement learning training to generate a global goal, and the A star algorithm is used to execute 5 local actions on the local map to follow the global goal to solve the multi-agent exploration task in the real world.
[0058] In the actual execution process, the communication between agents adopts a request sending mechanism. After all execution actions are completed, the latest local extraction features of other agents are obtained through the ROStopic.
[0059] In an embodiment of the present application, after collecting the exploration data during the execution of the environment exploration task, it further includes: storing the exploration data in the internal cache of each agent, and pushing the cache data in the internal cache to a preset centralized data buffer at preset time intervals until each agent has completed the environment exploration task, and generating the environment exploration result of the target area based on the data in the preset centralized data buffer.
[0060] It can be understood that the embodiments of the present application can extend a policy-based MARL algorithm, MAPPO, to an asynchronous setting, which is called asynchronous MAPPO. The original MAPPO assumes synchronous execution of all agents; at each time step, all agents take actions simultaneously, and the trainer waits for all new transitions and then inserts them into the centralized data buffer for reinforcement learning training. In asynchronous MAPPO, different agents may not take actions simultaneously (some agents may even get stuck and not return new observations at all), which makes it impossible for the trainer to collect transitions in the original synchronous manner. Therefore, the embodiments of the present application allow each agent to store its own transition data in a separate cache and periodically push the cached data to the centralized data buffer. Then, the embodiments of the present application can run the standard MAPPO training algorithm on this buffer until each agent has completed the environment exploration task, thereby generating the exploration task results of multiple agents. The pseudocode of asynchronous MAPPO is shown in Alg.1, as Figure 4 shown.
[0061] In an embodiment of the present application, the next exploration action is planned using the unexplored area and the position relationship, including: controlling the agent to wait for a preset random simulation step; after waiting for the preset random simulation step, querying the next macro action of the agent according to the unexplored area and the position relationship, and generating the next exploration action of the agent based on the next macro action.
[0062] The preset random simulation step can be set according to the actual situation without specific limitation.
[0063] It can be understood that the training process is executed in a grid world simulator, where agents execute local actions synchronously without considering the different execution costs of different actions. In addition, in reality, problems such as hardware failures and network congestion can cause agents to delay action execution. These problems create a huge gap between simulation and reality. To narrow this gap, the embodiments of the present application can apply action delay randomization during training to simulate the action delay challenges in the real world. Specifically, for each action generation step of each agent, before querying the next macro action, the embodiments of the present application can wait for a random simulation step from 3 to 5, and then generate the next exploration action of each agent.
[0064] According to the multi-agent environment exploration method proposed in the embodiments of the present application, by sending the environment exploration task and map information of the target area to multiple agents, using the local feature extractor of each agent to extract the local feature map, by extracting the local information features of each agent and combining the local features of each agent, action planning is performed to control each agent to asynchronously execute the corresponding exploration actions, and action delay randomization is used to better simulate different action delays in the real world. Thus, the problems in the related art that adopt a synchronous algorithm framework, all agents perform actions simultaneously, and all actions will be immediately executed at each time step, with low efficiency and inability to solve multi-robot exploration tasks in the real world are solved.
[0065] Next, a multi-agent environment exploration device proposed according to the embodiments of the present application will be described with reference to the accompanying drawings.
[0066] Figure 5 It is a block diagram of a multi-agent environment exploration device according to the embodiments of the present application.
[0067] The device according to the embodiments of the present application is applied to an agent, such as Figure 5 As shown, the multi-agent environment exploration device 10 includes: an acquisition module 100, a planning module 200, and a control module 300.
[0068] Among them, the acquisition module 100 is used to acquire the environment exploration task and map information of the target area; the planning module 200 is used to extract the local feature map from the map information, receive the local feature maps of other agents, determine the unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and plan the next exploration action by using the unexplored area and the positional relationship; the control module 300 is used to control the agent to execute the exploration action and collect the exploration data when performing the environment exploration task, where the exploration data is used to generate the environment exploration result of the target area, and the exploration actions are executed asynchronously between different agents.
[0069] In an embodiment of the present application, the planning module 200 is further used to extract the local feature map of the map information by using the local feature extractor of the agent, where the local feature extractor is composed of a three-layer convolutional neural network with shared weights; the planning module 200 is further used to control the agent to wait for a preset random simulation step; after waiting for the preset random simulation step, query the next macro action of the agent according to the unexplored area and the positional relationship, and generate the next exploration action of the agent based on the next macro action.
[0070] In one embodiment of the present application, the device 10 of the embodiments of the present application further includes: a cache module, configured to store the exploration data in the internal cache of the agent after collecting the exploration data during the execution environment exploration task, and push the cache data in the internal cache to a preset centralized data buffer every preset time period until the environment exploration task is completed.
[0071] It should be noted that the foregoing explanation of the embodiments of the multi-agent environment exploration method also applies to the multi-agent environment exploration device of this embodiment, and will not be elaborated here.
[0072] The multi-agent environment exploration device proposed according to the embodiments of the present application issues the environment exploration task and map information of the target area to multiple agents, extracts the local feature map by using the local feature extractor of each agent, extracts the local information features of each agent, and combines the local features of each agent to perform action planning, controls each agent to asynchronously execute the corresponding exploration actions, and uses action delay randomization to better simulate different action delays in the real world. Thus, the problems in the related art that adopt a synchronous algorithm framework, all agents perform actions simultaneously, and all actions will be immediately executed at each time step, with low efficiency and inability to solve the multi-robot exploration tasks in the real world are solved.
[0073] In addition, the embodiments of the present application also propose a multi-agent environment exploration system.
[0074] Figure 6 It is a block diagram of a multi-agent environment exploration system according to an embodiment of the present application.
[0075] As Figure 6 shown, the multi-agent environment exploration system 20 includes: multiple agents 400, a preset centralized data buffer 500, and a generation module 600.
[0076] Among them, each agent 400 is configured to obtain the environment exploration task and map information of the target area; extract the local feature map from the map information, and receive the local feature maps of other agents, determine the unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and use the unexplored area and the positional relationship to plan the next exploration action; control the agent to execute the exploration action, and collect the exploration data during the execution of the environment exploration task, where the exploration data is used to generate the environment exploration result of the target area, and different agents execute the exploration actions asynchronously; the preset centralized data buffer 500 is configured to receive the exploration data sent by each agent every preset time period; the generation module 600 is configured to generate the environment exploration result of the target area according to the exploration data of each agent.
[0077] The multi-agent environment exploration system proposed according to the embodiments of the present application sends the environment exploration task and map information of the target area to multiple agents, extracts the local feature maps by using the local feature extractors of each agent, extracts the local information features of each agent, and combines the local features of each agent to perform action planning, controls each agent to asynchronously execute the corresponding exploration actions, and randomizes the action delay to better simulate different action delays in the real world. Thereby, it solves the problems in the related art that adopt a synchronous algorithm framework, all agents perform actions simultaneously, and all actions will be immediately executed at each time step, with low efficiency and inability to solve multi-robot exploration tasks in the real world, etc.
[0078] Figure 7 It is a schematic structural diagram of the agent provided by the embodiments of the present application. The agent may include:
[0079] A memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702.
[0080] When the processor 702 executes the program, it implements the performance optimization method of the adaptive cruise control system provided in the above embodiments.
[0081] Further, the agent further includes:
[0082] A communication interface 707 for communication between the memory 701 and the processor 702.
[0083] The memory 701 is used to store a computer program executable on the processor 702.
[0084] The memory 701 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0085] If the memory 701, the processor 702, and the communication interface 707 are independently implemented, the communication interface 707, the memory 701, and the processor 702 can be interconnected through a bus and communicate with each other. The bus may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0086] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 707 are integrated on a single chip, the memory 701, the processor 702, and the communication interface 707 can communicate with each other through an internal interface.
[0087] The processor 702 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0088] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above multi-agent environment exploration method is implemented.
[0089] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0090] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0091] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0092] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, and the like.
[0093] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0094] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for multi-agent environmental exploration, characterized in that The method is applied to an agent, and the method includes the following steps: Obtain the environmental exploration task and map information of the target area; Extract the local feature map from the map information, and receive the local feature maps of other agents. Determine the unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and use the unexplored area and the positional relationship to plan the next exploration action; Control the agent to execute the exploration action, and collect exploration data during the execution of the environmental exploration task. The exploration data is used to generate the environmental exploration result of the target area. An asynchronous coordination resource manager is used to realize the asynchronous execution of exploration actions between different agents. The asynchronous coordination resource manager includes a policy based on a multi-tower convolutional neural network, asynchronous multi-agent proximal policy optimization MAPPO, and an action delay randomization mechanism. The policy of each agent is represented as the policy based on the multi-tower convolutional neural network; Asynchronous MAPPO is obtained by asynchronously expanding the multi-agent reinforcement learning algorithm based on the policy and is used for multi-agent reinforcement learning training. In the asynchronous multi-agent proximal policy optimization MAPPO; The action delay randomization mechanism is used to simulate zero-shot generalization in the real world; After collecting the exploration data during the execution of the environmental exploration task, it further includes: storing the exploration data in the internal cache of the agent, and pushing the cache data in the internal cache to a preset centralized data buffer every preset time interval until the environmental exploration task is completed; The extraction of the local feature map of the map information includes: using the local feature extractor of the agent to extract the local feature map of the map information, where the local feature extractor is composed of a three-layer convolutional neural network with shared weights.
2. The method according to claim 1, characterized in that The use of the unexplored area and the positional relationship to plan the next exploration action includes: Control the agent to wait for a preset random simulation step; After waiting for the preset random simulation step, query the next macro action of the agent according to the unexplored area and the positional relationship, and generate the next exploration action of the agent based on the next macro action.
3. An environment exploration device for multi-agent, characterized in that, The device is applied to an agent, and the device includes: An acquisition module, configured to acquire the environmental exploration task and map information of the target area; A planning module, configured to extract the local feature map from the map information, receive the local feature maps of other agents, determine the unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and use the unexplored area and the positional relationship to plan the next exploration action; A control module for controlling the agent to perform the exploration action and collecting exploration data when performing the environmental exploration task of the execution environment, wherein the exploration data is used to generate an environmental exploration result of the target area, and an asynchronous coordination resource manager is used to asynchronously execute exploration actions between different agents. The asynchronous coordination resource manager includes a policy based on a multi-tower convolutional neural network, asynchronous multi-agent proximal policy optimization (MAPPO), and an action delay randomization mechanism. The policy of each agent is represented as the policy based on the multi-tower convolutional neural network; the asynchronous MAPPO is obtained by asynchronously expanding a multi-agent reinforcement learning algorithm based on the policy and is used for multi-agent reinforcement learning training. In the asynchronous multi-agent proximal policy optimization (MAPPO), the action delay randomization mechanism is used to simulate zero-shot generalization in the real world; A cache module for storing the exploration data in the internal cache of the agent after collecting the exploration data when performing the environmental exploration task of the execution environment, and pushing the cache data in the internal cache to a preset centralized data buffer every preset time period until the environmental exploration task is completed; The planning module is further configured to: extract a local feature map of the map information by using a local feature extractor of the agent, wherein the local feature extractor is composed of a three-layer convolutional neural network with shared weights.
4. The apparatus according to claim 3, wherein The planning module is further configured to: control the agent to wait for a preset random simulation step; after waiting for the preset random simulation step, query the next macro action of the agent according to the unexplored area and the positional relationship, and generate the next exploration action of the agent based on the next macro action.
5. An agent, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent environmental exploration method according to claim 1 or 2.
6. A multi-agent environment exploration system, characterized in that, Comprising: Multiple agents, where each agent is used to obtain the environmental exploration task and map information of the target area; extract the local feature map from the map information, receive the local feature maps of other agents, determine the unexplored area and the positional relationship between the agent and other agents according to the local feature map of the agent and the local feature maps of other agents, and plan the next exploration action using the unexplored area and the positional relationship; control the agent to execute the exploration action and collect the exploration data during the execution of the environmental exploration task, where the exploration data is used to generate the environmental exploration result of the target area, and an asynchronous coordination resource manager is used to asynchronously execute the exploration actions between different agents, where the asynchronous coordination resource manager includes a policy based on a multi-tower convolutional neural network, asynchronous multi-agent proximal policy optimization MAPPO, and an action delay randomization mechanism, where the policy of each agent is represented as the policy based on the multi-tower convolutional neural network; asynchronous MAPPO is obtained by asynchronously expanding the multi-agent reinforcement learning algorithm based on the policy and is used for multi-agent reinforcement learning training, in the asynchronous multi-agent proximal policy optimization MAPPO; the action delay randomization mechanism is used to simulate zero-shot generalization in the real world; A preset centralized data buffer for receiving the exploration data sent by each agent at preset time intervals; A generation module for generating an environmental exploration result of the target area according to the exploration data of each agent.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the environmental exploration method for multiple agents as described in claim 1 or 2.
Citation Information
Patent Citations
Multi-robot collaborative exploration method, device and system with unknown initial state
CN113110455A
Autonomous exploration mapping system and exploration mapping method based on multi-robot cooperation
CN114859375A