Navigation method, device and equipment in multi-agent environment and storage medium thereof
By constructing a multi-agent environment based on real navigation logs, the problem of unsatisfactory simulation effects of multi-agent environments in existing technologies is solved, and more accurate multi-agent task execution is achieved.
Patent Information
- Application Number
- CN202410732759.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-12-09
AI Technical Summary
Existing multi-agent environments typically only target a single agent. Randomly generated multi-agent environments differ significantly from real physical environments, resulting in unsatisfactory simulation effects.
By collecting navigation log data from multiple robots in a real physical environment, a multi-agent environment is constructed. The historical navigation data of the robots is then integrated to generate a more realistic multi-agent environment for executing multi-agent tasks.
It improves the accuracy of multi-agent task execution, making navigation results based on multi-agent environments more accurate.
Smart Images

Figure CN121089697A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of robot navigation, and in particular, to a navigation method and device in a multi-agent environment and a storage medium thereof. BACKGROUND
[0002] A multi-agent environment for mobile navigation tasks is a virtual platform for testing and optimizing mobile robot navigation technology. The multi-agent environment can simulate various complex scenarios and conditions in the real world, allowing researchers to conduct in-depth research on navigation algorithms, path planning, and other aspects without actual robot hardware.
[0003] Currently, most multi-agent environments for mobile navigation tasks are only for a single agent. For multi-robot navigation simulation scenarios, the multi-agent environment is a randomly generated environment. The randomly generated multi-agent environment may be quite different from the real physical environment, resulting in suboptimal simulation results. SUMMARY
[0004] Embodiments of the present application provide a navigation method and device in a multi-agent environment and a storage medium thereof. A multi-agent environment for multiple robots is constructed based on actual navigation logs of a large number of robots. The multi-agent environment is generated based on real data, making the constructed multi-agent environment more realistic, and thus making the execution results of multi-agent tasks based on the multi-agent environment more accurate.
[0005] In a first aspect, embodiments of the present application provide a navigation method in a multi-agent environment, the method comprising:
[0006] obtaining navigation log data of a plurality of robots in a target time period;
[0007] parsing the navigation log data to obtain N frames of historical navigation data of each robot, the historical navigation data including robot pose, local map, and navigation planning path;
[0008] determining a first target robot from the plurality of robots, and fusing the historical navigation data of the plurality of robots to obtain a multi-agent environment based on the first target robot as a subjective viewpoint, the multi-agent environment including N frames of global maps and poses of the plurality of agents in the N frames of global maps, wherein each robot corresponds to an agent;
[0009] performing multi-agent navigation in the multi-agent environment to execute a multi-agent task.
[0010] In some embodiments, the performing multi-agent navigation in the multi-agent environment to execute a multi-agent task comprises:
[0011] replacing a historical navigation strategy of a target agent corresponding to a second target robot in the plurality of robots with a new navigation strategy, reasoning in the multi-agent environment using agents corresponding to the plurality of robots, to obtain a reasoning result, wherein the reasoning result includes a new navigation trajectory of the target agent, and the historical navigation strategy is a navigation strategy used by the second target robot to generate the historical navigation data.
[0012] In some embodiments, the reasoning in the multi-agent environment using agents corresponding to the plurality of robots, to obtain a reasoning result, comprises:
[0013] using the initial value of the target agent as an input of the new navigation strategy, performing M-step reasoning in the multi-agent environment using the new navigation strategy, to obtain a new navigation trajectory of the target agent, wherein the initial value of the target agent includes a reasoning start time, a reasoning start position, and a destination;
[0014] controlling other agents to run in the multi-agent environment according to historical navigation trajectories of corresponding robots, wherein the other agents are agents corresponding to remaining robots in the plurality of robots except the second target robot.
[0015] In some embodiments, the method further comprises:
[0016] displaying the following in real time in the multi-agent environment during the reasoning process: a new position, a historical position, and a new planned path of the target agent, and a historical position of the other agents, wherein the new position is a navigation position obtained according to the reasoning of the new navigation strategy, the historical position is a position indicated by the same time historical navigation data, and the new planned path is a path obtained according to the reasoning of the new navigation strategy.
[0017] In some embodiments, the reasoning in the multi-agent environment using agents corresponding to the plurality of robots, to obtain a reasoning result, comprises:
[0018] using the initial value of the target agent as an input of the new navigation strategy, performing M1-step reasoning in the multi-agent environment using the new navigation strategy, to obtain a new navigation trajectory of the target agent, wherein the initial value of the target agent includes a reasoning start time, a reasoning start position, and a destination;
[0019] replace a history navigation strategy of another agent with the new navigation strategy, use initial values of the another agent as inputs of the new navigation strategy, perform M2-step reasoning in the multi-agent environment using the new navigation strategy, and obtain a new navigation trajectory of the another agent, wherein the initial values of the another agent include a reasoning start time, a reasoning start position, and a destination.
[0020] In some embodiments, the method further includes:
[0021] displaying in real time in the multi-agent environment during the reasoning the new position, the history position, and the new planned path of the target agent, and the new position and the history position of the another agent, wherein the new position is a position obtained according to the new navigation strategy, the history position is a position indicated by the same time history navigation data, and the new planned path is a path obtained according to the new navigation strategy.
[0022] In some embodiments, the performing multi-agent navigation in the multi-agent environment to perform the multi-agent task includes:
[0023] training a to-be-trained navigation strategy in the multi-agent environment using a reinforcement learning method or an imitation learning method.
[0024] In some embodiments, the performing multi-agent navigation in the multi-agent environment to perform the multi-agent task includes:
[0025] In the multi-agent environment, the agents corresponding to the plurality of robots perform navigation playback according to history navigation trajectories of the plurality of robots.
[0026] In some embodiments, the fusing the history navigation data of the plurality of robots to obtain the multi-agent environment includes:
[0027] aligning the history navigation data of each frame of the plurality of robots;
[0028] fusing local maps of N frames of the plurality of robots to obtain a global map of each frame;
[0029] determining, according to the history navigation data of the other robots, poses of the other robots in the global map with reference to a pose of the first target robot in each frame.
[0030] In some embodiments, the new navigation strategy is a strategy learned through a reinforcement learning method or an imitation learning method.
[0031] In some embodiments, the method further includes: displaying in real time the positions of the agents corresponding to the plurality of robots in the multi-agent environment during the inference process.
[0032] In some embodiments, the method further includes: evaluating the new navigation strategy by comparing the new navigation trajectory of the target agent with the historical navigation trajectory.
[0033] In some embodiments, the method further includes:
[0034] The new navigation strategy is evaluated by comparing the new navigation trajectory and the historical navigation trajectory of the target intelligent agent, and a first evaluation result is obtained;
[0035] The new navigation strategy is evaluated by comparing the new navigation trajectory with the historical navigation trajectory of the other intelligent agents, resulting in a second evaluation result;
[0036] The new navigation strategy is evaluated based on the first evaluation result and the second evaluation result.
[0037] Secondly, embodiments of this application provide a navigation device in a multi-agent environment, comprising:
[0038] The acquisition module is used to acquire navigation log data of multiple robots within a target time period;
[0039] The parsing module is used to parse the navigation log data to obtain N frames of historical navigation data for each robot. The historical navigation data includes robot pose, local map, and navigation planning path.
[0040] The fusion module is used to determine a first target robot from the plurality of robots, and to fuse the historical navigation data of the plurality of robots from the first target robot's subjective perspective to obtain a multi-agent environment. The multi-agent environment includes N frames of global map and the poses of the plurality of agents in the N frames of global map, wherein each robot corresponds to one agent.
[0041] An execution module is used to perform multi-agent navigation in the multi-agent environment to execute multi-agent tasks.
[0042] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method as described in the first aspect above.
[0043] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium for storing a computer program, the computer program causing a computer to execute the method of the first aspect.
[0044] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the method of the first aspect.
[0045] The method, device, and storage medium provided by the embodiments of the present application for navigation in a multi-agent environment are as follows: navigation log data of a plurality of robots in a target time period is obtained, the navigation log data of each robot is analyzed to obtain N frames of historical navigation data of each robot, the historical navigation data including a robot pose, a local map, and a navigation planning path; a first target robot is determined from the plurality of robots, the first target robot is taken as a subjective visual angle, the historical navigation data of the plurality of robots is fused to obtain a multi-agent environment, the multi-agent environment including N frames of global maps and poses of a plurality of agents corresponding to the plurality of robots in the N frames of global maps, and multi-agent navigation is performed in the multi-agent environment to execute a multi-agent task. The method constructs a multi-agent environment of a plurality of robots according to actual navigation logs of a large number of robots, the multi-agent environment is generated based on real data, the multi-agent environment constructed is more real, and therefore an execution result of the multi-agent task based on the multi-agent environment is more accurate. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0047] Figure 1 A flowchart of the navigation simulation method of the plurality of robots provided by the first embodiment of the present application;
[0048] Figure 2 A flowchart of the navigation simulation method of the plurality of robots provided by the second embodiment of the present application;
[0049] Figure 3 A display schematic diagram of the simulation navigation of the multi-agent;
[0050] Figure 4 A flowchart of the navigation simulation method of the plurality of robots provided by the third embodiment of the present application;
[0051] Figure 5A structural schematic diagram of a navigation simulation device of multiple robots provided for Embodiment Four of the present application;
[0052] Figure 6 A structural schematic diagram of an electronic device provided for Embodiment Five of the present application. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be clearly and completely described in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0054] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product, or device.
[0055] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration, any embodiment or aspect described herein as "exemplary" or "for example" is not necessarily to be construed as preferred or advantageous over other embodiments or aspects. Rather, use of the words "exemplary" or "for example" is intended to present concepts in a concrete manner.
[0056] In the description of the embodiments of the present application, "a plurality of" means two or more, i.e., at least two, unless otherwise specified. "At least one" means one or more.
[0057] The multi-agent environment for mobile navigation tasks is a virtual environment created by using a computer or other technical means, which is used for testing and optimizing mobile robot navigation. The multi-agent environment can simulate or reproduce various complex scenes and conditions in the real world, so that researchers can conduct in-depth research on navigation algorithms, sensor integration, path planning, etc. without actual robot hardware.
[0058] By simulating different navigation scenarios, developers can test and adjust the navigation algorithm to ensure it provides accurate and reliable navigation services in various environments. Compared to field testing, using a multi-agent environment for navigation simulation can significantly reduce development costs and time. Developers can simulate different scenarios multiple times in a short period of time, quickly identify problems and make improvements.
[0059] A multi-agent environment typically includes the following key components:
[0060] Scene modeling: A multi-agent environment needs to be able to create realistic three-dimensional scenes, including indoor and outdoor environments. These scenes can contain various obstacles, terrain changes, and lighting conditions to simulate the complexity of the real world.
[0061] Sensor simulation: In the navigation task, sensors play a crucial role. The multi-agent environment needs to be able to simulate the output of various sensors, such as lidar, camera, depth camera, ultrasonic sensor, etc. These simulated data should be as close as possible to the performance characteristics of real sensors, so as to get accurate results when testing the navigation algorithm.
[0062] Interaction function: The multi-agent environment should support user interaction with virtual robots, such as setting start points, target points, obstacles, etc. In addition, users can also adjust navigation parameters, view sensor data, analyze navigation paths, etc.
[0063] Performance evaluation: The multi-agent environment should have performance evaluation functions, which can quantify the completion time, path length, collision times, etc. of the navigation task. This helps researchers understand the performance of the navigation algorithm in different scenarios, so as to optimize and improve it.
[0064] By building a multi-agent environment, developers can more efficiently develop and test mobile robot navigation technology, reduce development costs, and shorten development cycles. At the same time, the multi-agent environment can also provide important reference for actual robot deployment, improve the navigation performance of robots in the real world.
[0065] An agent is a computer system or entity with autonomy and interactivity, which can also be understood as a virtual robot. Autonomy refers to the ability of an agent to autonomously formulate plans and execute actions based on its own goals and environmental state without direct human intervention. Interactivity refers to the ability of an agent to interact with other agents, humans or the environment through communication and collaboration to complete tasks.
[0066] Agents run in a multi-agent environment to simulate the behavior of physical robots, for example, agents navigate in a multi-agent environment.
[0067] A multi-agent environment with high fidelity and strong scalability is essential for exploring multi-robot navigation algorithms and establishing a perfect algorithm evaluation system. Currently, multi-agent environments usually only evaluate a single agent. For multi-agent interaction scenarios, a multi-agent environment is usually randomly generated. The randomly generated multi-agent environment does not have a real data distribution, resulting in poor simulation results.
[0068] To solve the problems of the prior art, the embodiment of the present application provides a navigation method in a multi-agent environment. The navigation log data generated by a plurality of robots running in a real physical environment is collected, and a multi-agent environment is constructed based on the navigation log data of the plurality of robots. The multi-agent environment constructed based on the real navigation log data has higher fidelity, so that the execution result of the multi-agent task based on the multi-agent environment is more accurate.
[0069] The technical solutions of the present application will be described in detail in some embodiments. The embodiments described below can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments.
[0070] Figure 1 The flowchart of the navigation method in the multi-agent environment provided by the embodiment of the present application, the method of the embodiment can be executed by a simulation device, which can be a terminal or a server. The terminal device can be a mobile phone, a tablet computer, a desktop computer, a portable notebook computer, a personal digital assistant, etc. The server can be a server cluster or a single server, and the server can be a cloud server, etc.
[0071] As shown in Figure 1 The method provided by the embodiment includes the following steps:
[0072] S101, obtaining navigation log data of a plurality of robots in a target time period.
[0073] The navigation log data includes running log data of the plurality of robots collected and recorded at a preset collection frame rate in the same time period (i.e. in the target time period). The preset collection frame rate is, for example, 10 frames per second, 20 frames per second, or 30 frames per second, etc. The running log data of each robot is the running trajectory data of the robot moving and recording in a real physical environment.
[0074] The plurality of robots are located in the same space or the same map. For example, when the robots are applied in a warehouse, the plurality of robots are located in the same warehouse. When the robots are applied in an office area, a hotel, a supermarket, or other indoor space, the plurality of robots are located in the same indoor space. The plurality of robots move in the same space to generate motion trajectories or running trajectories.
[0075] The target time period can be in time granularity of hours or minutes, for example, the target time period can be one or more hours, or the target time period is 20 minutes, 30 minutes, 60 minutes, etc.
[0076] The running log data of each robot includes a robot ID, a plurality of collection time points, a robot pose, a local map, and navigation path planning information at each collection time point.
[0077] The pose of the robot includes a position and a pose, the position of the robot can be represented by a three-dimensional coordinate of the robot in the map, and the pose of the robot is used to represent the orientation of the robot.
[0078] The local map is a map area that can be seen by the robot at the current position, for example, a map area within a radius of 6 meters from the current position of the robot as the origin is the local map.
[0079] The navigation path planning can include the starting point and the ending point of the path, and the navigation path planning can be a path planned by an upper application.
[0080] Each robot can collect and record its own navigation log data according to the log configuration information during operation, and send the navigation log data to a simulation device, which can obtain the navigation log data of a plurality of robots in the same space.
[0081] S102, the navigation log data is parsed to obtain N frames of historical navigation data of each robot, and the historical navigation data includes a robot pose, a local map, and a navigation planning path.
[0082] The navigation log data includes a plurality of frames of data, the navigation log data can be stored according to a log structure, and the navigation log data can also include some data unrelated to navigation. By parsing the navigation log data, N frames of historical navigation data of each robot can be obtained, and the historical navigation data includes a robot pose, a local map, and a navigation planning path, each frame corresponding to a time point.
[0083] Optionally, the historical navigation data also includes a historical navigation trajectory, the historical navigation trajectory is a path passed by the robot from a starting point to a destination, and the historical navigation trajectory includes a plurality of position points, and the position of the robot at each frame can be taken as a position point of the historical navigation trajectory.
[0084] S103, a first target robot is determined from the plurality of robots, and the historical navigation data of the plurality of robots is fused from the first target robot as a subjective visual angle to obtain a multi-agent environment, the multi-agent environment includes N frames of global maps and poses of a plurality of agents in the N frames of global maps, wherein each robot corresponds to an agent.
[0085] The fusion of the historical navigation data of the plurality of robots includes map fusion and robot pose fusion. The map fusion refers to fusing the local maps of the plurality of robots into a global map to obtain a global map of each frame. The robot pose fusion refers to determining the poses of the plurality of robots in the global map of each frame.
[0086] Exemplarily, the historical navigation data of each frame of the plurality of robots is aligned, the local maps of N frames of the plurality of robots are fused to obtain a global map of each frame. According to the historical navigation data of other robots, the poses of the other robots in the global map are determined according to the pose of the first target robot in each frame to obtain the pose of the corresponding agent of each robot in the global map.
[0087] The first target robot is not particularly specified as a certain robot in the plurality of robots, and the first target robot can be any one of the plurality of robots. The simulation device selects any one of the plurality of robots as the first target robot, and fuses the historical navigation data of the plurality of robots from the perspective of the first target robot as the ego view.
[0088] The simulation device can fuse from the first frame of historical navigation data in the target time period according to the chronological order. The local maps of each frame of the plurality of robots can or can not have intersections. When fusing the maps, the simulation device can use the local maps of multiple frames according to the set fusion algorithm to obtain a global map of each frame. The global map is a complete map that is continuous in space, and the global map includes the local map of each robot.
[0089] Exemplarily, the fusion algorithm includes but is not limited to weighted average, maximum likelihood estimation, Bayesian filtering, etc. The simulation device obtains a global map of each frame through map fusion.
[0090] When fusing the poses of the plurality of robots, the pose of the first target robot is taken as the ego view, and the pose and time of the first target robot are taken as the initial state or reference to determine the pose of other robots relative to the first target robot at the same time. The pose of the other robots relative to the first target robot is the pose of the other robots in the global map.
[0091] The layout map of each robot can contain some dynamic obstacles that change in real time. The dynamic obstacles are equivalent to static obstacles. For example, for a robot, pedestrians, animals, vehicles, etc. in the local map are dynamic obstacles, and walls, doors, columns, furniture, etc. in the local map are static obstacles. Dynamic obstacles are moving, so the positions of dynamic obstacles in the global map change dynamically in each frame after map fusion, while the positions of static obstacles are fixed.
[0092] Through map fusion and robot pose fusion, the global map changes and the state changes of the robot itself can be obtained within any frame number starting from any time point covered by the navigation log data and with any robot as the subjective view.
[0093] S104, multi-agent navigation is performed in the multi-agent environment to perform a multi-agent task.
[0094] In one application scenario, the multi-agent environment is used to simulate and evaluate a new navigation strategy, also referred to as a new navigation algorithm. For example, the new navigation strategy is used to replace the historical navigation strategy of a target agent corresponding to a second target robot in a plurality of robots, and the agents corresponding to the plurality of robots are used to reason in the multi-agent environment to obtain a reasoning result, which includes a new navigation trajectory of the target agent and navigation trajectories of other agents. The historical navigation strategy is the navigation strategy used by the second target robot to generate historical navigation data.
[0095] In this embodiment, the multi-agent environment can simulate the simultaneous operation of a plurality of robots, i.e., the agents corresponding to the plurality of robots simultaneously reason in the multi-agent environment. The agent corresponding to the robot can be understood as a virtual robot that is used to replace the physical robot to operate in the multi-agent environment. The reasoning of the agent in the multi-agent environment can be understood as the operation of the agent in the environment according to the navigation strategy. The agent needs to constantly calculate the next path and execute the path during the operation. The process of calculating and executing the path by the agent once can be understood as one reasoning.
[0096] In this embodiment, the navigation simulation of the plurality of robots is performed from the perspective of the agent corresponding to the second target robot. The second target robot is any one of the plurality of robots, i.e., the navigation strategy of any agent corresponding to the robot can be replaced.
[0097] The new navigation strategy refers to a newly generated navigation strategy used for evaluation, and the historical navigation strategy refers to the navigation strategy used by each robot to form the navigation log data or the historical navigation data.
[0098] The navigation strategy (a new navigation strategy or a historical navigation strategy) can be a rule-based navigation algorithm, a navigation algorithm obtained based on reinforcement learning (RL), a navigation algorithm obtained based on imitation learning (IL), or a navigation algorithm obtained in another manner. Embodiments of the present application do not limit the manner of obtaining the navigation strategy.
[0099] The navigation strategy is used to assist the agent in planning a navigation path, so that the agent reaches the destination as soon as possible, or so that the agent avoids obstacles as much as possible, or so that the agent is as far away from the obstacles as possible.
[0100] The use of the new navigation strategy to replace the historical navigation strategy of the target agent corresponding to the second target robot includes: associating the target agent corresponding to the second target robot with the new navigation strategy. After the association, the target agent will use the new navigation strategy for navigation.
[0101] In this embodiment, there are multiple agents in the multi-agent environment. In addition to the target agent corresponding to the second target robot as the main perspective, other agents also interact with the multi-agent environment. The interaction of the other agents with the multi-agent environment includes two modes: (1) a feedback-free mode, in which the other agents use the historical navigation strategy to reason in the multi-agent environment; and (2) a feedback mode, in which the other agents use the new navigation strategy to reason in the multi-agent environment.
[0102] The target agent forms a new navigation trajectory by reasoning in the multi-agent environment, and evaluates the new navigation strategy according to the new navigation trajectory and the historical navigation trajectory of the target agent. The starting point and the destination of the new navigation trajectory and the historical navigation trajectory are the same, and the historical navigation trajectory of the target agent can be obtained according to the navigation log data or the historical navigation data of the second target robot.
[0103] Optionally, the positions of the agents corresponding to the multiple robots in the multi-agent environment are displayed in real time during the reasoning process, so that the tester can intuitively and in real time understand the positions of the agents.
[0104] In another application scenario, a reinforcement learning method or an imitation learning method is used to train a to-be-trained navigation strategy in the multi-agent environment. The to-be-trained navigation strategy can be the new navigation strategy described above or another navigation strategy. Exploring a new navigation strategy through reinforcement learning or imitation learning can improve the intelligence of the navigation algorithm.
[0105] Reinforcement learning consists of two parts: an agent and an environment. During the reinforcement learning process, the agent and the environment interact with each other. After the agent obtains a state in the environment, it outputs an action based on the state. The action is executed in the environment, and the environment outputs the next state and the reward brought by the action according to the action taken by the agent. The purpose of the agent is to obtain as much reward as possible from the environment.
[0106] The environment is used to provide state information and reward feedback, and is an external system affected by the action of the agent. The state is information describing the current condition of the environment, the action is an operation performed by the agent in the environment, and the reward is an evaluation of the action performed by the agent by the environment, which is a scalar value.
[0107] In a navigation scenario, the action is to move to a specified position or to avoid a certain obstacle and then move to a specified position, the state of the environment can be the change of the environment, and the state of the environment changes as the position of the agent changes during the movement of the agent. The environment calculates the reward corresponding to the action each time the agent performs a movement operation.
[0108] The environment can start simulation exploration with any state as an initial value. For example, it is assumed that the agent walks randomly in a multi-agent environment. If the agent encounters an obstacle each time it walks, the environment performs a penalty, and a negative reward can be generated as a penalty. If the agent does not encounter an obstacle, the environment generates a positive reward. During the exploration process, the trained navigation strategy is updated to maximize the expected reward.
[0109] Imitation training is to train the agent to replicate the continuous actions of an expert, so as to achieve the purpose of imitation. In the embodiment of the present application, the continuous actions of the expert are continuous historical actions corresponding to the historical navigation trajectories of the robots. In an exemplary manner, a neural network is initialized, and in the inference process, the neural network is used to infer the next action, but the action is not executed according to the inferred action, but is executed according to the historical action in the historical navigation trajectory. When updating the parameters of the neural network, the parameters of the neural network are updated to be as close to the historical action as possible, so as to achieve the goal of imitation learning. The neural network trained through imitation learning is a navigation strategy close to the historical navigation strategy.
[0110] In another application scenario, in a multi-agent environment, the agents corresponding to the plurality of robots perform navigation playback according to the historical navigation trajectories of the plurality of robots, that is, the agents corresponding to the plurality of robots perform the historical actions of the plurality of robots in the multi-agent environment according to the historical actions corresponding to the historical navigation trajectories of the plurality of robots, to realize playback of the navigation process of the plurality of robots. The reproducibility of the historical state of the plurality of robots is realized through the navigation playback, so as to provide a fair benchmark for effect evaluation of different methods, and a tester can find a bad case according to the historical state of the plurality of robots, and perform problem troubleshooting, strategy optimization, and the like.
[0111] In this embodiment, navigation log data of a plurality of robots in a target time period is obtained, N frames of historical navigation data of each robot are obtained by analyzing the navigation log data, the historical navigation data including robot poses, local maps, and navigation planning paths; a first target robot is determined from the plurality of robots, and the historical navigation data of the plurality of robots is fused to obtain a multi-agent environment, taking the first target robot as a subjective view, the multi-agent environment including N frames of global maps and poses of a plurality of agents corresponding to the plurality of robots in the N frames of global maps, and multi-agent navigation is performed in the multi-agent environment to execute a multi-agent task. The method constructs a multi-agent environment of a plurality of robots according to actual navigation logs of a large number of robots, the multi-agent environment is generated based on real data, so that the constructed multi-agent environment is more realistic, and thus the execution result of the multi-agent task based on the multi-agent environment is more accurate.
[0112] Figure 2 A flowchart of a navigation method in a multi-agent environment provided by Embodiment Two of the present application is mainly used for simulating and evaluating a new navigation strategy, as shown in FIG. 2, the method provided by this embodiment includes the following steps: Figure 2
[0113] S201, navigation log data of a plurality of robots in a target time period is obtained.
[0114] S202, N frames of historical navigation data of each robot are obtained by analyzing the navigation log data, the historical navigation data including robot poses, local maps, and navigation planning paths.
[0115] S203, a first target robot is determined from the plurality of robots, and the historical navigation data of the plurality of robots is fused to obtain a multi-agent environment, taking the first target robot as a subjective view.
[0116] The multi-agent environment includes N frames of global maps and poses of a plurality of agents in the N frames of global maps, wherein each robot corresponds to an agent.
[0117] The specific implementation of steps S201-S203 can refer to the related description of Embodiment 1, which will not be repeated here.
[0118] S204, using the new navigation strategy to replace the historical navigation strategy of the target agent corresponding to the second target robot in the plurality of robots, using the initial value of the target agent as the input of the new navigation strategy, using the new navigation strategy to perform M-step reasoning in the multi-agent environment to obtain a new navigation trajectory of the target agent.
[0119] The initial value of the target agent is used as the input of the new navigation strategy, and the initial value includes a reasoning start time, a reasoning start position, and a destination. The target agent starts running from the reasoning start position, and the start time of the running is the reasoning start time. The target agent uses the new navigation strategy for navigation during the running.
[0120] The reasoning start time can be any one of the time points in the target time period. The time point can be represented by a specific time value, or can be represented by a frame number corresponding to the time point. For example, the target time period is 6:00-8:00 on April 25, and the start time point is any one of the time points in 6:00-8:00. Assuming that 2000 frames of historical navigation data are collected in 6:00-8:00, the 2000 frames of historical navigation data can be numbered in chronological order. The reasoning start time can be determined according to the frame number.
[0121] The target agent can perform reasoning at a predetermined time interval, for example, once per second or once per frame.
[0122] The target agent uses the new navigation strategy and the multi-agent environment to perform one-step reasoning each time, obtains the navigation position after each step of reasoning, and controls the target agent to move to the navigation position. The navigation positions after M-step reasoning form the new navigation trajectory of the target agent.
[0123] Assuming that the target agent reasons once per frame, the target agent starts reasoning at the first frame, and the pose of the target agent at the first frame, the frame number, and the destination are used as the input of the new navigation strategy. The new navigation strategy reasons to obtain the navigation position of the target agent at the second frame, and controls the target agent to move to the navigation position at the second frame. Then, the navigation position at the second frame, the frame number, and the destination are used as the input of the next step of reasoning to obtain the navigation position of the target agent at the third frame. In this way, the navigation position of the target agent at each frame is obtained. The navigation position of the target agent at each frame refers to the position of the target agent in the global map of each frame of the multi-agent environment.
[0124] S205, controlling other agents to run in the multi-agent environment according to the historical navigation trajectory of the corresponding robot.
[0125] wherein the other agents are agents corresponding to the remaining robots in the plurality of robots other than the second target robot. Assuming that navigation log data of 10 robots are obtained, the multi-agent environment can simultaneously simulate navigation of the 10 robots, and assuming that only the historical navigation strategy of one robot is replaced, there are 9 other agents in addition to the main perspective agent in the multi-agent environment, and for each of the 9 other agents, each agent is controlled to run in the multi-agent environment according to the historical navigation trajectory of the corresponding robot.
[0126] The interaction mode of the other agents and the multi-agent environment in this embodiment is a non-feedback mode, that is, the other agents use the historical navigation strategy to reason in the multi-agent environment, and the other agents strictly run according to the historical navigation trajectory and do not produce new feedback to the behavior of the agent of the main perspective.
[0127] For example, agent 1 is the main perspective, and in the historical navigation data, agent 1 and agent 2 do not collide at time t, after agent 1 uses the new navigation strategy to reason, agent 1 moves according to the navigation position planned by the new navigation strategy at time t, and agent 2 still moves according to the historical navigation trajectory, resulting in a collision between agent 1 and agent 2 at time t. In the non-feedback mode, agent 2 does not produce new feedback to the new collision behavior, that is, the new collision behavior does not change the historical navigation trajectory of agent 2, and agent 2 still runs according to the historical navigation trajectory.
[0128] S206, in the reasoning process, the following contents are displayed in the multi-agent environment in real time: the new position, the historical position and the new planned path of the target agent, and the historical position of the other agents.
[0129] wherein the new position is a navigation position obtained by reasoning according to the new navigation strategy, the historical position is a position indicated by the historical navigation data at the same time, and the new planned path is a path obtained by reasoning according to the new navigation strategy.
[0130] The simulation device can visually display the dynamically changing global map and the real-time changing positions of all robots in a frame-by-frame playing mode, so as to facilitate the test personnel to intuitively and immediately understand the positions of the robots. By comparing the new position and the historical position of the target agent at the same time, the test personnel can intuitively see the difference between the reasoning results of the new navigation strategy and the historical navigation strategy.
[0131] For example, the new position of the target agent can be represented by a red rectangle, the historical position of the target agent can be represented by a pink rectangle, the historical position of the other agents can be represented by a green rectangle, and the new planned path of the target agent can be represented by a blue thick line.
[0132] Optionally, the identifier or number of the agent is displayed at the position of each agent, so as to facilitate the user to distinguish different agents.
[0133] It can be understood that this is only an example, and the new position of the target agent, the historical position of the target agent, and the historical position of other agents can be distinguished in other manners.
[0134] Reference Figure 3 , Figure 3 A display schematic diagram of multi-agent simulation navigation is shown in FIG. 1, which includes four robot agents in the multi-agent environment. The agent 1 is the target agent, i.e., the ego agent. The new position, the historical position, and the new planning path of the agent 1, and the historical position of the other three agents are displayed in the multi-agent environment. Figure 3
[0135] S207, comparing the new navigation trajectory of the target agent with the historical navigation trajectory to evaluate the new navigation strategy.
[0136] The new navigation trajectory is a path planned according to the new navigation strategy, and the historical navigation trajectory is a path planned according to the historical navigation strategy. The starting point and the destination of the new navigation trajectory and the historical navigation trajectory are the same. By comparing the new navigation trajectory with the historical navigation trajectory, the new navigation strategy is evaluated.
[0137] For example, the navigation time, the path length, and the collision times of the new navigation trajectory and the historical navigation trajectory are compared to evaluate the new navigation strategy. The navigation time is the completion time of the navigation task, i.e., the time required from the starting point to the destination. The path length is the distance of the road along which the navigation trajectory passes.
[0138] In an implementation manner, when the navigation time corresponding to the new navigation trajectory is less than the navigation time corresponding to the historical navigation trajectory, and the path length corresponding to the new navigation trajectory is less than or similar to the path length corresponding to the historical navigation trajectory, it is considered that the new navigation strategy is better than the historical navigation strategy.
[0139] In this embodiment, the effect indicators of single agents can be compared more fairly by using the feedback-free mode. Optionally, the historical navigation strategies of multiple agents can be replaced in sequence, and the new navigation strategy is evaluated according to the reasoning results of the multiple agents.
[0140] In this embodiment, the historical navigation strategy of the target agent is replaced by a new navigation strategy, the target agent uses the new navigation strategy to perform M-step reasoning in the multi-agent environment, and a new navigation trajectory of the target agent is obtained. Meanwhile, the other agents run in the multi-agent environment according to the corresponding historical navigation trajectory of the robot. The new navigation strategy is evaluated by comparing the new navigation trajectory and the historical navigation trajectory of the target agent. This method can replace the historical navigation strategy of any robot with any new navigation strategy for simulation evaluation, making the navigation simulation of multiple robots more flexible.
[0141] Figure 4 The flowchart of the multi-robot navigation simulation method provided in Embodiment Three of the present application is mainly used for simulating and evaluating a new navigation strategy. As shown in Figure 4 The method provided in this embodiment includes the following steps:
[0142] S301, obtaining navigation log data of a plurality of robots in a target time period.
[0143] S302, analyzing the navigation log data to obtain N frames of historical navigation data of each robot, the historical navigation data including robot pose, local map and navigation planning path.
[0144] S303, determining a first target robot from the plurality of robots, and fusing the historical navigation data of the plurality of robots to obtain a multi-agent environment from the perspective of the first target robot.
[0145] The multi-agent environment includes N frames of global maps and poses of the plurality of agents in the N frames of global maps, wherein each robot corresponds to an agent.
[0146] The specific implementation of steps S301-S303 is described in Embodiment One, which will not be repeated here.
[0147] S304, replacing the historical navigation strategy of a target agent corresponding to a second target robot in the plurality of robots with a new navigation strategy, taking the initial value of the target agent as the input of the new navigation strategy, and using the new navigation strategy to perform M1-step reasoning in the multi-agent environment to obtain a new navigation trajectory of the target agent.
[0148] The initial value of the target agent includes reasoning start time, reasoning start position and destination. The specific implementation of step S304 is described in Embodiment Two, which will not be repeated here.
[0149] S305, replace the historical navigation strategy of the other agent with the new navigation strategy, take the initial value of the other agent as the input of the new navigation strategy, use the new navigation strategy to perform M2-step reasoning in the multi-agent environment, and obtain a new navigation trajectory of the other agent.
[0150] The initial value of the other agent includes a reasoning start time, a reasoning start position, and a destination. The reasoning start times of the target agent and the other agent can be the same or different. For example, agent 1 is the target agent, agent 1 starts reasoning at frame 0, agent 2 starts reasoning at frame 10, and agent starts reasoning at frame 50.
[0151] M1 and M2 can be the same or different, and the reasoning step number is related to the reasoning start position and the destination of the agent.
[0152] In this embodiment, the interaction mode of the other agent with the multi-agent environment is a feedback mode, that is, the other agent uses the new navigation strategy to reason in the multi-agent environment. In this embodiment, the agents of the plurality of robots all use the new navigation strategy to simultaneously perform simulated navigation in the multi-agent environment, so that the new navigation strategy can be evaluated in the multi-agent case.
[0153] The feedback mode and the non-feedback mode each have advantages and disadvantages. The non-feedback mode can more fairly compare the effect indicators of single agents, but ignores the problem of possible game between agents. The feedback mode can effectively evaluate the historical navigation strategy in the multi-agent case, but the distorted strategy adopted by the agent can cause the cumulative error of the entire environment to become large, thereby affecting the evaluation of the historical navigation strategy by the single agent.
[0154] S306, in the reasoning process, the following contents are displayed in the multi-agent environment in real time: the new position, the historical position, and the new planned path of the target agent, and the new position and the historical position of the other agent.
[0155] The new position is the position obtained by reasoning according to the new navigation strategy, the historical position is the position indicated by the historical navigation data at the same time, and the new planned path is the path obtained by reasoning according to the new navigation strategy.
[0156] Optionally, in order to avoid too much cluttered display, the historical position of the other agent can also not be displayed.
[0157] S307, compare the new navigation trajectory of the target agent with the historical navigation trajectory to evaluate the new navigation strategy and obtain a first evaluation result; and compare the new navigation trajectory of the other agent with the historical navigation trajectory to evaluate the new navigation strategy and obtain a second evaluation result.
[0158] In this embodiment, for multiple agents in a multi-agent environment, the new navigation trajectory of each agent is compared with the historical navigation trajectory of the agent, and the new navigation strategy is evaluated from the navigation time, path length, collision times and other indicators of the new navigation trajectory and the historical navigation trajectory.
[0159] S308, evaluating the new navigation strategy according to the first evaluation result and the second evaluation result.
[0160] The evaluation results of the multiple agents are combined to evaluate the new navigation strategy, which considers the game problem that may be generated between the multiple agents and is closer to the real scene.
[0161] In this embodiment, the historical navigation strategy of each agent corresponding to the multiple robots in the multi-agent environment is replaced with the new navigation strategy, the multiple agents simultaneously use the new navigation strategy to perform multi-step reasoning in the multi-agent environment, the new navigation trajectory of each agent is obtained, the new navigation strategy is evaluated by comparing the new navigation trajectory of each agent with the historical navigation trajectory of the agent, and the new navigation strategy is comprehensively evaluated according to the evaluation results of the multiple agents.
[0162] To better implement the navigation method in the multi-agent environment of the embodiments of the present application, the embodiments of the present application also provide a navigation device in a multi-agent environment. Figure 5 As shown in the structural schematic diagram of the navigation device in the multi-agent environment provided by Embodiment Four of the present application, Figure 5 The navigation device 100 in the multi-agent environment can include:
[0163] The acquisition module 11 is configured to acquire navigation log data of multiple robots in a target time period.
[0164] The analysis module 12 is configured to analyze the navigation log data to obtain N frames of historical navigation data of each robot, wherein the historical navigation data includes robot pose, local map and navigation planning path.
[0165] The fusion module 13 is configured to determine a first target robot from the multiple robots, take the first target robot as a subjective view, and fuse the historical navigation data of the multiple robots to obtain a multi-agent environment, wherein the multi-agent environment includes N frames of global maps and poses of the multiple agents in the N frames of global maps, and each robot corresponds to an agent.
[0166] The execution module 14 is configured to perform multi-agent navigation in the multi-agent environment to execute a multi-agent task.
[0167] In an implementation manner, the execution module 14 is specifically configured to:
[0168] replace a historical navigation strategy of a target agent corresponding to a second target robot in the plurality of robots with a new navigation strategy, use the agents corresponding to the plurality of robots to reason in the multi-agent environment, and obtain a reasoning result, wherein the reasoning result includes a new navigation trajectory of the target agent, and the historical navigation strategy is a navigation strategy used by the second target robot to generate the historical navigation data.
[0169] In an implementation manner, the execution module 14 is specifically configured to:
[0170] use the initial value of the target agent as an input of the new navigation strategy, perform M-step reasoning in the multi-agent environment using the new navigation strategy, and obtain a new navigation trajectory of the target agent, wherein the initial value of the target agent includes a reasoning start time, a reasoning start position, and a destination.
[0171] control other agents to run in the multi-agent environment according to historical navigation trajectories of corresponding robots, wherein the other agents are agents corresponding to remaining robots in the plurality of robots except the second target robot.
[0172] In an implementation manner, the device further includes:
[0173] a display module configured to display, in real time in the multi-agent environment during the reasoning process, the following contents: a new position, a historical position, and a new planned path of the target agent, and a historical position of the other agents, wherein the new position is a navigation position obtained according to the new navigation strategy, the historical position is a position indicated by the same time historical navigation data, and the new planned path is a path obtained according to the new navigation strategy.
[0174] In an implementation manner, the execution module 14 is specifically configured to:
[0175] use the initial value of the target agent as an input of the new navigation strategy, perform M1-step reasoning in the multi-agent environment using the new navigation strategy, and obtain a new navigation trajectory of the target agent, wherein the initial value of the target agent includes a reasoning start time, a reasoning start position, and a destination.
[0176] replace historical navigation strategies of other agents with the new navigation strategy, use initial values of the other agents as inputs of the new navigation strategy, perform M2-step reasoning in the multi-agent environment using the new navigation strategy, and obtain new navigation trajectories of the other agents, wherein the initial values of the other agents include reasoning start times, reasoning start positions, and destinations.
[0177] In an implementation manner, the device further includes:
[0178] a display module configured to display, in real time during the reasoning process, the following in the multi-agent environment: a new position of the target agent, a historical position of the target agent, and a new planned path of the target agent, a new position of the other agents, and a historical position of the other agents, wherein the new position is a position obtained by reasoning according to the new navigation strategy, the historical position is a position indicated by historical navigation data at the same time, and the new planned path is a path obtained by reasoning according to the new navigation strategy.
[0179] In an implementation manner, the execution module 14 is specifically configured to train the to-be-trained navigation strategy in the multi-agent environment by using a reinforcement learning method or an imitation learning method.
[0180] In an implementation manner, the execution module 14 is specifically configured to perform, in the multi-agent environment, navigation playback of the agents corresponding to the plurality of robots according to historical navigation trajectories of the plurality of robots.
[0181] In an implementation manner, the fusion module 13 is specifically configured to:
[0182] align the historical navigation data of each frame of the plurality of robots;
[0183] fuse the local maps of N frames of the plurality of robots to obtain a global map of each frame;
[0184] reference to the pose of the first target robot at each frame, determine poses of the other robots in the global map according to the historical navigation data of the other robots.
[0185] In an implementation manner, the new navigation strategy is a strategy learned by using a reinforcement learning method or an imitation learning method.
[0186] In an implementation manner, the apparatus further includes:
[0187] a display module configured to display, in real time during the reasoning process, positions of the agents corresponding to the plurality of robots in the multi-agent environment.
[0188] In an implementation manner, the apparatus further includes:
[0189] an evaluation module configured to evaluate the new navigation strategy by comparing the new navigation trajectory of the target agent and the historical navigation trajectory of the target agent.
[0190] In an implementation manner, the apparatus further includes an evaluation module configured to:
[0191] The new navigation strategy is evaluated by comparing the new navigation trajectory of the target agent with the historical navigation trajectory, to obtain a first evaluation result;
[0192] The new navigation strategy is evaluated by comparing the new navigation trajectory of the other agent with the historical navigation trajectory, to obtain a second evaluation result;
[0193] The new navigation strategy is evaluated according to the first evaluation result and the second evaluation result.
[0194] The device of the embodiment can be used to execute any method described in Embodiment One to Embodiment Three, and the specific implementation manner refers to the description of the method embodiment, which will not be described here.
[0195] It should be understood that the device embodiment and the method embodiment can correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, it will not be described here.
[0196] The device 100 of the embodiment of the application is described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions of software, or in the form of a combination of hardware and software modules. Specifically, each step of the method embodiment in the embodiment of the application can be completed by integrated logic circuits of hardware in a processor and / or instructions in the form of software, and the steps of the method disclosed in the embodiment of the application can be directly embodied as hardware code processor execution completion, or executed by a combination of hardware and software modules in the code processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiment.
[0197] The embodiment of the application also provides an electronic device. Figure 6 A structural schematic diagram of the electronic device provided in Embodiment Five of the application is shown in FIG. 3, which can include: Figure 6
[0198] The memory 31 is used to store a computer program and transmit the program code to the processor 32. In other words, the processor 32 can call and run the computer program from the memory 31 to implement the navigation simulation method of multiple robots provided in the embodiment of the application.
[0199] For example, the processor 32 can be used to execute the navigation simulation method of multiple robots provided in the above method embodiment according to the instructions in the computer program.
[0200] In some embodiments of the application, the processor 32 can include, but is not limited to, a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, and the like.
[0201] In some embodiments of the application, the memory 31 includes, but is not limited to, a volatile memory and / or a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0202] In some embodiments of the application, the computer program can be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method provided by the application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the electronic device.
[0203] As Figure 6As shown, the electronic device 300 can further include a transceiver 33, a display screen 34, etc., and the processor 32 is electrically connected with the transceiver 33 and the display screen 34, respectively.
[0204] The processor 32 can control the transceiver 33 to communicate with other devices, specifically, can send information or data to other devices, or receive information or data sent by other devices. The transceiver 33 can include a transmitter and a receiver. The transceiver 33 can further include an antenna, and the number of the antenna can be one or more.
[0205] The display screen 34 can be used to display a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface. The display screen 34 can be a touch display screen, which can include a display panel and a touch panel. The display panel can be used to display information input by a user or provided to a user and various graphical user interfaces of a computer device, which can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations (such as operations of a user using a finger, a stylus, or any suitable object or accessory on or near the touch panel) of a user on or near the touch panel, and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel can include two parts, a touch detection device and a touch controller. The touch detection device detects the touch position of a user and detects signals generated by a touch operation, and transmits the signals to the touch controller; the touch controller receives touch information from the touch detection device, and converts the touch information into touch coordinates, and then sends the touch coordinates to the processor 32, and can receive commands from the processor 32 and execute the commands. The touch panel can cover the display panel, and when the touch panel detects a touch operation on or near the touch panel, the touch panel transmits the touch operation to the processor 32 to determine the type of the touch event, and then the processor 32 provides corresponding visual output on the display panel according to the type of the touch event.
[0206] It can be understood that the electronic device structure shown in the foregoing embodiments does not constitute a limitation on the electronic device, and can include more or fewer components than those shown, or combine certain components, or different component arrangements. For example, the electronic device 300 can further include a camera module, a wireless fidelity (WIFI) module, a positioning module, a Bluetooth module, a display, a controller, etc., which will not be described here. Figure 6
[0207] It should be understood that various components in the electronic device are connected through a bus system, wherein the bus system includes, in addition to a data bus, a power supply bus, a control bus, and a state signal bus.
[0208] The application further provides a computer storage medium, which stores a computer program. When the computer program is executed by a computer, the computer is enabled to perform the method of the method embodiments.
[0209] The application further provides a computer program product, which includes a computer program stored in a computer readable storage medium. A processor of an electronic device reads the computer program from the computer readable storage medium. The processor executes the computer program, so that the electronic device performs the corresponding procedures of the navigation simulation method of the multi-robot provided by the method embodiments. For brevity, details are not repeated here.
[0210] In several embodiments provided in the application, it should be understood that the disclosed system, device and method can be implemented by other manners. For example, the device embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0211] The modules illustrated as separate components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place or can be distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. For example, the functional modules in each embodiment of the application can be integrated into one processing module, or each module can be physically separated, or two or more modules can be integrated into one module.
[0212] The above is only a specific implementation of the application, but the protection scope of the application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the application, which should be covered within the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims.
Claims
1. A method of navigation in a multi-agent environment, the method comprising: The method comprises: obtaining navigation log data of a plurality of robots in a target time period; parsing the navigation log data to obtain N-frame historical navigation data of each robot, the historical navigation data comprising robot pose, local map and navigation planning path; determining a first target robot from the plurality of robots, and fusing the historical navigation data of the plurality of robots to obtain a multi-agent environment from the perspective of the first target robot, the multi-agent environment comprising N-frame global map and pose of a plurality of agents in the N-frame global map, wherein each robot corresponds to an agent; performing multi-agent navigation in the multi-agent environment to execute a multi-agent task.
2. The method of claim 1, wherein, The method of performing multi-agent navigation in the multi-agent environment to execute a multi-agent task comprises: replacing a historical navigation strategy of a target agent corresponding to a second target robot in the plurality of robots with a new navigation strategy, reasoning in the multi-agent environment using agents corresponding to the plurality of robots to obtain a reasoning result, the reasoning result comprising a new navigation trajectory of the target agent, the historical navigation strategy being a navigation strategy used by the second target robot to generate the historical navigation data.
3. The method of claim 2, wherein, The method of reasoning in the multi-agent environment using agents corresponding to the plurality of robots to obtain a reasoning result comprises: inputting initial values of the target agent into the new navigation strategy, performing M-step reasoning in the multi-agent environment using the new navigation strategy to obtain a new navigation trajectory of the target agent, the initial values of the target agent comprising reasoning start time, reasoning start position and destination; controlling other agents to run in the multi-agent environment according to historical navigation trajectories of corresponding robots, wherein the other agents are agents corresponding to remaining robots in the plurality of robots except the second target robot.
4. The method of claim 3, wherein, The method further comprises: displaying the following contents in the multi-agent environment in real time during the reasoning process: new position, historical position and new planning path of the target agent, and historical position of the other agents, wherein the new position is a navigation position obtained according to the new navigation strategy, the historical position is a position indicated by the same time historical navigation data, and the new planning path is a path obtained according to the new navigation strategy.
5. The method of claim 2, wherein, The method of reasoning in the multi-agent environment using agents corresponding to the plurality of robots to obtain a reasoning result comprises: inputting initial values of the target agent into the new navigation strategy, performing M1-step reasoning in the multi-agent environment using the new navigation strategy to obtain a new navigation trajectory of the target agent, the initial values of the target agent comprising reasoning start time, reasoning start position and destination; replace a history navigation strategy of another agent with the new navigation strategy, use initial values of the another agent as inputs of the new navigation strategy, perform M2-step reasoning in the multi-agent environment using the new navigation strategy, and obtain a new navigation trajectory of the another agent, wherein the initial values of the another agent include a reasoning start time, a reasoning start position, and a destination.
6. The method of claim 5, wherein, Further comprising: displaying the following in real time in the multi-agent environment during the reasoning process: a new position, a history position, and a new planned path of the target agent, and a new position and a history position of the another agent, wherein the new position is a position obtained according to the new navigation strategy, the history position is a position indicated by the same time history navigation data, and the new planned path is a path obtained according to the new navigation strategy.
7. The method of claim 1, wherein, The multi-agent navigation in the multi-agent environment to perform the multi-agent task comprises: training a to-be-trained navigation strategy in the multi-agent environment using a reinforcement learning method or an imitation learning method.
8. The method of claim 1, wherein, The multi-agent navigation in the multi-agent environment to perform the multi-agent task comprises: In the multi-agent environment, the agents corresponding to the plurality of robots perform navigation playback according to the history navigation trajectories of the plurality of robots.
9. The method according to any one of claims 1-7, characterized in that, Fusing the history navigation data of the plurality of robots to obtain a multi-agent environment from a subjective perspective of the first target robot, comprising: aligning the history navigation data of each frame of the plurality of robots; fusing local maps of N frames of the plurality of robots to obtain a global map of each frame; determining the pose of the other robots in the global map according to the history navigation data of the other robots, with the pose of the first target robot in each frame as a reference.
10. The method of any one of claims 2-6, wherein, The new navigation strategy is a strategy learned through a reinforcement learning method or an imitation learning method.
11. The method of any one of claims 2-6, wherein, Further comprising: displaying the positions of the agents corresponding to the plurality of robots in the multi-agent environment in real time during the reasoning process.
12. The method of claim 2, wherein, Further comprising: evaluating the new navigation strategy by comparing the new navigation trajectory of the target agent with the history navigation trajectory.
13. The method of claim 5, wherein, Further comprising: evaluating the new navigation strategy by comparing the new navigation trajectory of the target agent with the history navigation trajectory, obtaining a first evaluation result; evaluating the new navigation strategy by comparing the new navigation trajectory of the another agent with the history navigation trajectory, obtaining a second evaluation result; evaluating the new navigation strategy according to the first evaluation result and the second evaluation result.
14. A navigation device in a multi-agent environment, characterized by comprising: an acquisition module configured to acquire navigation log data of a plurality of robots in a target time period; an analysis module configured to analyze the navigation log data to obtain N frames of history navigation data of each robot, wherein the history navigation data includes a robot pose, a local map, and a navigation planning path; an analysis module configured to analyze the navigation log data to obtain N frames of history navigation data of each robot, wherein the history navigation data includes a robot pose, a local map, and a navigation planning path; a fusion module configured to determine a first target robot from the plurality of robots, fuse historical navigation data of the plurality of robots to obtain a multi-agent environment with respect to a first target robot as a subjective view, the multi-agent environment comprising N frames of global maps and poses of the plurality of agents in the N frames of global maps; an execution module configured to perform multi-agent navigation in the multi-agent environment to execute a multi-agent task.
15. An electronic device, comprising: comprising: a processor and a memory, the memory being configured to store a computer program, and the processor being configured to invoke and run the computer program stored in the memory to execute the method of any one of claims 1 to 13.
16. A computer readable storage medium characterized by: a computer program product configured to cause a computer to execute the method of any one of claims 1 to 13. a computer program product configured to cause a computer to execute the method of any one of claims 1 to 13.