Multi-robot collaborative exploration method based on intention reasoning and related equipment

By building an intention reasoning mechanism and hierarchical map fusion in a multi-robot system, the problem that robots cannot effectively share information is solved, more efficient environment exploration and collaboration is achieved, and the exploration ability of multi-robot systems in complex environments is improved.

CN120295303APending Publication Date: 2025-07-11PENG CHENG LAB
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510338183.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the process of collaborative exploration of multiple robots, the robot can only partially observe environmental information and cannot effectively share information, resulting in low communication efficiency and poor collaboration, which limits the system's exploration ability in complex environments.

Method used

By exchanging intention reasoning information between multiple robots, a hierarchical map fusion and intention reasoning mechanism is built, a local observation map is built using the server and global coordinate transformation is carried out, and information encoding and decoding is combined with a variational autoencoder and a recurrent neural network to generate potential spatial representations, realizing intention sharing and dynamic path planning among robots.

Benefits of technology

It improves the exploration efficiency and collaboration capabilities of multi-robot systems in complex environments, ensures the integrity and consistency of environmental information, reduces communication load and path conflicts, and improves the efficiency and effectiveness of collaborative exploration of multi-robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295303A_ABST
    Figure CN120295303A_ABST
Patent Text Reader

Abstract

The invention provides a multi-robot collaborative exploration method based on intention reasoning and related equipment, and the method comprises the steps: obtaining the sensing data of each robot, constructing a local observation map of each robot according to the sensing data of each robot, and estimating the pose information of each robot, obtaining a global observation map of each robot, and fusing the plurality of global observation maps to obtain a global fusion map; each robot extracts a spatial feature map from the global observation map, pre-estimates own detection intention information according to exploration information shared by other robots, and obtains a global exploration target according to fusion of the exploration intention information and the spatial feature map; and each robot generates a target exploration path for the corresponding global exploration target based on the global fusion map, and explores according to the target exploration path, so that information sharing among multiple robots is realized, and the exploration capability of a multi-robot system in a complex environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of multi-robot exploration, and in particular, to a multi-robot collaborative exploration method based on intention reasoning and related devices. Background Art

[0002] During the multi-robot collaborative exploration process, the local observation maps and pose information obtained by each robot through its own sensors in the traditional solution are limited and difficult to comprehensively reflect the state of the entire environment. When attempting to integrate the observation information of multiple robots, due to the lack of an effective information sharing mechanism between multiple robots, it is difficult for robots to transmit key data in a timely manner, resulting in problems such as low communication efficiency and poor collaboration between multiple robots, which limits the exploration ability of the multi-robot system in complex environments. Summary of the Invention

[0003] The embodiments of the present application provide a multi-robot collaborative exploration method based on intention reasoning and related devices, which can achieve information sharing by exchanging intention reasoning information between multiple robots, and thus can effectively solve the problems of low communication efficiency and poor collaboration caused by the fact that robots can only partially observe environmental information and cannot effectively share information during the multi-robot collaborative exploration process, and improve the exploration ability of the multi-robot system in complex environments.

[0004] In a first aspect, the present application provides a multi-robot collaborative exploration method based on intention reasoning, which is applied to a multi-robot collaborative exploration system. The multi-robot collaborative exploration system includes a server and multiple robots. The method includes: the server obtains the sensing data of each robot, and respectively constructs a local observation map of each robot and estimates the pose information of each robot according to the sensing data of each robot; the server respectively performs global coordinate transformation processing based on the historical observation information of each robot and the corresponding local observation map to obtain a global observation map of each robot; the server fuses multiple global observation maps to obtain a global fusion map, and sends the global observation map and the global fusion map to each robot; each robot extracts a spatial feature map from the global observation map, estimates its own exploration intention information according to the exploration information shared by other robots, and fuses the exploration intention information and the spatial feature map to obtain a global exploration target; each robot generates a target exploration path for its corresponding global exploration target based on the global fusion map, and conducts exploration according to the target exploration path.

[0005] In some embodiments, the server obtains the sensing data of each robot, and respectively constructs a local observation map for each robot and estimates the pose information of each robot based on the sensing data of each robot, including: the server obtains the sensing data of each robot, and obtains a plurality of current-time maps and previous-time maps centered on the robot according to the sensing data of each robot, and calculates the pose change of each robot according to the current-time map and the previous-time map; the server respectively performs spatial feature aggregation on the corresponding current-time map and the previous-time map based on the pose change of each robot to obtain the local observation map of each robot.

[0006] In some embodiments, the server respectively performs global coordinate transformation processing on the historical observation information of each robot and the corresponding local observation map to obtain the global observation map of each robot, including: the server respectively obtains a plurality of historical observation maps centered on the robot based on the historical observation information of each robot, and converts the plurality of historical observation maps and the local observation map into maps in the same global coordinate system, and performs exploration boundary clipping processing on the maps in the same global coordinate system to obtain the global observation map of each robot.

[0007] In some embodiments, estimating its own detection intention information according to the exploration information shared by other robots includes: each robot obtains an intention inference module deployed by the server, and encodes the exploration information shared by other robots based on the variational autoencoder in the intention inference module to generate a latent space representation; wherein, the server inputs the exploration information transmitted by other robots into the variational autoencoder, and trains the encoder parameters and decoder parameters of the variational autoencoder with the relative entropy of the variational autoencoder as the optimization target; each robot decodes the latent space representation to obtain its own detection intention information.

[0008] In some embodiments, training the encoder parameters and decoder parameters of the variational autoencoder with the relative entropy of the variational autoencoder as the optimization target includes: the server obtains the approximate latent distribution and approximate posterior distribution output by the variational autoencoder based on the multi-layer perception mechanism, and calculates the relative entropy of the variational autoencoder according to the difference between the approximate latent distribution and the approximate posterior distribution, and trains the encoder parameters and decoder parameters of the variational autoencoder with minimizing the relative entropy as the optimization target.

[0009] In some embodiments, after estimating its own detection intention information based on the exploration information shared by other robots, it also includes: each robot obtains its corresponding hidden state information, and aggregates the hidden state information with the exploration intention information through a preset fully connected neural network to generate a communication vector for each robot; each robot sends its corresponding communication vector as exploration information to other robots, so that the other robots cyclically update their respective hidden state information based on the received exploration information.

[0010] In some embodiments, the global exploration target is obtained based on the fusion of the exploration intention information and the spatial feature map, including: each robot obtains a target exploration area from a discretized grid corresponding to the spatial feature map, and determines the relative position coordinates corresponding to the exploration intention information from the target exploration area; each robot obtains a global exploration target based on the relative position coordinates and the coordinate information of the target exploration area.

[0011] In some embodiments, each robot generates a target exploration path for the corresponding global exploration target based on the global fusion map, and explores according to the target exploration path, including: each robot divides the corresponding global exploration target into local targets based on the global fusion map to obtain a target sequence consisting of multiple local targets; each robot generates a target exploration path according to the target sequence, and explores according to the target exploration path.

[0012] In some embodiments, each robot performs local target division for the corresponding global exploration target based on the global fusion map to obtain a target sequence composed of multiple local targets, including: each robot calculates the shortest planned path from the robot's current position to the global target using a fast marching method based on the fused global map; each robot extracts multiple local targets on the shortest planned path at preset distance intervals to form a target sequence.

[0013] In some embodiments, the exploration according to the target exploration path includes: each robot determines the current local target from the corresponding target exploration path, and parses the visual input information, relative spatial distance and target relative angle corresponding to the current local target through a pre-trained visual encoder; each robot outputs a navigation action instruction based on the corresponding visual input information, relative spatial distance and target relative angle, and performs exploration actions according to the navigation action instruction.

[0014] In a second aspect, the present application provides an electronic device, including: at least one processor; at least one memory for storing at least one program; when the at least one program is executed by the at least one processor, the method for multi-robot collaborative exploration based on intention reasoning described in any one of the embodiments in the first aspect is implemented.

[0015] In a third aspect, the present application provides a computer-readable storage medium storing computer-executable instructions for executing the method for multi-robot collaborative exploration based on intention reasoning described in any one of the embodiments in the first aspect.

[0016] A method for multi-robot collaborative exploration based on intention reasoning and related devices proposed in the present application can construct a hierarchical map fusion and intention reasoning mechanism by exchanging intention reasoning information between multiple robots, so as to improve the exploration efficiency and collaboration ability of the multi-robot system; in the specific implementation process of the present application, the server can respectively construct local observation maps based on the sensing data of each robot, and generate a fused global map through global coordinate transformation and max pooling operations to ensure the integrity and consistency of environmental information; each robot can extract spatial features from the global observation map, and use a variational autoencoder to encode and decode the exploration information of other robots to generate a latent space representation for inferring collaborative intentions, and dynamically fuse intentions and spatial features in combination with a recurrent neural network to generate a global exploration target; on this basis, each robot can plan a target sequence composed of local targets based on the fused global map, execute exploration actions through visual encoding and navigation instructions, and realize implicit transmission and iterative update of intentions between robots through an intention reasoning module, reduce communication load and path conflicts, and at the same time, by using hierarchical map fusion and dynamic path planning, it can ensure that multiple robots efficiently cover unknown areas in a complex environment, so as to solve the problems of low communication efficiency and poor collaboration in traditional methods, and further improve the efficiency and effect of collaborative exploration between multiple robots, and improve the exploration ability of the multi-robot system in a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the method for multi-robot collaborative exploration based on intention reasoning provided in an embodiment of the present application;

[0018] Figure 2 It is a schematic structural diagram of a multi-robot collaborative exploration framework based on intention reasoning in the method for multi-robot collaborative exploration based on intention reasoning provided in an embodiment of the present application;

[0019] Figure 3 It is a flowchart of constructing the local observation map of each robot and estimating the pose information of each robot in the method for multi-robot collaborative exploration based on intention reasoning provided in an embodiment of the present application;

[0020] Figure 4 In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application, it is a flowchart for obtaining the global observation map of each robot;

[0021] Figure 5 In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application, it is a schematic diagram of the framework structure of the multi-robot intention reasoning sensor;

[0022] Figure 6 In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application, it is a flowchart for estimating the detection intention information;

[0023] Figure 7 In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application, it is a flowchart for updating the hidden state information;

[0024] Figure 8 In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application, it is a flowchart for obtaining the global exploration target;

[0025] Figure 9 In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application, it is a flowchart for generating a target exploration path and exploring according to the target exploration path;

[0026] Figure 10 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0027] To make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0028] In some embodiments, although functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the flowchart. Terms such as first and second in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence.

[0029] In addition, unless otherwise clearly defined and limited, the term "connected / linked" should be understood in a broad sense. For example, it can be a fixed connection or a movable connection, or a detachable connection or a non-detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection or can communicate with each other; it can be directly connected or indirectly connected through an intermediate medium.

[0030] In the description of the embodiments of the present application, the descriptions referring to terms such as "one embodiment / embodiment mode", "another embodiment / embodiment mode", "certain embodiments / embodiment modes", "in the above embodiments / embodiment modes", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least two embodiments or embodiment modes disclosed in the present application. In the disclosure of the present application, the schematic expressions of the above terms do not necessarily refer to the same embodiment or embodiment mode. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a sequence different from that in the flowchart.

[0031] Currently, in the process of multi-robot collaborative exploration, the local observation maps and pose information obtained by each robot through its own sensors in the traditional solution are limited and difficult to comprehensively reflect the state of the entire environment. When trying to integrate the observation information of multiple robots, due to the lack of an effective information sharing mechanism between multiple robots, it is difficult for robots to transmit key data in a timely manner, which in turn leads to problems such as low communication efficiency and poor cooperation between multiple robots, restricting the exploration ability of multi-robot systems in complex environments.

[0032] Based on this, the embodiments of the present application provide a multi-robot collaborative exploration method and related devices based on intention reasoning, which can realize information sharing by exchanging intention reasoning information between multiple robots, and can effectively solve the problems that robots can only partially observe environmental information and cannot perform effective information sharing during the multi-robot collaborative exploration process, resulting in low communication efficiency and poor cooperation, and improve the exploration ability of multi-robot systems in complex environments.

[0033] The following further describes the embodiments of the present application with reference to the accompanying drawings.

[0034] Refer to Figure 1 , Figure 1 is a flowchart of a multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application; in some embodiments, the present application provides a multi-robot collaborative exploration method based on intention reasoning, which is applied to a multi-robot collaborative exploration system. The multi-robot collaborative exploration system includes a server and multiple robots. The method at least includes the following steps:

[0035] Step S110, the server obtains the sensing data of each robot, and respectively constructs the local observation map of each robot and estimates the pose information of each robot according to the sensing data of each robot;

[0036] In some embodiments, the multi-robot collaborative exploration system can be understood as a system composed of multiple robots and a server. These robots can work collaboratively in an unknown environment. The server can be used to collect the sensing data of the robots, construct local observation maps, estimate the pose information of the robots, perform global coordinate transformation, fuse global observation maps, and other tasks. The robots then conduct autonomous exploration based on the data processed by the server.

[0037] In some embodiments, the sensing data can be understood as the data obtained by the robots in the environment through various sensors. This data contains the perception information of the robots about the surrounding environment and is the basis for map construction and pose estimation. The local observation map can be understood as a map centered on each robot constructed based on its own sensing data, which reflects the layout of the local environment around the robot and information such as obstacles, and is the robot's preliminary understanding of the surrounding environment. The pose information can be understood as the position and orientation of the robot, that is, the specific coordinate position and orientation information of the robot in the environment, which is crucial for the robot's positioning and path planning in the map.

[0038] Step S120, the server performs global coordinate transformation processing based on the historical observation information of each robot and the corresponding local observation map to obtain the global observation map of each robot.

[0039] In some embodiments, the historical observation information can be understood as the observation data of the robot at past times, including the previously constructed local observation maps, etc. This information helps the server understand the exploration history of the robot and provides a basis for global coordinate transformation. The global coordinate transformation processing can be understood as the process of converting the local observation map centered on the robot to a unified global coordinate system, enabling the map data of different robots to be fused and processed in the same reference system, and providing a unified environmental perception for multi-robot collaborative exploration.

[0040] Step S130, the server fuses the multiple global observation maps to obtain a global fusion map, and sends the global observation map and the global fusion map to each robot.

[0041] In some embodiments, the global observation map can be understood as the observation map of a single robot represented in the global coordinate system after global coordinate transformation, which integrates the robot's historical exploration information and provides a basis for subsequent global fusion map construction and determination of the robot's exploration goals; the global fusion map can be understood as the comprehensive map obtained by the server by fusing the global observation maps of multiple robots, which integrates the observation information of all robots, can more comprehensively reflect the layout of the entire environment, and provides a global reference for the exploration path planning of each robot; the spatial feature map can be understood as the feature information extracted by the robot from the global observation map and generated through technologies such as convolutional neural networks. These features can highlight the key information in the map, such as the positions of obstacles and passable areas, and provide a basis for subsequent intention reasoning and goal determination.

[0042] Step S140: Each robot extracts a spatial feature map from the global observation map, estimates its own detection intention information based on the exploration information shared by other robots, and fuses the exploration intention information and the spatial feature map to obtain a global exploration goal.

[0043] In some embodiments, the exploration information can be understood as the information shared by the robot with other robots during the exploration process, including its own observation data, intention information, etc. This information helps other robots understand its exploration status and intention and achieve effective collaborative exploration; the detection intention information can be understood as the intention of the robot estimated by the intention reasoning module during the exploration process, that is, the possible next exploration direction or goal of the robot. This information is generated based on the shared exploration information of other robots and reflects the collaborative intention among the robots.

[0044] Step S150: Each robot generates a target exploration path based on the global fusion map for its corresponding global exploration goal and explores according to the target exploration path.

[0045] In some embodiments, the global exploration goal can be understood as the goal obtained by the robot by fusing its own detection intention information and the spatial feature map, which is the long-term navigation goal of the robot in the global environment and guides the exploration direction of the robot in the environment; the target exploration path can be understood as the specific exploration route generated by the robot based on the global fusion map and the global exploration goal. This path planning takes into account the obstacles and passable areas in the environment to ensure that the robot can reach the exploration goal efficiently and safely.

[0046] Among them, it can be understood that during the multi-robot collaborative exploration process, due to factors such as the sensor range and communication bandwidth of the robots, only partial environmental information can be observed, and it is difficult to perform effective information sharing, resulting in low communication efficiency and poor collaboration. To solve these problems, this application proposes a multi-robot collaborative exploration method based on intention reasoning and related devices, aiming to improve the exploration efficiency and collaboration ability of the multi-robot system in complex environments. It can construct a hierarchical map fusion and intention reasoning mechanism by exchanging intention reasoning information between multiple robots to enhance the exploration efficiency and collaboration ability of the multi-robot system; in the specific implementation process of this application, the server can respectively construct local observation maps based on the sensing data of each robot, and generate a fused global map through global coordinate transformation and max-pooling operations to ensure the integrity and consistency of environmental information; each robot can extract spatial features from the global observation map, encode and decode the exploration information of other robots using a variational autoencoder, generate a latent space representation to infer the collaborative intention, and dynamically fuse the intention and spatial features in combination with a recurrent neural network to generate a global exploration goal; on this basis, each robot can plan a sequence of goals composed of local goals based on the fused global map, execute exploration actions through visual encoding and navigation instructions, and achieve implicit transmission and iterative update of intentions between robots through the intention reasoning module, reducing communication load and path conflicts. At the same time, by using hierarchical map fusion and dynamic path planning, it can ensure that multiple robots efficiently cover unknown areas in complex environments to solve the problems of low communication efficiency and poor collaboration in traditional methods, thereby improving the efficiency and effect of collaborative exploration between multiple robots and enhancing the exploration ability of the multi-robot system in complex environments.

[0047] Reference Figure 2 , Figure 2 FIG. is a schematic structural diagram of a multi-robot collaborative exploration framework based on intention reasoning in the multi-robot collaborative exploration method provided by an embodiment of this application. This figure shows the system architecture and information flow of the multi-robot collaborative exploration method based on intention reasoning. Among them, corresponding to the above steps S110 to S150, for robots 1 to N, each robot can input specific information such as the global map, position, trajectory, and goal centered on itself, and calculate the relative coordinates and global goals according to its own position and goal.

[0048] In the CNN feature extractor, a convolutional neural network can be used to extract spatial features from the input global map to generate feature maps, and these feature maps are used to represent key information in the environment, such as obstacle positions and passable areas.

[0049] In the intention reasoning module, each robot receives the exploration information shared by the other N - 1 robots, encodes and decodes the received information through a variational auto - encoder to generate a latent space representation, infers the intentions of the other robots, aggregates the robot's own hidden state information and intention reasoning information to generate a communication vector, and sends it to other robots;

[0050] In the RNN (Recurrent Neural Network), the RNN can be used to update the robot's hidden state information, fuse intention reasoning information and spatial features, and process the fused information to generate a global exploration goal;

[0051] In the action generator, the processed features can be mapped to the action space to determine the next exploration area and specific location of the robot. The global goal is input into the local planner module. The robot uses the rapid marching method to plan a path to achieve the long - term goal in the fused global map, generates a sequence of short - term sub - goals, and combines the obtained RGB images of the corresponding environmental information. Given a short - term sub - goal, the local policy module outputs navigation actions based on visual input, relative spatial distance, and relative angle to the sub - goal, realizing the efficient cooperative exploration of multiple robots in an unknown environment.

[0052] It can be understood that this application can obtain sensing data from each robot, which includes the robot's observation information about the surrounding environment. Based on the sensing data of each robot, a local observation map centered on the robot is constructed, and the pose information of each robot is estimated. Further, based on the historical observation information of each robot and the corresponding local observation map, these local maps are transformed into a unified global coordinate system to obtain the global observation map of each robot, and the global observation maps of multiple robots are fused to obtain a global fusion map. This fusion process ensures that all robots share a unified environmental model, so that the robots can make subsequent exploration decisions based on the global observation map and the global fusion map.

[0053] Furthermore, each robot extracts a spatial feature map from the received global observation map. These feature maps are used to represent key information in the environment, such as obstacle positions and passable areas. The robot can estimate its own exploration intention information based on the exploration information shared by other robots. By using the variational autoencoder in the intention reasoning module to encode and decode the exploration information of other robots, a latent space representation is generated, thereby obtaining its own exploration intention information. Then, the exploration intention information is fused with the spatial feature map to obtain the global exploration goal. This goal is the long-term navigation goal of the robot in the global environment. Furthermore, the robot can generate a target exploration path based on the global fusion map and the global exploration goal. This path planning process can consider obstacles and passable areas in the environment, enabling the robot to explore according to the generated target exploration path and output navigation actions, such as moving forward and turning, based on short-term sub-goals.

[0054] Among them, it can be understood that based on the intention reasoning module, robots can share exploration information through the message receiving and sending module. Each robot receives the exploration information sent by other robots and sends its own exploration information to other robots, so that each robot processes the received exploration information through the intention reasoning module, updates its own hidden state information, and generates navigation actions through the action generator. Furthermore, a preset map fuser in the system fuses the local maps of multiple robots into a global map and executes specific exploration actions based on the global map according to the generated navigation action instructions.

[0055] In some embodiments, the multi-robot collaborative exploration method based on intention reasoning is applied to a multi-robot collaborative exploration system. The system may further include a neural SLAM module, a map refinement module, and a multi-robot intention reasoning and perception module. Corresponding to steps S110 to S150 above, this application can construct a 2D map and estimate the robot pose information by executing the neural SLAM module to prevent the accumulation of sensor errors; convert the local map into a global map through the map refinement module and fuse the global maps of all robots to obtain a fused global map; use the multi-robot intention reasoning and perception module to generate a global goal, plan a path through the local planner module, generate a sequence of short-term sub-goals, and output navigation actions according to the short-term sub-goals by the local policy module to guide the robot exploration, so as to be applicable to multi-robot systems including robots and autonomous vehicles, and can effectively improve the efficiency and effect of multi-robot collaborative exploration.

[0056] It can be understood that based on Figure 2Multi-robot collaborative exploration framework structure. In this application, each robot can first use the neural SLAM module to construct its own 2D local map and estimate its own pose. Then, the robots run the intention reasoning module respectively, and infer each other's intentions and strategies based on the observation data of other robots received. Subsequently, the robots run the message generation module, pack their latest observation information and the intention reasoning information obtained from the intention reasoning module into a compact message format, and send it to other robots. After receiving messages from other robots, each robot runs the intention reasoning module again to process the new information and fuse it with its own observation information. Then, the robots use the action generator to calculate the optimal navigation action or path planning scheme according to the fused spatial feature map. Further, each robot executes specific operations according to the instructions provided by the action generator, such as going to a specific area for further exploration or returning to the base to report the discovered situation, thereby enabling the multi-robot system to share information more efficiently, understand and adapt to each other's behavior patterns, and thus improve the exploration efficiency and collaboration ability of the entire team in an unknown environment.

[0057] In some embodiments, corresponding to step S110, this application constructs a 2D map and estimates the current pose information of the robot by executing the neural SLAM module to prevent the accumulation of sensor errors, obtains the sensing data of each robot, and constructs the local observation map of each robot and estimates the pose information of each robot respectively according to the sensing data of each robot. The specific process can be as follows:

[0058] The neural SLAM module (f SLAM ) receives the current RGB observation o i of robot i, the current and last sensor pose data pose′ i of the robot, and the last robot pose and map estimate map i , and outputs the updated map and the current robot pose estimate where θ S represents the training parameter of the neural SLAM module. The neural SLAM module consists of two learning components, a map builder and a pose estimator. The map builder (f Map ) outputs a top-down 2D spatial map P t ego ∈ [0, 1] 2×V×V (where V is the field of view), predicting obstacles and explored areas in the current observation. The pose estimator (f PE ) is based on the past pose estimate and the last two self-centered map predictions to predict the robot's pose

[0059] Among them, it can be understood that the neural SLAM module mainly compares the pose change of the current ego-centric map prediction transformed to the current frame with the previous ego-centric map prediction, transforms the ego-centric map into a geocentric map according to the pose estimation given by the pose estimator, and then aggregates it with the previous spatial map to obtain the current map.

[0060] Reference Figure 3 , Figure 3 is a flowchart for constructing the local observation map of each robot and estimating the pose information of each robot in the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application; in some embodiments, the server obtains the sensing data of each robot, and respectively constructs the local observation map of each robot and estimates the pose information of each robot according to the sensing data of each robot, including at least the following steps:

[0061] Step S310, the server obtains the sensing data of each robot, and obtains multiple current-time maps and previous-time maps centered on the robot according to the sensing data of each robot, and calculates the pose change of each robot according to the current-time map and the previous-time map;

[0062] Step S320, the server performs spatial feature aggregation on the corresponding current-time map and previous-time map based on the pose change of each robot to obtain the local observation map of each robot.

[0063] In some embodiments, the current-time map and the previous-time map respectively refer to the local maps centered on the robot constructed at the current time and the previous time. By comparing the maps at these two times, the pose change of the robot can be calculated, providing a basis for the construction of the local observation map; it can be understood that by comparing the current-time map and the previous-time map, the change in the position and pose of the robot between two consecutive times is obtained, and then the process of integrating the spatial features in the current-time map and the previous-time map of the robot. The local observation map obtained in this way can more accurately reflect the characteristics of the robot's surrounding environment; further, the historical observation map can be understood as a set of observation maps centered on the robot constructed at multiple past times. These maps record the exploration history of the robot, provide the necessary data support for the global coordinate transformation process, and then perform an operation of cropping the boundary of the map transformed to the global coordinate system to remove the unexplorable boundary areas in the map, making the map more focused on the exploitable part and improving the practicality and efficiency of the map.

[0064] In some embodiments, the present application can execute steps S310 to S320 through a map refinement module to convert the self - centered local map of each robot into a self - centered global map. The map refinement module can first obtain the sensing data of each robot and construct the current - moment map and the previous - moment map centered on the robot based on this data. By combining all the past local maps centered on the robot, a global map centered on the robot is obtained. Then, according to the pose estimation, the global maps of all robots are transformed into the same coordinate system. At the same time, to ensure that the feature extractor in the multi - robot information aggregation planner only focuses on the feasible part of the map and generates a more concentrated action space, the unexplorable boundaries of the map are cropped and enlarged to obtain the refined self - centered global map of each robot.

[0065] Reference Figure 4 , Figure 4 is a flowchart for obtaining the global observation map of each robot in the multi - robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application. In some embodiments, the server performs global coordinate transformation processing based on the historical observation information of each robot and the corresponding local observation map respectively to obtain the global observation map of each robot, including at least the following steps:

[0066] Step S410: Obtain multiple historical observation maps centered on the robot based on the historical observation information of each robot respectively, and convert the multiple historical observation maps and the local observation map into maps in the same global coordinate system.

[0067] Step S420: Perform exploration boundary cropping processing on the maps in the same global coordinate system to obtain the global observation map of each robot.

[0068] In some embodiments, the present application can execute steps S410 to S420 through a map fusion module. Input the self - centered global map of each robot into the map fusion module to obtain a fused global map. The map fusion module can integrate these maps by applying a max - pooling operator at each pixel position, perform exploration boundary cropping processing on the maps in the same global coordinate system to obtain the global observation map of each robot. The global observation map can be used in the local planner in subsequent steps to generate better short - term local goals when performing local planning.

[0069] Reference Figure 5 , Figure 5In the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of this application, it is a schematic diagram of the framework structure of the multi-robot intention reasoning sensor. Among them, from robot 1 to robot N, each robot inputs specific information such as the global map, position, trajectory, and target of its own center; CNN feature extractor: uses a convolutional neural network to extract spatial features from the input global map and generate a feature map, which is used to represent key information in the environment, such as obstacle positions and passable areas; in the intention reasoning module, each robot receives the exploration information shared by the other N-1 robots, encodes and decodes the received information through a variational autoencoder (IRM) to generate a latent space representation, infers the intentions of other robots, and aggregates the hidden state information and intention reasoning information of the robot itself through a message generation module to generate a communication vector and send it to other robots; RNN is used for intention embedding, updating the hidden state information of the robot, fusing intention reasoning information and spatial features, and can further process the fused information through a fully connected neural network to generate a global exploration target; in the action generator, the processed features can be projected into the action space through CNN to determine the next exploration area and specific position of the robot, for example, the area and point (xy), and generate a navigation action instruction.

[0070] It can be understood that the CNN feature extractor in the figure shares weights among all robots to improve the consistency of feature extraction. The intention reasoning module is independent for each robot, and the weights are not shared among all robots to adapt to individual differences. Furthermore, through the multi-robot intention reasoning sensor, all robot information can be integrated to achieve efficient information sharing and intention reasoning, and improve the collaborative exploration efficiency.

[0071] In some embodiments, corresponding to Figure 5 , it can be understood that when this application estimates its own detection intention information based on the intention reasoning sensor according to the exploration information shared by other robots, a convolutional neural network can be applied as a feature extractor to extract a spatial feature map from the observation information of each robot; secondly, an intention reasoning module is set for each robot. This module uses a variational autoencoder to infer the strategy and features of the sender, that is, the intention reasoning information; then, in order to achieve effective communication between robots, a recursive method is used to exchange information between robots, and the hidden state information and intention reasoning information of the robot are compressed and encoded into compact information through a message generation module and sent to other robots, thereby improving the exploration efficiency between robots.

[0072] Further, the information output by the message generation module is connected to the spatial feature map of the robot and input into a recurrent neural network to obtain a fused spatial feature map, that is, the updated hidden state information of the robot. The fused spatial feature map is input into the action generator to obtain a global target, thereby realizing efficient multi-robot collaborative exploration.

[0073] Reference Figure 6 , Figure 6 FIG. is a flowchart for estimating detection intention information in the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application; in some embodiments, the detection intention information of itself is estimated according to the exploration information shared by other robots, including at least the following steps:

[0074] Step S610, each robot obtains an intention reasoning module deployed by the server, and encodes the exploration information shared by other robots based on the variational autoencoder in the intention reasoning module to generate a latent space representation;

[0075] Step S620, each robot decodes the latent space representation to obtain its own detection intention information.

[0076] In some embodiments, the intention reasoning module is obtained by the server inputting the exploration information transmitted by other robots into the variational autoencoder, and training the encoder parameters and decoder parameters of the variational autoencoder with the relative entropy of the variational autoencoder as the optimization target.

[0077] In some embodiments, training the encoder parameters and decoder parameters of the variational autoencoder with the relative entropy of the variational autoencoder as the optimization target includes: the server obtains the approximate latent distribution and approximate posterior distribution output by the variational autoencoder based on a multi-layer perception mechanism, and calculates the relative entropy of the variational autoencoder according to the difference between the approximate latent distribution and the approximate posterior distribution, and uses minimizing the relative entropy as the optimization target to train the encoder parameters and decoder parameters of the variational autoencoder.

[0078] Among them, the intention reasoning module can be trained on the server, and the server deploys the trained intention reasoning module on the robot to encode and decode the exploration information shared by other robots to generate its own detection intention information. The variational autoencoder consists of an encoder and a decoder, which is used to learn the latent distribution of data, encode the exploration information of other robots to generate a latent space representation, and then obtain its own detection intention information through decoding.

[0079] It can be understood that the relative entropy is the optimization target used when optimizing the parameters of the variational autoencoder, that is, minimizing the relative entropy of the variational autoencoder, which is calculated by comparing the difference between the approximate latent distribution and the approximate posterior distribution, and guides the training direction of the encoder and decoder parameters.

[0080] Reference Figure 7 , Figure 7 In a multi - robot collaborative exploration method based on intention reasoning provided in an embodiment of the present application, the flowchart of updating hidden state information; in some embodiments, after estimating its own detection intention information according to the exploration information shared by other robots, it further includes at least the following steps:

[0081] Step S710, each robot obtains its corresponding hidden state information, and aggregates the hidden state information and the exploration intention information through a preset fully - connected neural network to generate a communication vector for each robot;

[0082] Step S720, each robot sends its corresponding communication vector as exploration information to other robots, so that other robots respectively update their hidden state information cyclically based on the received exploration information.

[0083] In some embodiments, the hidden state information is the representation of the robot's own state during the exploration process, including information such as the robot's current observation and intention. It is aggregated with the exploration intention information through a preset fully - connected neural network to generate a communication vector; the communication vector is the vector generated after aggregating the robot's own hidden state information and exploration intention information, and is sent to other robots as exploration information for realizing information sharing and collaborative exploration among robots.

[0084] In some embodiments, the discretized grid is to divide the spatial feature map into discrete grid cells, and each grid cell corresponds to a region in the environment. The robot can determine the target exploration region and the relative position coordinates from these grids; the relative position coordinates are the coordinate positions relative to the robot's current position in the target exploration region, which are used to determine the specific position of the global exploration target and guide the robot's navigation direction in the global environment.

[0085] In some embodiments, the present application executes the above method steps S610 to S620, steps S710 to S720 through a multi - robot intention reasoning perceptron; it can be understood that correspondingly Figure 5 , the intention reasoning perceptron can generate a global target for each robot, that is, a long - term navigation target. The specific process may include:

[0086] The intention inference perceptron can obtain the input maps of other robots from the exploration information shared by other robots, and transform the input maps of each robot into the same global coordinate system. The input maps can include the current position, motion trajectory, previous goals, goal history, occupancy map, and obstacle map of the robot. Furthermore, a convolutional neural network is used to generate an H×H×B feature map, where H corresponds to the discretization level of the scene and B is the number of channels. Additionally, extra embedding features are introduced, including the relative position of the robot and the relative position of the previous global goal.

[0087] Furthermore, the variational autoencoder in the intention inference module can learn a density function p(z|x), where z is a latent variable following a normal distribution and x is from the dataset X. However, the true posterior distribution p θ (z|x) is unknown. Therefore, the variational autoencoder uses a multi-layer perceptron to approximate the distribution q φ (z|x); at the same time, the variational autoencoder also uses a multi-layer perceptron to approximate p θ (x|z); calculate the KL divergence between the two distributions:

[0088]

[0089] The above equation can be derived as:

[0090]

[0091] The right side of the above equation is called the evidence lower bound. Since the second term on the left side is greater than or equal to zero, the inequality can be obtained:

[0092]

[0093] Among them, the KL divergence is the relative entropy of the variational autoencoder. It can be understood that in this application, the encoder q and decoder p of the variational autoencoder are parameterized by φ and θ respectively. Assuming that the information transmitted between robots is not lost, the information c transmitted by other robots at time step t t is input into the encoder, and the output is an m-dimensional Gaussian distribution N(z; μ, σ 2 ). Among them, the transmitted information c t is a C×(N - 1) communication tensor, C is the length of the communication vector, and N is the number of robots. The decoder is used to predict p(x t+1 |z), where x t+1 =(o t+1 ,r t+1 ) is the observed information and reward for inferring other robots at time step t + 1. Therefore, the information c transmitted by other robots at time step t t can be passed through q φ(z|c t ) and p θ (z|x t+1 ) are projected onto the latent space Z. The optimization objective is D KL (q φ (z|c t )||p θ (z|x t+1 ), which is equivalent to optimizing the evidence lower bound to implement steps S610 to S620, as derived below:

[0094]

[0095] Furthermore, steps S710 to S720 can be implemented through the message generation module. For the message generation module, a fully connected neural network is adopted. This module aggregates the hidden state information of the robot with the intention information of all other robots obtained by the intention reasoning module to obtain a communication vector. For example, for robot i at time step t, the message generation module is used to generate a communication vector and send it to other robots; furthermore, the intention reasoning module is used to infer the intentions of other robots The hidden state of the robot is tracked using a recurrent neural network as follows:

[0096]

[0097] h t+1 = RNN(o t + c t+1 ).

[0098] In summary, the process is to enable each robot to decode the latent space representation based on the communication vector to obtain its own exploration intention information.

[0099] Refer to Figure 8 , Figure 8 which is the flowchart for obtaining the global exploration goal in the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of this application; in some embodiments, the global exploration goal is obtained by fusing the exploration intention information and the spatial feature map, including at least the following steps:

[0100] Step S810, each robot obtains the target exploration area from the discretized grid corresponding to the spatial feature map, and determines the relative position coordinates corresponding to the exploration intention information from the target exploration area;

[0101] Step S820, each robot obtains the global exploration goal based on the relative position coordinates and the coordinate information of the target exploration area.

[0102] In some embodiments, it can be understood that after obtaining its own detection intention information, the present application executes the above method steps S810 to S820 through an action generator. The action generator mainly fuses the spatial feature map according to the exploration intention information to output a global target. The exploration intention information corresponds to the updated hidden state information [h of the robot. t+1 ; it can be understood that corresponding to step S810, in order to generate an accurate global target, the present application can first select a target exploration area from the H×H discretized grid, and then output a relative position coordinate, that is, the relative position of the global target within the selected area, so that each robot combines the relative position coordinate with the global coordinate information of the target exploration area based on the exploration intention information, calculates the global exploration target, and thus guides the robot to the exploration position.

[0103] Reference Figure 9 , Figure 9 is a flowchart of generating a target exploration path and exploring according to the target exploration path in the multi-robot collaborative exploration method based on intention reasoning provided by an embodiment of the present application; in some embodiments, each robot generates a target exploration path based on the global fusion map for its respective global exploration target, and explores according to the target exploration path, including at least the following steps:

[0104] Step S910, each robot performs local target division for its respective global exploration target based on the global fusion map, and obtains a target sequence composed of multiple local targets;

[0105] Step S920, each robot generates a target exploration path according to the target sequence, and explores according to the target exploration path.

[0106] In some embodiments, the present application can execute the above method steps S910 to S920 through a local planner module. The global target is input into the local planner module. The robot uses a fast marching method to plan a path to achieve the long-term goal in the fused global map, and generates a short-term sub-goal sequence to be input into the local planner module. The local planner module

[0107] In some embodiments, each robot performs local target division for its respective global exploration target based on the global fusion map, and obtains a target sequence composed of multiple local targets, including: each robot uses a fast marching method to calculate the shortest planned path from the current position of the robot to the global target based on the fused global map; each robot extracts multiple local targets at preset distance intervals on the shortest planned path to form a target sequence.

[0108] It can be understood that in the present application, the global exploration goal is decomposed into multiple local goals, and these local goals form a goal sequence in a certain order, which facilitates the robot to gradually achieve the global exploration goal. The goal sequence is an ordered sequence composed of multiple local goals. The robot gradually generates the target exploration path according to this sequence and conducts exploration. Further, the shortest planned path from the current position of the robot to the global goal can be calculated based on the rapid marching method. During the path planning process, the unexplored area is regarded as free space to ensure the feasibility of the path, and local goals are extracted according to a preset distance interval on the shortest planned path. These local goals form a goal sequence to guide the robot to gradually move towards the global goal.

[0109] In some embodiments, the present application can use the rapid marching method to calculate, according to the current spatial map map i the shortest path from the current position of the robot to the long-term goal During the local planning process, the unexplored area is regarded as free space. On the shortest planned path, the farthest coordinate point within 0.25 m around the robot is the short-term goal coordinate. After a short-term sub-goal is given, the local policy module outputs navigation actions according to the visual input, the relative spatial distance, and the relative angle with the sub-goal.

[0110] In some embodiments, exploring according to the target exploration path includes: each robot determines the current local goal from its corresponding target exploration path, and parses the visual input information, the relative spatial distance, and the target relative angle corresponding to the current local goal through a pre-trained visual encoder; each robot outputs a navigation action instruction based on its corresponding visual input information, relative spatial distance, and target relative angle, and executes the exploration action according to the navigation action instruction.

[0111] It can be understood that the visual encoder is a pre-trained neural network model, which can be used to parse the visual input information corresponding to the local goal, convert the visual data into a feature representation that the robot can understand and process, and provide a visual basis for the generation of the navigation action instruction. The relative spatial distance is the spatial distance between the current position of the robot and the local goal, and the target relative angle is the relative angle between the current position of the robot and the local goal. Together with the relative spatial distance, they are used to determine the direction and position of the robot relative to the goal, provide key information for the generation of the navigation action instruction, so that the robot can output a control instruction according to the visual input information, relative spatial distance, and target relative angle, and drive the robot to execute specific exploration actions, such as moving forward, turning, etc., to achieve navigation on the target exploration path.

[0112] In some embodiments, the local policy module can take the current RGB observation (o i ) of the robot and the short-term goal as input and outputs a navigation action where θ L is a parameter of the local policy. The short-term target coordinates are converted into distances and angles relative to the current position of the robot, and then input to the local policy to output a navigation action; it can be understood that the local policy uses a recurrent neural network and includes a pre-trained ResNet18 as a visual encoder.

[0113] Some embodiments of the present application provide a computer device, Figure 10 which is a schematic structural diagram of the computer device provided by an embodiment of the present application. Refer to Figure 10 , the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the multi-robot collaborative exploration method based on intention reasoning in any one of the above embodiments. For example, it executes the method steps S110 to S150 described above Figure 1 in, Figure 3 the method steps S310 to S320 in, Figure 4 the method steps S410 to S420 in, Figure 6 the method steps S610 to S620 in, Figure 7 the method steps S710 to S720 in, Figure 8 the method steps S810 to S820 in, Figure 9 the method steps S910 to S920 in.

[0114] The computer device 1000 of the embodiment of the present application includes one or more processors 1010 and a memory 1020, Figure 10 in which one processor 1010 and one memory 1020 are taken as examples.

[0115] The processor 1010 and the memory 1020 can be connected through a bus or other means, Figure 10 in which connection through a bus is taken as an example.

[0116] The memory 1020, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory 1020 can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1020 optionally includes a memory 1020 remotely set relative to the processor 1010, and these remote memories can be connected to the computer device 1000 through a network. At the same time, examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0117] In some embodiments, when the processor executes the computer program, it executes the multi-robot collaborative exploration method based on intention reasoning according to any one of the above embodiments at a preset interval.

[0118] Those skilled in the art can understand that Figure 10 the device structure shown in

[0119] does not limit the computer device 1000, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 10 In the computer device 1000 shown in

[0120] Based on the hardware structure of the above computer device 1000, various embodiments of the multi-robot collaborative exploration device based on intention reasoning of the present application are proposed. At the same time, the non-transitory software programs and instructions required to implement the multi-robot collaborative exploration method based on intention reasoning of the above embodiments are stored in the memory. When executed by the processor, the multi-robot collaborative exploration method based on intention reasoning of the above embodiments is executed.

[0121] The embodiments of the present application also provide a computer-readable storage medium, which stores computer-executable instructions for executing the above multi-robot collaborative exploration method based on intention reasoning, and can enable the above one or more processors to execute the multi-robot collaborative exploration method based on intention reasoning according to any one of the above embodiments. For example, execute the method steps S110 to S150 described above Figure 1 in Figure 3 the method steps S310 to S320 in Figure 4 the method steps S410 to S420 in Figure 6 the method steps S610 to S620 in Figure 7 the method steps S710 to S720 in Figure 8 the method steps S810 to S820 in Figure 9 the method steps S910 to S920 in

[0122] An embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the multi-robot collaborative exploration method based on intention reasoning in any of the above embodiments. For example, execute the Figure 1 method steps S110 to S150 in Figure 3 method steps S310 to S320 in Figure 4 method steps S410 to S420 in Figure 6 method steps S610 to S620 in Figure 7 method steps S710 to S720 in Figure 8 method steps S810 to S820 in Figure 9 method steps S910 to S920 in; The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed to multiple network nodes. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0123] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer-readable storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that the communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above-mentioned implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A multi-robot collaborative exploration method based on intention reasoning, characterized in that Applied to a multi-robot collaborative exploration system, the multi-robot collaborative exploration system includes a server and multiple robots, and the method includes: The server obtains the sensing data of each robot, and respectively constructs a local observation map of each robot and estimates the pose information of each robot according to the sensing data of each robot; The server respectively performs global coordinate transformation processing based on the historical observation information of each robot and the corresponding local observation map to obtain the global observation map of each robot; The server fuses the multiple global observation maps to obtain a global fusion map, and sends the global observation map and the global fusion map to each robot; Each robot extracts a spatial feature map from the global observation map, estimates its own detection intention information according to the exploration information shared by other robots, and fuses the exploration intention information and the spatial feature map to obtain a global exploration target; Each robot generates a target exploration path for the corresponding global exploration target based on the global fusion map, and explores according to the target exploration path.

2. The multi-robot collaborative exploration method based on intention reasoning according to claim 1, wherein The server obtains the sensing data of each robot, and respectively constructs a local observation map of each robot and estimates the pose information of each robot according to the sensing data of each robot, including: The server obtains the sensing data of each robot, obtains multiple current moment maps and previous moment maps centered on the robot according to the sensing data of each robot, and calculates the pose change of each robot according to the current moment map and the previous moment map; The server respectively performs spatial feature aggregation on the corresponding current moment map and the previous moment map based on the pose change of each robot to obtain the local observation map of each robot.

3. The multi-robot collaborative exploration method based on intention reasoning according to claim 2, wherein The server respectively performs global coordinate transformation processing based on the historical observation information of each robot and the corresponding local observation map to obtain the global observation map of each robot, including: The server respectively obtains multiple historical observation maps centered on the robot based on the historical observation information of each robot, converts the multiple historical observation maps and the local observation map into maps in the same global coordinate system, and performs exploration boundary clipping processing on the maps in the same global coordinate system to obtain the global observation map of each robot.

4. The multi-robot collaborative exploration method based on intention reasoning according to claim 1, wherein The estimating its own detection intention information according to the exploration information shared by other robots includes: Each robot obtains an intention inference module deployed by the server, encodes the exploration information shared by other robots based on the variational autoencoder in the intention inference module to generate a latent space representation; Wherein, the intention inference module is obtained by the server inputting the exploration information transmitted by other robots into the variational autoencoder and training the encoder parameters and decoder parameters of the variational autoencoder with the relative entropy of the variational autoencoder as the optimization target; Each robot decodes the latent space representation to obtain its own detection intention information.

5. The multi-robot collaborative exploration method based on intention reasoning according to claim 4, wherein, Training the encoder parameters and decoder parameters of the variational auto - encoder with the relative entropy of the variational auto - encoder as the optimization objective includes: The server obtains the approximate latent distribution and approximate posterior distribution output by the variational auto - encoder based on the multi - layer perception mechanism, calculates the relative entropy of the variational auto - encoder according to the difference between the approximate latent distribution and the approximate posterior distribution, and trains the encoder parameters and decoder parameters of the variational auto - encoder with minimizing the relative entropy as the optimization objective.

6. The multi-robot collaborative exploration method based on intention reasoning according to claim 4, characterized in that After estimating its own detection intention information based on the exploration information shared by other robots, it further includes: Each robot obtains its corresponding hidden state information, and aggregates the hidden state information and the exploration intention information through a preset fully - connected neural network to generate a communication vector for each robot. Each robot sends the corresponding communication vector as exploration information to other robots, so that the other robots respectively update their hidden state information cyclically based on the received exploration information.

7. The multi-robot collaborative exploration method based on intention reasoning according to claim 1, characterized in that Fusing the exploration intention information and the spatial feature map to obtain the global exploration target includes: Each robot obtains the target exploration area from the discretized grid corresponding to the spatial feature map, and determines the relative position coordinates corresponding to the exploration intention information from the target exploration area. Each robot obtains the global exploration target according to the relative position coordinates and the coordinate information of the target exploration area.

8. The multi-robot collaborative exploration method based on intention reasoning according to claim 1, characterized in that Each robot generates a target exploration path for its corresponding global exploration target based on the global fusion map and explores according to the target exploration path, including: Each robot performs local target division for its corresponding global exploration target based on the global fusion map to obtain a target sequence composed of multiple local targets. Each robot generates a target exploration path according to the target sequence and explores according to the target exploration path.

9. The multi-robot collaborative exploration method based on intention reasoning according to claim 8, characterized in that Each robot performs local target division for its corresponding global exploration target based on the global fusion map to obtain a target sequence composed of multiple local targets, including: Each robot uses the fast marching method based on the global fusion map to calculate the shortest planned path from the current position of the robot to the global target. Each robot extracts multiple local targets at preset distance intervals on the shortest planned path to form a target sequence.

10. The multi-robot collaborative exploration method based on intention reasoning according to claim 9, wherein Exploring according to the target exploration path includes: Each robot determines the current local target from its corresponding target exploration path, and parses the visual input information, relative spatial distance, and target relative angle corresponding to the current local target through a pre - trained visual encoder. Each robot outputs a navigation action instruction based on its corresponding visual input information, relative spatial distance, and target relative angle, and executes exploration actions according to the navigation action instruction.

11. An electronic device, characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When at least one of the programs is executed by at least one of the processors, the multi-robot collaborative exploration method based on intention reasoning according to any one of claims 1 to 10 is implemented.

12. A computer-readable storage medium storing computer-executable instructions for executing the multi-robot collaborative exploration method based on intention reasoning according to any one of claims 1 to 10.

Citation Information

Cited By

  • Man-machine cooperative work system and method and storage medium

    CN121783175A

  • Map-free navigation method based on global grid memory and access popularity

    CN122237630A

  • A map-free navigation method based on global grid memory and access heat

    CN122237630B