A human-machine multi-node collaborative semantic laser SLAM system and method
Through the human-machine multi-node collaborative semantic laser SLAM system, the high-precision global map is built using drones, unmanned vehicles and human collaboration, which solves the problem of insufficient map construction accuracy and efficiency in large-scale environments of traditional SLAM, and realizes efficient semantic information fusion and dynamic noise removal.
Patent Information
- Application Number
- CN202310255617.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-03-16
AI Technical Summary
When building maps and positioning in large-scale unknown environments, traditional SLAM systems cannot effectively distinguish similar but different goals, and a single robot cannot meet the mapping requirements of large-scale environments. The existing multi-machine collaborative mapping method has shortcomings in environmental generalization and optimization.
The human-machine multi-node collaborative semantic laser SLAM system is adopted to use the cooperation of drones, unmanned vehicles and human collaboration ends to integrate maps using semantic information to build a global map, including drone nodes for vertical three-dimensional space construction, unmanned vehicle nodes for ground construction, human terminals for interaction and correction, and server terminals for data calculation and integration.
It improves the accuracy and speed of large-scale environmental map construction, can select optimization strategies based on different scenarios, eliminate dynamic noise, and improves map construction accuracy and efficiency.
Smart Images

Figure CN116358520B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of synchronous positioning and mapping, and in particular relates to a human-machine multi-node collaborative semantic laser SLAM system and method. Background Art
[0002] Simultaneous Localization and Mapping (SLAM) refers to the exploration of an unknown environment using specific sensors without prior knowledge of the environment, building a model of the surrounding unknown environment during movement, obtaining an environmental map, and estimating its own movement and position. SLAM is widely used in a variety of scenarios, including indoor sweeping robots in conventional scenarios, high-precision mapping and positioning for outdoor autonomous driving, as well as unconventional cave environment exploration, disaster detection and warning, and underwater riverbed mapping, with a wide range of application scenarios. However, traditional SLAM generally assumes a static small environment, and the resulting map has a low semantic level, making it impossible to distinguish between similar but different targets. In addition, traditional SLAM is generally carried on a single subject and uses a single sensor to explore the environment. During this process, people generally perform remote operations and observe the mapping effects at a distance from the location environment. However, in actual environments, it is often necessary to map and locate large-scale unknown environments. A single machine can no longer meet the requirements, and large-scale multi-machine collaborative mapping has emerged. However, in existing large-scale collaborative mapping, the same type of unmanned vehicles or multiple drones of the same category are generally used for collaboration. The node settings are relatively fixed, with poor optimizability and poor generalization to different environments. Summary of the Invention
[0003] The present invention discloses a human-machine multi-node collaborative semantic laser SLAM system and method, which can effectively solve the technical problems involved in the background technology.
[0004] To achieve the above object, the technical solution of the present invention is:
[0005] A human-machine multi-node collaborative semantic laser SLAM system, comprising:
[0006] The node side is used to map the regional vertical three-dimensional space and ground environment;
[0007] The collaborative end is used to measure the true value of map geometry and evaluate the mapping effect of the new framework;
[0008] The server is used to calculate the data sent from the node and the collaboration end for map fusion.
[0009] As a preferred improvement of the present invention, the node end includes at least one drone node and one unmanned vehicle node; the drone node is used to map the longitudinal three-dimensional space of the area, and the unmanned vehicle node is used to map the environment at a lower altitude on the ground.
[0010] As a preferred improvement of the present invention, the cooperation end is composed of humans equipped with interactive devices, and the number of the cooperation end is at least one.
[0011] A human-machine multi-node collaborative semantic laser SLAM method based on any one of the systems described above, comprising the following steps:
[0012] Step 1: Set up the drone node, unmanned vehicle node, collaboration terminal, and server terminal;
[0013] Step 2: Collect point cloud data from UAV nodes and unmanned vehicle nodes, extract semantic information, and construct local semantic maps respectively;
[0014] Step 3: The server performs semantic feature matching and generates a map fusion matrix based on the pose and semantic information;
[0015] Step 4: Fuse the local semantic maps through the map fusion matrix to build a global map;
[0016] Step 5: The server transmits the global map to the drone node, the unmanned vehicle node, and the collaboration end;
[0017] Step 6: The server refreshes the mileage of the autonomous vehicle node and corrects the global map pose using its newly extracted features;
[0018] Step 7: The collaboration end collects the true geometric value of the map area and uploads it to the server to assist in optimizing the global map.
[0019] As a preferred improvement of the present invention, the number of the drone node, the unmanned vehicle node and the collaborative end is at least one.
[0020] As a preferred improvement of the present invention, the collaboration end collects the area map through a wearable camera or laser equipment.
[0021] As a preferred improvement of the present invention, a collaborative map is constructed by using drone nodes, unmanned vehicle nodes, and collaborative terminals. The ideal model for constructing the collaborative map is:
[0022]
[0023] Among them, i represents the corresponding node, d i represents the corresponding average point cloud density, f(i) represents the degree of overlap between the edge of the map and other nodes, w(i) represents the prior weight, N represents the total number of nodes, MN Indicates the ideal mapping state.
[0024] As a preferred improvement of the present invention, each node forms different local semantic maps through dynamic matching and static elimination of the semantic information of each frame and uploads them to the server. A map fusion matrix is established through similarity matching of local maps. Maps with the same semantic instance change the fusion matrix according to their relative postures. Then, map fusion is completed by estimating the consistency of known continuous posture changes and discrete semantic information of different nodes.
[0025] As a preferred improvement of the present invention, in step six, a filter is used to obtain the correction of the posture, and on this basis, the map is further fused by measuring the true geometric value of the area through the collaborative end to complete the optimization of the map.
[0026] The beneficial effects of the present invention are as follows:
[0027] 1. The map fusion strategy proposed in this paper is mainly based on the semantic features extracted from local maps. The map fusion matrix is established using the similarity evaluation of semantic information. Before the map fusion, the matrix is multiplied according to the relative pose of each different position and the collaborative noise to generate a collaborative global map based on semantic information. This strategy fully utilizes the advantages of collaboration and semantic SLAM, integrates more local information, and significantly improves the accuracy of scene map construction.
[0028] 2. Through the cooperation of various nodes and human units, the problem that traditional SLAM cannot complete large-scale map exploration is solved, the map construction speed is improved, and different optimization criteria can be selected according to different application scenarios to achieve the best mapping strategy;
[0029] 3. By using collaborative semantic SLAM, semantic information is more sensitive to feature matching in dynamic processes. In the process of building a complete map, dynamic noise can be effectively eliminated and mapping accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a system framework diagram of the present invention;
[0031] Figure 2 It is a workflow diagram of the present invention;
[0032] Figure 3 This is a diagram showing the map fusion principle of the present invention;
[0033] Figure 4 This is the map optimization effect diagram of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0036] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referenced. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this disclosure, "plurality" means at least two, such as two or three, unless otherwise specifically defined.
[0037] In the present invention, unless otherwise specified or limited, the terms "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can mean fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0038] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0039] See also Figure 1 As shown in the figure, the present invention proposes a human-machine multi-node collaborative semantic laser SLAM system, including a node end 1, a collaboration end 3, and a server end 2. The node end 1 includes at least one drone node 11 and an unmanned vehicle node 12. The drone node 11 is used to map the vertical three-dimensional space of the area; the unmanned vehicle node 12 is used to explore the lower altitude environment of the location and cover a larger ground area for mapping.
[0040] The collaboration end 3 is composed of humans equipped with interactive equipment, and is responsible for follow-up observation and secondary correction of key areas or areas with low mapping confidence, so as to measure the geometric true value of the map and evaluate the mapping effect of the new framework.
[0041] The server 2 is used to calculate the data transmitted from the node end and the collaboration end to perform map fusion.
[0042] See also Figures 2 to 4 As shown, the present invention also proposes a human-machine collaborative semantic laser SLAM method, comprising the following steps:
[0043] Step 1: Set up the node, collaboration, and server.
[0044] The node terminals include at least one drone node and one unmanned vehicle node. The collaborative semantic framework proposed in this invention can operate with a minimum of one server terminal, one collaboration terminal, and two node terminals. The number of collaboration terminals and node terminals can be increased based on this. If a node terminal is damaged and unable to operate, it can be removed from the system through self-inspection to avoid being affected. During node status initialization and update, each node performs a system self-inspection to obtain the number of collaboration terminals, drone nodes, and unmanned vehicle nodes.
[0045] Through collaborative mapping between the node side, the collaboration side, and the server side, we must first determine the ideal model for collaborative mapping. The ideal model refers to the ideal result of single-node mapping, the ideal overlapping boundary of collaborative mapping, and the mapping effect of the focus area meets the requirements. Considering that the average point cloud density distribution of a single node in the actual space should not be lower than the set threshold, the higher the average density, the better; the overlap degree of different nodes at the edge of the mapping should not be higher than the set threshold, nor should there be no intersection; the mapping effect of the focus area and the difficult area should meet the set threshold conditions, and the larger the weight, the better. Therefore, the corresponding ideal model of collaborative mapping is shown in formula (1):
[0046]
[0047] Among them, i represents the corresponding node, d i represents the corresponding average point cloud density, f(i) represents the degree of overlap between the edge of the map and other nodes, w(i) represents the prior weight, N represents the total number of nodes, M N Indicates the ideal mapping state.
[0048] Step 2: Collect point cloud data of UAV nodes and unmanned vehicle nodes through lidar, use lightweight semantic segmentation network to extract semantic information, and construct local semantic maps respectively.
[0049] Step 3: The server performs semantic feature matching and generates a map fusion matrix based on the position and semantic information of each node obtained using the laser SLAM algorithm.
[0050] Each node first independently describes the semantic model of its observations, dynamically matching and statically removing semantic information from each frame to form local semantic maps from different viewpoints. Each node then uploads the local semantic map to the server, which leverages its computing resources to build a map fusion matrix by similarity matching of local maps. Maps with the same semantic instance undergo changes in the fusion matrix based on their relative poses. By estimating the consistency of known continuous pose changes and discrete semantic information from different nodes, the map fusion strategy is implemented using the semantic information from different nodes.
[0051] Step 4: Fuse the local semantic maps through the map fusion matrix to establish a global map.
[0052] Step 5: The server transmits the global map to the drone node, the unmanned vehicle node, and the collaboration end.
[0053] Step 6: The server refreshes the mileage of the unmanned vehicle node and corrects the global map pose using its newly extracted features.
[0054] Based on the LOAM algorithm, a correction method is proposed that uses a collaborative lidar odometry to update and optimize the pose, and then feeds it back to the mapping end. A high-frequency, low-precision pose estimation node and a low-frequency, high-precision mapping correction node are set up separately to obtain high-precision lidar odometry information and an environmental point cloud map. The unmanned vehicle node unit recollects point cloud data, inputs it into the lidar odometry for frame-to-frame pose transformation estimation, and then inputs it into the server-side fusion node. Matching is performed on the generated point cloud map, and a filter is used to obtain a pose correction. Based on this correction, the map fusion is iterated to complete the optimization.
[0055] Step 7: The collaboration end collects the true geometric value of the map area and uploads it to the server to assist in optimizing the global map.
[0056] After the unmanned vehicle node corrects the global map pose, the human-worn camera or laser equipment is used to follow and observe key areas or areas with low map confidence, make secondary corrections, calculate errors and evaluate effects to obtain the optimal map.
[0057] The beneficial effects of the present invention are as follows:
[0058] 1. The map fusion strategy proposed in this paper is mainly based on the semantic features extracted from local maps. The map fusion matrix is established using the similarity evaluation of semantic information. Before the map fusion, the matrix is multiplied according to the relative pose of each different position and the collaborative noise to generate a collaborative global map based on semantic information. This strategy fully utilizes the advantages of collaboration and semantic SLAM, integrates more local information, and significantly improves the accuracy of scene map construction.
[0059] 2. Through the cooperation of various nodes and human units, the problem that traditional SLAM cannot complete large-scale map exploration is solved, the map construction speed is improved, and different optimization criteria can be selected according to different application scenarios to achieve the best mapping strategy;
[0060] 3. By using collaborative semantic SLAM, semantic information is more sensitive to feature matching in dynamic processes. In the process of building a complete map, dynamic noise can be effectively eliminated and mapping accuracy can be improved.
[0061] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A human-machine multi-node collaborative semantic laser SLAM method, characterized in that: The following steps are involved: Step 1: Set up the drone node, unmanned vehicle node, collaboration terminal, and server terminal; Step 2: Collect point cloud data from UAV nodes and unmanned vehicle nodes, extract semantic information, and construct local semantic maps respectively; Step 3: The server performs semantic feature matching and generates a map fusion matrix based on the pose and semantic information. Specifically, each node dynamically matches and statically eliminates the semantic information of each frame to form different local semantic maps and uploads them to the server. The map fusion matrix is established by similarity matching of local maps. Maps with the same semantic instance undergo changes in the fusion matrix based on their relative pose. Map fusion is then completed by estimating the consistency of the known continuous pose changes and discrete semantic information of different nodes. Step 4: Fuse the local semantic maps through the map fusion matrix to build a global map; Step 5: The server transmits the global map to the drone node, the unmanned vehicle node, and the collaboration end; Step 6: The server refreshes the mileage of the autonomous vehicle node and corrects the global map pose using its newly extracted features; Step 7: The collaboration end collects the true geometric value of the map area and uploads it to the server to assist in optimizing the global map.
2. The method according to claim 1, wherein: The number of the drone node, unmanned vehicle node and collaboration terminal is at least one.
3. The method according to claim 1, wherein: The collaboration end collects the area map through a wearable camera or laser equipment.
4. The method according to claim 1, wherein: Point cloud data is acquired through lidar, and the lightweight semantic segmentation network U-Net is used to extract semantic information.
5. The method according to claim 1, wherein: The ideal model for collaborative mapping is as follows: Among them, i represents the corresponding node, d i represents the corresponding average point cloud density, f(i) represents the degree of overlap between the edge of the map and other nodes, w(i) represents the prior weight, N represents the total number of nodes, M N Indicates the ideal mapping state.
6. The method according to claim 1, wherein: In step six, a filter is used to obtain the correction of the posture. On this basis, the map is further fused by measuring the true geometric value of the area through the collaborative end to complete the optimization of the map.
7. A human-machine multi-node collaborative semantic laser SLAM system for executing the method according to any one of claims 1 to 6, characterized in that: include: The node side is used to map the regional vertical three-dimensional space and ground environment; The collaborative end is used to measure the true value of map geometry and evaluate the mapping effect of the new framework; The server is used to calculate the data sent from the node and the collaboration end for map fusion.
8. The system according to claim 7, characterized in that: The node end includes at least one drone node and one unmanned vehicle node; the drone node is used to map the longitudinal three-dimensional space of the area, and the unmanned vehicle node is used to map the environment at a lower altitude on the ground.
9. The system according to claim 7, characterized in that: The cooperation end is composed of humans equipped with interactive devices, and the number of the cooperation end is at least one.
Citation Information
Patent Citations
Map construction method, apparatus and storage medium
US20230083965A1