Multi-robot loopback detection method and system based on bimodal verification and storage medium
Through the multi-task deep semantic enhancement network and conflict graph screening method based on deep learning, the robustness and accuracy problems of loop detection in multi-robot collaborative SLAM are solved, redundant loops are eliminated, and detection efficiency is improved.
Patent Information
- Application Number
- CN202510753652.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-10
AI Technical Summary
The loop detection in existing multi-robot collaborative SLAM lacks robustness and accuracy in dynamic scenes, and there are redundant candidate loops, resulting in high back-end graph optimization overhead.
A multi-task deep semantic enhancement network based on deep learning is used to extract the semantic and geometric features of the image, distinguish dynamic and static areas, give dynamic areas a lower weight, and screen candidate loops through geometric consistency, semantic similarity and spatiotemporal continuity. A conflict graph is constructed to screen the maximum independent set and eliminate redundant loops.
This improves the accuracy and robustness of loop detection, reduces redundant loops, and reduces backend graph optimization overhead.
Smart Images

Figure CN120765732A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to simultaneous localization and mapping of robots, in particular to a multi-robot loop detection method, system and storage medium based on double-mode verification. BACKGROUND
[0002] Multi-robot cooperative simultaneous localization and mapping (SLAM) can improve work efficiency and is suitable for large-scale complex scenes such as urban canyons, underground spaces and large indoor environments. The precision and robustness of loop detection in multi-robot cooperative SLAM directly affect the map fusion quality and trajectory optimization effect. However, the existing technology has the following key defects:
[0003] The existing method relies on extracted geometric features for feature fusion matching, but geometric features are sensitive to light changes and view angle differences and are prone to failure in dynamic scenes, resulting in low cross-view matching success rate in dynamic scenes, and thus low robustness of subsequent loop detection. The existing loop detection relies too much on image similarity and ignores the spatiotemporal constraints, resulting in insufficient loop detection accuracy. Secondly, the number of loops in multi-robot cooperative SLAM is greatly increased compared to single-robot SLAM candidate loops, and there are a large number of redundant candidate loops. If not optimized, it will greatly increase the backend graph optimization overhead. SUMMARY
[0004] The purpose of the present application is to provide a multi-robot loop detection method, system and storage medium based on double-mode verification that can accurately identify loops and eliminate redundant loops in dynamic environments.
[0005] Technical solution: The multi-robot loop detection method based on double-mode verification according to the present application is characterized by comprising the following steps:
[0006] S1. Constructing a multi-task deep semantic enhancement network based on deep learning for extracting semantic features and geometric features of images;
[0007] S2. According to the extracted semantic features, distinguishing dynamic and static regions in the image and assigning a lower weight to the dynamic region, and calculating the similarity of the multi-robot image according to the weighted region for preliminary fusion;
[0008] S3. Selecting cross-robot candidate loops that meet geometric consistency, semantic similarity and spatiotemporal continuity from the multi-robot preliminary fused images;
[0009] S4. Constructing a conflict graph according to the cross-robot candidate loops selected in step S3 to select a maximum independent set and eliminate redundant loops.
[0010] Based on the above technical solution, a multi-task deep semantic enhancement network is constructed to extract the semantic features of the image to distinguish the dynamic and static areas in the image, and assign lower weights to the dynamic areas. This can avoid the interference of dynamic areas during cross-view matching and improve the matching accuracy. On this basis, candidate loops are screened by geometric consistency, semantic similarity and spatiotemporal continuity. Not only image similarity, that is, semantic similarity, but also geometric consistency and spatiotemporal continuity are considered, which can improve the accuracy of loop detection. Finally, among the screened candidate loops, a conflict graph is constructed and the maximum independent set is screened as the final detected loop, which can eliminate a large number of redundant loops and reduce the back-end graph optimization overhead.
[0011] Preferably, step S4 includes the following sub-steps:
[0012] S4.1. Construct a conflict graph of candidate loop nodes based on the nodes in the inter-robot candidate loops selected in step S3 and select the maximum independent set.
[0013] S4.2. Based on the loop candidate nodes filtered out in step S4.1, candidate loop edges are constructed. Based on this, a conflict graph of the candidate loop edges is constructed and a maximum independent set is filtered out. The candidate loop edges in the maximum independent set are the loop edges finally detected between robots.
[0014] By constructing conflict graphs twice and screening them in batches, the first conflict graph of loop candidate nodes is constructed to screen out the most conflict-free loop candidate nodes, that is, to eliminate redundant loop candidate nodes. On this basis, the conflict graph of candidate loop edges is constructed for the second time to screen out the most conflict-free loop edges as the final detected loop edges, and redundant loop edges are further proposed. After these two steps, redundant loops can be greatly reduced and the efficiency of back-end graph optimization can be improved.
[0015] The multi-robot loop detection system based on dual-modal verification of the present invention comprises:
[0016] Network building module: used to build a multi-task deep semantic enhancement network based on deep learning to extract semantic and geometric features of images;
[0017] Preliminary fusion module: This module is used to distinguish dynamic and static areas in the image based on the extracted semantic features and assign a lower weight to the dynamic area. The multi-robot images are preliminarily fused based on the similarity calculated based on the weighted areas.
[0018] Preliminary screening module: used to screen out cross-robot candidate loops that meet geometric consistency, semantic similarity, and spatiotemporal continuity based on the preliminarily fused images of multiple robots;
[0019] Final screening module: used for screening out the maximum independent set from the conflict graph screened out by the cross-robot candidate loop according to the preliminary screening module to eliminate redundant loops.
[0020] The computer readable storage medium storing one or more programs of the application comprises one or more programs including instructions, which when executed by a computing device, cause the computing device to perform any of the above methods.
[0021] Beneficial effects: the multi-task deep semantic enhancement network jointly learns semantic and geometric features, and based on the semantic features extracted by the multi-task deep semantic enhancement network, a lower weight is assigned to the dynamic area to avoid interference of the dynamic area, thereby improving the robustness and success rate of cross-view matching; the candidate loop is screened through geometric consistency, semantic similarity and spatiotemporal continuity, thereby breaking the false loop bottleneck of the traditional image similarity-based method and significantly improving the accuracy of loop detection; the conflict graph-based screening of the maximum independent set can eliminate a large number of redundant loops, thereby reducing the optimization overhead of the backend graph while ensuring the detection accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a flowchart of the method;
[0023] Figure 2 is a schematic diagram of the multi-task deep semantic enhancement network structure;
[0024] Figure 3 is a flowchart of the preliminary fusion module of the application;
[0025] Figure 4 is a flowchart of the preliminary screening module of the application;
[0026] Figure 5 is a flowchart of the final screening module of the application. DETAILED DESCRIPTION
[0027] As shown in the figure, the multi-robot loop detection method based on dual-modal verification of the application comprises the following steps:
[0028] S1, constructing a multi-task deep semantic enhancement network based on deep learning for extracting semantic features and geometric features of images;
[0029] The multi-task deep semantic enhancement network adopts a shared encoder (shared encoding network) and a double-branch structure, and the double-branch structure is respectively a semantic decoding network and a geometric feature decoding network; the shared encoding network is composed of three convolutional layers connected in sequence, and each layer structure is composed of a convolution kernel with a size of 3*3, a step of 2, a padding of 1, a batch normalization (BatchNorm) and a ReLU activation function. The encoding process realizes down-sampling through a convolution step, and the backbone encoding output (256*60*80) is shared by the semantic decoding network and the geometric feature decoding network.
[0030] The semantic decoding network is composed of three transposed convolutional layers connected in sequence, and the structure is composed of a convolution kernel with a size of 3*3, a step of 2, a padding of 1, a batch normalization (BatchNorm) and a ReLU activation function. The final output realizes Softmax classification through a cross-entropy loss function.
[0031] The geometric feature decoding network is composed of an adaptive average pooling layer and a fully connected layer. The adaptive average pooling layer compresses the feature map of 256*60*80 into 256*1*1, and the fully connected layer can convert the input of 256 dimensions into the output of 128 dimensions, and then the ReLU activation outputs the geometric descriptor.
[0032] After inputting the RGB image into the multi-task deep semantic enhancement network, deep visual features are extracted by the shared encoder, and are respectively input into the semantic decoding network and the geometric feature decoding network to generate semantic prediction and 128-dimensional global descriptors for loop matching and graph optimization.
[0033] To realize the collaborative optimization of the semantic and geometric two tasks, a joint loss function is introduced to train the multi-task deep semantic enhancement network. Specifically:
[0034] In the semantic decoding network, a standard cross-entropy loss is used for pixel-level supervision, The calculation formula is
[0035]
[0036] Where N is the total number of pixels, C is the number of categories, and are the true label and the predicted value of the i-th pixel as the category c;
[0037] In the geometric feature decoding network, a triplet loss (Triplet Margin Loss) is used for feature optimization, and the discriminability of the descriptor is optimized by constructing an (anchor, positive, negative) triplet, The calculation formula is
[0038]
[0039] where T is the number of triplets, denotes the feature of the i-th anchor image, denotes the feature of the i-th positive sample, denotes the feature of the i-th negative sample, and a is the margin constant.
[0040] The joint loss function of the final multi-task deep semantic enhancement network is as follows, which is used to simultaneously optimize the semantic prediction accuracy and the geometric description discriminativeness:
[0041]
[0042] where, is the loss function of the multi-task deep semantic enhancement network; λ is a hyperparameter used to control the contribution of the two tasks (i.e., the semantic decoding network and the geometric feature decoding network), which is set to 0.6 in the experiment, aiming to strengthen the semantic supervision dominance.
[0043] S2, according to the extracted semantic features, the dynamic and static regions in the image are distinguished and the dynamic regions are given a lower weight, and the multi-robot image is preliminarily fused according to the weighted regions.
[0044] The traditional feature matching method is easily disturbed by moving objects in a dynamic scene, resulting in a large number of false matches. The semantic-guided local matching mechanism distinguishes static and dynamic regions by introducing semantic information and assigns different attention weights to different regions, thereby suppressing false matches in dynamic regions.
[0045] The mechanism is based on the following core idea: using the prediction results of the semantic decoding network to identify dynamic objects, and attenuating the feature matching weight of the dynamic region, that is, generating a semantic attention weight for the dynamic region, and assigning a lower matching weight to the dynamic region, and keeping a high weight for the static region, forming an attention mask.
[0046] The formula for calculating the semantic attention weight is:
[0047]
[0048] In the formula, w(i) represents the semantic attention weight of image pixel point i, w(i) = a0 when pixel point i belongs to a dynamic region, otherwise w(i) = 1, a0 ∈ [0, 1] is a dynamic region attenuation coefficient.
[0049] When calculating the feature similarity, the semantic attention weight is combined to reduce the influence of the dynamic region on the matching result, and when calculating the cosine similarity of different images, the semantic attention weight is combined, and the specific calculation formula is:
[0050]
[0051] wherein Sim(f i ,f j ) is the cosine similarity considering semantic attention weight; f i and f j are feature descriptors of images i and j respectively; f i (p) represents the feature of image i at pixel point p, f j (q) represents the feature of image j at pixel point q; w(p) and w(q) are semantic attention weights of pixel points p and q respectively.
[0052] The preliminary fusion of multi-robot images based on the cosine similarity considering semantic attention weight can avoid the interference of dynamic areas, improve the robustness and success rate of cross-view matching, and the subsequent loop detection based thereon can also improve the robustness of loop detection.
[0053] S3, screening out cross-robot candidate loops that meet the geometric consistency, semantic similarity and spatio-temporal continuity from the preliminary fused images of multi-robots;
[0054] In loop detection, the mis-matched loop edges will lead to divergence of graph optimization, and the existing technology excessively relies on image similarity, so it is necessary to screen effective candidate loops through multi-dimensional constraints to reduce the mis-selected loops, which is specifically screened through the following three dimensions:
[0055] (1) The Mahalanobis distance is used to measure the consistency of feature point space distribution to exclude geometric outliers caused by noise, and the Mahalanobis distance of different robot loop frames is calculated by the following formula
[0056]
[0057] wherein D M (x,y) is the Mahalanobis distance, x and y are feature point space coordinate vectors of different robot loop frames, and Σ -1 is the information matrix;
[0058] When D M (x,y)≤τ g , τ g is the geometric threshold value corresponding to the Mahalanobis distance, that is, when D M (x,y) is not greater than the corresponding threshold value, it is considered to meet the consistency of space distribution, that is, to meet the geometric consistency.
[0059] (2) The cosine similarity of deep features is used to verify the scene consistency of cross-robot view, that is, the semantic similarity, and the calculation formula of the cosine similarity is
[0060]
[0061] CosSim(f a ,f b ) is the cosine similarity, f a and f b are the global geometric feature vectors of loop closure frames a and b respectively, and loop closure frames a and b are key frames generated by different robot front-end SLAM systems in multi-robot collaborative mapping.
[0062] When CosSim(f a ,f b ) ≥ 0.6τ s , τ s is the semantic similarity threshold, that is, CosSim(f a ,f b ) is not less than the corresponding threshold (0.6τ s , the coefficient before τ s can be adjusted according to actual conditions) is considered to be semantic similarity.
[0063] (3) When the following determination formula is satisfied, it is determined that the loop closure frames of different robots or each k frame (k is a constant, which can be set to different values according to actual conditions) before and after them satisfy the spatio-temporal continuity:
[0064]
[0065] Where τ s is the semantic similarity threshold, and v represents the effective pairing set; wherein CosSim(f a ,f b ) ≥ 0.8τ s indicates a single strong match, that is, there is at least one pair of high semantic similarity, excluding isolated strong matches; indicates the overall matching quality, that is, the average similarity meets the standard and there are at least two effective pairs, in order to enhance the spatio-temporal robustness;
[0066] and are the set of loop closure frame a and each k frame before and after it and the set of loop closure frame b and each k frame before and after it, respectively, and the specific expression formula is
[0067]
[0068] Where v i is the loop closure frame i, v i-k and v i+k are the k frames before and after the loop closure frame i respectively, and k is a constant.
[0069] The loop closure that needs to satisfy the constraints of the above three dimensions at the same time can be used as a candidate loop closure.
[0070] The pairing process according to the above judgment formula mainly involves first extracting the global features f of each frame in the window, then calculating the similarity of different frame pairs, screening out candidate frame pairs that meet the conditions, and selecting frame pairs that meet the single-pair strong match and overall matching quality to form a valid pairing set. The number of elements in this set must be ≥2. Only the loop frames of two different robots that meet these conditions can be considered to meet spatiotemporal continuity.
[0071] S4, constructing a conflict graph based on the inter-robot candidate loops screened out in step S3, screening out the maximum independent set and eliminating redundant loops;
[0072] S4.1. Construct a conflict graph of candidate loop nodes based on the nodes in the inter-robot candidate loops selected in step S3 and select the maximum independent set.
[0073] The expression formula of the conflict graph of loop candidate nodes is
[0074] G=(V,E)
[0075] Where G is the conflict graph of the loop candidate nodes; V is the vertex set of the conflict graph, which is expressed as V = {l1,l2,...,l n}, l n is the nth loop candidate node, n is a constant; E is the edge set of the conflict graph, and its expression is E={(l i ,l j )|Overlap(l i ,l j )>τ conflict}, where Overlap(l i ,l j ) is loop l i and l j The proportion of shared feature points; τ conflict is the conflict threshold, the specific value of which can be set according to actual conditions. In this embodiment, 0.3 is used.
[0076] S4.2. Based on the loop candidate nodes filtered out in step S4.1, candidate loop edges are constructed. Based on this, a conflict graph of the candidate loop edges is constructed and a maximum independent set is filtered out. The candidate loop edges in the maximum independent set are the loop edges finally detected between robots.
[0077] The conflict graph of candidate loop edges is expressed as
[0078] G c =(V c ,E c )
[0079] Among them, G c is the conflict graph of candidate loop edges; V cis the vertex set of the conflict graph, and its expression is V c ={e1,e2,...,e n}, e n is the nth candidate loop edge, n is a constant; E c is the edge set of the conflict graph, and its expression is in, and Represent the candidate loop edge e i and e j There are time conflicts, spatial conflicts and geometric contradictions between them. Specifically, when the two loop time windows overlap, for example, the loop time window overlap rate is greater than 30% (the specific overlap rate can be adjusted according to actual conditions), the two candidate loop edges are considered to have a time conflict. When the overlap rate of the two loop coverage areas is greater than the threshold (the threshold is set according to actual conditions), they are considered to have a spatial conflict. When the posture transformation error of the two loops is greater than the threshold (the threshold is set according to actual conditions), they are considered to have a geometric conflict.
[0080] The goal of maximum independent set (MIS) screening is to find the conflict-free subset with the largest number of vertices in the conflict graph. The core of this method is to maximize the retention of valid loops while ensuring that the loop candidates are conflict-free (such as spatial overlap, temporal redundancy, geometric contradiction, etc.). Both S4.1 and S4.2 use greedy algorithms to screen the maximum independent set. Of course, other algorithms can also be used to screen the maximum independent set.
[0081] The multi-robot loop detection system based on dual-modal verification of the present invention includes:
[0082] Network building module: used to build a multi-task deep semantic enhancement network based on deep learning to extract semantic and geometric features of images;
[0083] Preliminary fusion module: This module is used to distinguish dynamic and static areas in the image based on the extracted semantic features and assign a lower weight to the dynamic area. The multi-robot images are preliminarily fused based on the similarity calculated based on the weighted areas.
[0084] Preliminary screening module: used to screen out cross-robot candidate loops that meet geometric consistency, semantic similarity, and spatiotemporal continuity based on the preliminarily fused images of multiple robots;
[0085] Final screening module: used to construct a conflict graph based on the cross-robot candidate loops screened by the preliminary screening module to screen out the maximum independent set and eliminate redundant loops.
[0086] The computer-readable storage medium storing one or more programs according to the present invention includes one or more programs including instructions, which, when executed by a computing device, enable the computing device to perform any of the above methods.
Claims
1. A multi-robot loop closure detection method based on bimodal verification, characterized in that: The following steps are involved: S1. Build a multi-task deep semantic enhancement network based on deep learning to extract semantic and geometric features of images; S2. Distinguish dynamic and static areas in the image based on the extracted semantic features and assign lower weights to dynamic areas. Multi-robot images are initially fused based on similarity calculated based on the weighted areas. S3. Screen out cross-robot candidate loops that satisfy geometric consistency, semantic similarity, and spatiotemporal continuity based on the preliminarily fused images of multiple robots; S4. Construct a conflict graph based on the inter-robot candidate loops screened out in step S3, and select the largest independent set to eliminate redundant loops.
2. The multi-robot loop closure detection method based on bimodal verification according to claim 1, characterized in that: In step S1, the multi-task deep semantic enhancement network adopts a shared encoder and a dual-branch structure. The shared encoder includes three convolutional layers, a batch normalization layer and a ReLU activation function connected in sequence. The dual-branch structure is a semantic decoding network and a geometric feature decoding network. The semantic decoding network structure is the same as that of the shared encoder. The geometric feature decoding network includes an adaptive average pooling layer and a fully connected layer connected in sequence.
3. The multi-robot loop closure detection method based on bimodal verification according to claim 2, characterized in that: The loss function calculation formula of the multi-task deep semantic enhancement network in step S1 is: in, is the loss function of the multi-task deep semantic enhancement network, λ is a hyperparameter, and They are the loss functions of the semantic decoding network and the geometric feature decoding network respectively; The calculation formula is Where N is the total number of pixels, C is the number of categories, and are the true label and predicted value of the i-th pixel in category c respectively; The calculation formula is Where T is the number of triplets, represents the features of the i-th group of anchor images, represents the positive sample features of group i, represents the negative sample feature of the i-th group, and α is the interval constant.
4. The multi-robot loop closure detection method based on bimodal verification according to claim 1, characterized in that: In step S3, the loop closure condition is determined to be satisfied according to the following conditions: the Mahalanobis distance of different robot loop closure frames is not greater than the corresponding threshold, the cosine similarity of different robot loop closure frames is not less than the corresponding threshold, and the cosine similarity of different robot loop closure frames or several frames before and after them is greater than the corresponding threshold.
5. The multi-robot loop closure detection method based on bimodal verification according to claim 4, characterized in that: The calculation formula of the Mahalanobis distance is: Among them, D M (x,y) is the Mahalanobis distance, x and y are the spatial coordinate vectors of the feature points of different robot loop frames, ∑ -1 is the information matrix; The formula for calculating cosine similarity is: Among them, CosSim(f a ,f b ) is the cosine similarity, f a and f b are the global geometric feature vectors of loop frames a and b, respectively. Loop frames a and b belong to different robots. The formula for determining space-time continuity is: Among them, τ s is the semantic similarity threshold, and The loop frame a and the set of several frames before and after it and the loop frame and the set of several frames before and after it. The specific expression formula is Among them, v i Loopback frame i, v i-k and v i+k are the kth frames before and after loop frame i, respectively, and k is a set constant.
6. The multi-robot loop closure detection method based on bimodal verification according to claim 1, characterized in that: The step S4 comprises the following sub-steps: S4.
1. Construct a conflict graph of candidate loop nodes based on the nodes in the inter-robot candidate loops selected in step S3 and select the maximum independent set. S4.
2. Based on the loop candidate nodes filtered out in step S4.1, candidate loop edges are constructed. Based on this, a conflict graph of the candidate loop edges is constructed and a maximum independent set is filtered out. The candidate loop edges in the maximum independent set are the loop edges finally detected between robots.
7. The multi-robot loop closure detection method based on bimodal verification according to claim 6, characterized in that: The conflict graph of the loop candidate nodes in step S4.1 is expressed as G = (V, E) Among them, G is the conflict graph of the loop candidate node; V is the vertex set of the conflict graph, and its expression is V={l1,l2,...,l n }, l n is the nth loop candidate node, n is a constant; E is the edge set of the conflict graph, and its expression is E={(l i ,l j )|Overlap(l i ,l j )>τ conflict }, where Overlap(l i ,l j ) is loop l i and l j The proportion of shared feature points, τ conflict is the conflict threshold.
8. The multi-robot loop closure detection method based on bimodal verification according to claim 6, characterized in that: The expression formula of the conflict graph of the candidate loop edge in step S4.2 is: G c =(V c ,E c ) Among them, G c is the conflict graph of candidate loop edges; V c is the vertex set of the conflict graph, which is expressed as V={e1,e2,...,e n }, e n is the nth candidate loop edge, n is a constant; E c is the edge set of the conflict graph, and its expression is in, and Represent the candidate loop edge e i and e j There are time conflicts, space conflicts and geometric contradictions between them.
9. A multi-robot loop detection system based on dual-modal verification, characterized in that: The system comprises: Network building module: used to build a multi-task deep semantic enhancement network based on deep learning to extract semantic and geometric features of images; Preliminary fusion module: This module is used to distinguish dynamic and static areas in the image based on the extracted semantic features and assign a lower weight to the dynamic area. The multi-robot images are preliminarily fused based on the similarity calculated based on the weighted areas. Preliminary screening module: used to screen out cross-robot candidate loops that meet geometric consistency, semantic similarity, and spatiotemporal continuity based on the preliminarily fused images of multiple robots; Final screening module: used to construct a conflict graph based on the cross-robot candidate loops screened by the preliminary screening module to screen out the maximum independent set and eliminate redundant loops.
10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 8.