A method for generating a map, a computing device, and a computer storage medium.
By detecting loop closures in the target depth image and candidate images, a target factor graph model is generated, which solves the problem of insufficient map accuracy in existing technologies and achieves more accurate map generation and robot autonomous navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI HAOHAI STARRY SKY ROBOT CO LTD
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Maps generated by existing technologies lack sufficient accuracy in terms of scale and global consistency, making it difficult to meet the accuracy requirements of practical applications.
By determining the target depth image, identifying target dynamic objects, generating target candidate images, and performing loop closure detection, a target factor graph model is generated based on the detection results to represent the constraint relationship between variables in the map under loop closure conditions, and finally, the target map is generated.
It achieves more accurate map generation, improves map scale and global consistency, and enhances the accuracy of robot autonomous navigation.
Smart Images

Figure CN122492962A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of map construction technology, and in particular to a map generation method, computing device, and computer storage medium. Background Technology
[0002] With the rapid development of robotics technology, simultaneous localization and mapping (SLAM) technology has become a core capability for realizing autonomous navigation of robots.
[0003] Maps generated by existing technologies lack accuracy in terms of scale and global consistency, making it difficult to meet the accuracy requirements of practical applications. Summary of the Invention
[0004] The purpose of this application is to provide a map generation method, computing device, and storage medium that can accurately determine the target factor map model, thereby helping to generate target maps more accurately through the target factor map.
[0005] To achieve the above objectives, embodiments of this application provide a map generation method, including: Determine the target depth image; Based on the target depth image, determine the target candidate image; Based on the target depth image and the target candidate image, loop closure detection is performed; When it is determined that the loop state is in effect, a target factor graph model is generated based on the detection results. The target factor graph model is used to characterize the constraint relationship between variables in the map under the loop state. A target map is generated based on the target factor graph model and the target depth image.
[0006] In one embodiment, acquiring the target depth image includes: Acquire environmental images; The environmental image is converted into an initial depth image using a preset neural network model. The initial depth image is used to represent the distance information of at least one element in the environmental image. Based on the environmental image and the initial depth image, identify the target dynamic object; The region where the target dynamic object is located is processed to determine the target depth image.
[0007] In one embodiment, determining the target candidate image based on the target depth image includes: Determine at least one feature point in the target depth image; Based on at least one feature point in the target depth image, determine multiple candidate images of adjacent frames; The feature vectors of the target depth image and the plurality of candidate images are compared; Based on the comparison results, the target candidate image is determined from the multiple candidate images.
[0008] In one embodiment, the step of performing loop closure detection based on the target depth image and the target candidate image includes: The target depth image is matched with the target candidate image to determine the matching loop closure point pairs that match between the target depth image and the target candidate image; Based on the determined loop closure matching point pairs, determine whether the current state is in a loop closure state.
[0009] In one embodiment, determining whether the current state is in a loop based on the determined loop closure matching point pair includes: The number of loop closure matching point pairs is obtained, and it is determined whether the number of loop closure matching point pairs is greater than a preset number threshold and whether the loop closure matching point pairs are in a preset distribution state. If so, then it is determined that it is in a loop state; If not, then it is determined that the system is not in a loop state.
[0010] In one embodiment, generating a target factor graph model based on the detection results when it is determined that the loop is closed includes: Obtain the initial factor graph model; Based on at least one of the loop closure matching point pairs, a loop closure relative pose constraint between the target depth image and the target candidate image is determined, wherein the loop closure relative pose constraint characterizes the relative positional relationship between the target depth image and the target candidate image; The first target constraint edge is determined based on the relative pose constraint of the loop, and the initial factor graph model is updated based on the first target constraint edge to determine the target factor graph model.
[0011] In one embodiment, generating a target factor graph model based on the detection results when it is determined that the loop is closed further includes: The initial scaling factor of the image is determined by performing plane fitting based on the target depth image; A global depth factor is determined based on the depth value of at least one of the loop-matching point pairs and the initial scaling factor, wherein the global depth factor characterizes the relationship between the actual size of the object and the current map scale. Based on the global depth factor, determine the second target constraint edge, and update the initial factor graph model based on the second target constraint edge to determine the target factor graph model.
[0012] In one embodiment, determining the global depth factor based on the depth value of at least one of the loop-matching point pairs and the initial scaling factor includes: Obtain a first depth value of at least one loop-matching point in the target depth image, and a second depth value of at least one loop-matching point in the target candidate image; Perform pose transformation on at least one loop matching point corresponding to the second depth value to determine the third depth value; Determine the depth error value based on the first depth value and the third depth value; The initial scaling factor is adjusted based on the depth error value to determine the global depth factor.
[0013] This application also provides a computing device, specifically including: a processor and a memory for storing executable instructions; the processor is configured to execute the instructions for performing the map generation method as described above.
[0014] This application also provides a computer-readable storage medium storing a computer program, wherein when the instructions in the computer-readable storage medium are executed by a processor of a computing device, the computing device is able to implement any of the map generation methods described above.
[0015] This application provides a map generation method, computing device, and computer-readable storage medium. The method includes: determining a target depth image; determining target candidate images based on the target depth image; performing loop closure detection based on the target depth image and the target candidate images; when a loop closure is detected, generating a target factor graph model based on the detection results, wherein the target factor graph model is used to characterize the constraint relationships between variables in the map under the loop closure state; and generating a target map based on the target factor graph model and the target depth image. The technical solution of this application, through loop closure detection, can accurately determine the target factor graph model, thereby facilitating more accurate generation of the target map. Attached Figure Description
[0016] Figure 1 This is a schematic flowchart illustrating the map generation method provided in an embodiment of the present invention.
[0017] Figure 2 This is a schematic diagram illustrating the specific process of determining the target depth image provided in an embodiment of the present invention.
[0018] Figure 3 This is a schematic diagram illustrating the specific process of determining the target factor graph model provided in an embodiment of the present invention.
[0019] Figure 4 This is a schematic diagram illustrating the specific process of determining the target map provided in an embodiment of the present invention.
[0020] Figure 5This is a schematic diagram of the structure of a computing device provided in an embodiment of the present invention.
[0021] Processor 510, memory 511, network interface 512, bus system 513. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0023] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0024] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, can be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" as used herein are to be interpreted as inclusive, or mean any one or any combination thereof. Therefore, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B, and C". Exceptions to this definition will only occur if the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0025] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0026] It should be noted that step designations such as S101 and S102 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S102 first and then S101, etc., but these should all be within the protection scope of this application.
[0027] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0028] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0029] like Figure 1 As shown, the map generation method provided in this application embodiment can be implemented using software and / or hardware. This embodiment takes the map generation method applied to a server as an example. The map generation method provided in this application embodiment includes the following steps: Step S101: Determine the target depth image.
[0030] Optionally, a depth map is a two-dimensional image where the value of each pixel represents the distance between the scene point corresponding to that pixel and the camera, and visually displays the depth information of different locations in the scene in the form of an image. For example, in a depth map, white or bright areas may represent objects closer to the camera, while black or dark areas may represent parts farther away from the camera, and the depth difference is reflected by pixel grayscale values or color encoding.
[0031] Optionally, the current environment image can be acquired by a monocular color camera, and the depth information of the acquired monocular color image can be inferred by a pre-trained depth estimation neural network to obtain a depth image corresponding to the environment image. At the same time, an inference acceleration engine can be used to optimize the forward inference of the depth estimation neural network to improve the depth estimation efficiency, thereby determining the target depth image in real time.
[0032] In one embodiment, acquiring a target depth image includes: Acquire environmental images; The environmental image is converted into an initial depth image through a pre-defined neural network model. The initial depth image is used to represent the distance information of at least one element in the environmental image. Identify dynamic objects based on environmental and initial depth images; The region containing the target dynamic object is processed to determine the target depth image.
[0033] Optionally, image information of the surrounding environment can be acquired by a monocular color camera. Then, the environmental image is converted into an initial depth image using a preset neural network model. To improve real-time performance, an inference acceleration engine can be used to optimize the network, which can quickly convert the monocular image into a depth map. This initial depth image can represent the distance information of at least one element in the environmental image.
[0034] Optionally, to improve the purity of the constructed image, a lightweight semantic segmentation model can be integrated. For example, the semantic segmentation model can be used to identify the categories of dynamic objects in the image, such as pedestrians and vehicles. At the same time, for areas not covered by the segmentation model, motion vectors can be calculated using optical flow. If the depth values of a certain region change drastically over multiple consecutive frames, or if the motion vector calculated by optical flow is large, it is considered a dynamic region.
[0035] Optionally, the processing methods for identified dynamic objects include: directly removing the depth points of the corresponding area; or, if an area is determined to be a dynamic area, reducing its map update weight when building the map, and even directly removing areas with extremely drastic depth changes or excessive motion vectors, so as to ensure that the final target depth image can be used to build a static map with higher purity.
[0036] Specifically, such as Figure 2 As shown, the specific processing steps for determining the target depth image include: Step S201: Input environment image.
[0037] Optionally, images of the surrounding environment can be acquired using a monocular color camera to provide raw data for the subsequent generation of target depth images. These images contain various information about the scene and serve as the starting point for the entire processing flow.
[0038] Step S202: Image normalization / scaling processing.
[0039] Optionally, the input environmental image can be normalized and scaled. Normalization can bring the image data into a uniform scale range, which helps the neural network to train stably and infer efficiently. Scaling adjusts the image size according to the input requirements of the subsequent lightweight inference network to adapt to the network structure and improve processing efficiency.
[0040] Step S203: The lightweight inference network is used for processing.
[0041] Alternatively, lightweight inference networks can be used for processing, such as pre-trained deep estimation neural networks optimized by an inference acceleration engine, and lightweight semantic segmentation models can be used in dynamic object recognition. By using lightweight networks, the computational load can be effectively reduced while ensuring a certain level of accuracy, thus meeting real-time requirements.
[0042] Step S204: Non-maximum suppression and mask morphology optimization.
[0043] Optionally, non-maximum suppression can be used for optimization. Specifically, the predicted masks are first sorted in descending order according to their confidence scores, and masks with high confidence scores are retained first. Then, the remaining masks are compared with the high-scoring masks in terms of overlap, and low-scoring masks that highly overlap with the high-scoring masks are removed. This removes duplicate and redundant masks, reduces computational overhead, avoids mask boundary conflicts, and ensures that only one accurate predicted mask is retained for the same dynamic object, thereby improving the stability and accuracy of subsequent processing.
[0044] By performing morphological operations, such as erosion and dilation, on the mask after nonmaximum suppression, the mask shape is optimized to more accurately fit the boundary of dynamic objects, providing more precise mask information for subsequent processing.
[0045] Step S205: Depth change detection and optical flow consistency verification.
[0046] Optionally, based on the principle that depth changes in static scenes are caused by camera motion and are geometrically predictable, N consecutive frames of depth maps and poses are recorded. A predicted depth map is obtained after pose compensation, and the depth residual is calculated. If the proportion of pixels with depth residuals greater than a threshold exceeds a certain percentage in multiple consecutive frames, the region is marked as a dynamic region.
[0047] Optionally, the actual optical flow between adjacent frames is compared with the geometric motion field derived from the camera pose and depth. The actual optical flow is calculated using a lightweight optical flow network or a traditional algorithm. The optical flow is then predicted by using the pose increment and depth map. The optical flow residual is calculated. If the residual is greater than a threshold and is consistent in two consecutive frames, it is determined to be a dynamic region.
[0048] Step S206: Output dynamic object mask and static region mask.
[0049] Optionally, the results of nonmaximum suppression, mask morphology optimization, depth change detection, and optical flow consistency verification can be combined to output clear dynamic object masks and static region masks, thus distinguishing between dynamic and static regions in the scene.
[0050] Step S207: Determine the target depth image of the static region.
[0051] Optionally, based on the static region mask, the depth information corresponding to the static region is extracted from the initial depth image to determine the target depth image, which is then used to construct a clean static map and avoid interference from dynamic objects in map construction.
[0052] Please continue reading. Figure 1 Step S102: Determine the target candidate image based on the target depth image.
[0053] Optionally, deep learning global descriptors are extracted from keyframes, and target candidate images are quickly retrieved from historical keyframes using vector retrieval indexes. These target candidate images are historical image frames that are similar to the target depth image at the feature level, and therefore have potential matching possibilities in loop closure detection. This provides possible matching objects for subsequent loop closure detection, reduces the search space, and improves detection efficiency.
[0054] In one embodiment, determining a target candidate image based on a target depth image includes: Determine at least one feature point in the target depth image; Based on at least one feature point in the target depth image, determine multiple candidate images of adjacent frames; The feature vectors of the target depth image and multiple candidate images are compared. Based on the comparison results, the target candidate image is determined from multiple candidate images.
[0055] Optionally, a feature point detection algorithm can be used to perform a full-image scan of the target depth image to extract unique and stable feature points. These feature points are used to determine key local information of the image, preventing feature matching failures due to subtle image changes. After acquiring feature points, redundant, blurry, or low-contrast invalid feature points (such as feature points in blurred edge areas or false feature points generated by noise interference) are removed, retaining valid feature points with clear texture and well-defined coordinate information to ensure the accuracy of subsequent feature tracking and matching. Typically, by calculating parameters such as the feature point's response value and contrast, a filtering threshold is set to select at least one feature point that meets the requirements. Generally, multiple feature points can be extracted to improve data reliability.
[0056] Optionally, key information for each feature point after filtering can be recorded, including pixel coordinates, depth information, local descriptors, etc.
[0057] Optionally, when determining multiple candidate images for adjacent frames, a feature tracking algorithm (such as optical flow tracking) is first used to match feature points in adjacent frames of the image sequence based on the determined feature points in the target depth image, tracking the positional changes of the same feature point in adjacent frames. By analyzing the displacement changes of the feature points, the relative motion parameters between adjacent frames are estimated, including the camera's translational displacement difference and rotation angle. When the detected camera motion (including translational displacement difference and rotation angle) exceeds a preset threshold, or the inter-frame time interval reaches a preset standard (such as 10 milliseconds), the current frame is determined as a candidate image. The selection scope of candidate images mainly focuses on adjacent keyframes and adjacent ordinary frames of the target depth image, ensuring that the candidate images and the target images have strong temporal and spatial correlation, reducing interference from irrelevant images.
[0058] Optionally, feature vectors are used to represent the global information of an image. By analyzing the feature vectors of the target depth image and the candidate image, the correlation between them can be quantified. Specifically: First, the image undergoes preprocessing operations, including scaling (adjusting the image to a uniform size to avoid scale differences affecting feature extraction) and normalization (standardizing the image pixel values to eliminate the influence of illumination variations and pixel intensity differences, making feature extraction more stable). Second, the preprocessed image is input into a predefined deep convolutional neural network. Through the network's forward propagation process, high-level semantic features of the image are extracted, capturing the global information of the image rather than local details, thus improving the representational power of the feature vectors. Then, the features output by the neural network are aggregated using predefined aggregation methods, such as local aggregation descriptor vector alignment and generalized average pooling, to integrate scattered feature information into a unified feature set, solving the problems of inconsistent feature dimensions and scattered information, and enhancing the discriminative power of the feature vectors.
[0059] Finally, the aggregated features are normalized to map the feature vectors onto a unit sphere, eliminating the influence of feature vector scale differences and ensuring the accuracy of subsequent similarity comparisons. The final output is a fixed-dimensional feature vector that uniquely represents the global features of the corresponding image.
[0060] Optionally, after extracting the feature vectors of the target depth image and all candidate images, a standardized similarity calculation method is used to compare the similarity between the feature vector of the target depth image and the feature vector of each candidate image one by one. When comparing the feature vectors between the target depth image and the candidate images, cosine similarity or Euclidean distance can be used for comparison. Optionally, based on the comparison results, the candidate images with similarity exceeding a preset threshold among multiple candidate images are determined as the target candidate images, or the candidate images with the strongest correlation / most similarity can also be determined as the target candidate images.
[0061] In this way, by extracting target depth image feature points, filtering adjacent candidate images, and combining a standardized feature vector extraction and comparison process, the target candidate image can be determined efficiently and accurately, which helps to improve the reliability of inter-frame association.
[0062] Step S103: Perform loop closure detection based on the target depth image and the target candidate image.
[0063] Optionally, loop closure detection refers to the process of verifying the feature correlation and consistency between two images to determine whether the current camera position corresponding to the target depth image is the same as the historical camera position corresponding to the target candidate image. Specifically, based on the feature vectors extracted from both images, their global and local features can be compared more precisely, and depth information can be used for verification to eliminate interference from illumination, scale changes, and noise, ultimately determining whether the image is in a loop closure state.
[0064] By using loop closure detection, the drift error accumulated during camera pose estimation can be effectively corrected, avoiding excessive pose deviations over long-term operation; it reduces invalid and repeated mapping, improving the consistency and accuracy of map construction; it reduces computational redundancy, optimizes the overall system operating efficiency, and enhances the stability and reliability of image processing and pose estimation.
[0065] In one embodiment, loop closure detection is performed based on the target depth image and the target candidate image, including: Match the target depth image with the target candidate image to identify the matching loop closure point pairs in the target depth image and the target candidate image; Based on the determined loop closure matching point pairs, determine whether the current state is in a loop closure state.
[0066] Optionally, local feature points and their descriptors are extracted from the target depth image and the target candidate image, respectively. Initial matching point pairs are obtained through brute-force matching or fast approximate nearest neighbor matching. Then, geometric verification is performed using epipolar constraints (based on the fundamental matrix or the essential matrix). The number of interior points that satisfy the epipolar geometry is counted. If the number of interior points exceeds a preset threshold (e.g., no less than 20), the target candidate image and the effective interior points are retained as preliminary matching point pairs, and most of the mismatches are eliminated.
[0067] Furthermore, using the initial matched point pairs and camera intrinsic parameters, the relative pose transformation between two frames is solved by combining the perspective N-point pose solving algorithm or the five-point relative pose estimation algorithm with the random sampling consensus algorithm. The pose is checked to see if it meets the physical constraints, including whether the rotation angle and translation distance are within a reasonable range. At the same time, the reprojection error is verified (the error of the 3D points of the candidate frame projected onto the current frame is less than the threshold). After the verification is completed, the remaining matched point pairs are determined as loop closure matched point pairs.
[0068] In one embodiment, determining whether the current state is in a loopback state based on the determined loopback matching point pairs includes: Obtain the number of loop closure matching pairs, determine whether the number of loop closure matching pairs is greater than a preset threshold, and whether the loop closure matching pairs are in a preset distribution state; If so, then it is determined that it is in a loop state; If not, then it is determined that the system is not in a loop state.
[0069] Optionally, a preset threshold can be determined based on the currently detected scene type, image complexity, etc., typically set to 20 frames. Thus, if the number of acquired loop closure matching point pairs exceeds the preset threshold, it is determined that the matching degree between the target depth image and the target candidate image is high, indicating that the current state is a loop closure. Otherwise, it is determined that the current state is not a loop closure, and image acquisition and detection processing is required.
[0070] Step S104: When it is determined that the loop state is in effect, a target factor graph model is generated based on the detection results. The target factor graph model is used to characterize the constraint relationship between variables in the map under the loop state.
[0071] Optionally, the target factor graph model is an online, incremental method for constructing and solving factor graphs, widely used in scenarios requiring real-time processing of streaming data, such as robot synchronous localization and mapping. The factor graph is a bipartite graph containing two types of nodes: variable nodes, representing the state variables to be estimated, such as the robot's pose at various times (x_1, x_2, ...), landmark positions, etc.; and factor nodes, representing constraints between variables derived from sensor measurements (such as odometry, loop closure detection, GPS, depth consistency factors, etc.).
[0072] In one embodiment, upon determining that a loop closure state is in place, a target factor graph model is generated based on the detection results, including: Obtain the initial factor graph model; Based on at least one loop closure matching point pair, determine the loop closure relative pose constraint between the target depth image and the target candidate image. The loop closure relative pose constraint characterizes the relative positional relationship between the target depth image and the target candidate image. The first target constraint edge is determined based on the relative pose constraint of the loop, and the initial factor graph model is updated based on the first target constraint edge to determine the target factor graph model.
[0073] Optionally, a preset initial factor graph model is obtained. The initial factor graph model is a mathematical data structure containing variable nodes (such as camera pose, ground position, etc.). Figure 3 The system uses dimensional points and factor nodes (such as inter-frame pose constraints and observation constraints) and implements incremental optimization based on a Bayesian tree structure, updating only locally rather than re-optimizing the entire graph.
[0074] Optionally, the relative pose constraint of the loop closure is determined based on the loop closure matching point pairs. Using the geometrically verified loop closure matching point pairs, combined with camera intrinsic parameters and depth information, the relative rotation and translation relationship between the target depth image and the target candidate image is calculated to form the relative pose constraint of the loop closure. This constraint is a strong geometric constraint, representing that the two frames are at the same or adjacent positions in real physical space.
[0075] Optionally, the first target constraint edge is generated based on the loop closure constraint, and the target factor graph is updated. The aforementioned loop relative pose constraint is transformed into the first target constraint edge in the factor graph, which is then added as a new global constraint factor to the initial factor graph. Under the incremental update mechanism, only the affected variable nodes in the Bayesian tree are relinearized and locally updated, without optimizing the entire graph trajectory. Finally, a target factor graph model containing the loop global constraint is formed, providing a constraint basis for subsequent backend optimization to correct pose drift.
[0076] In one embodiment, the specific processing method for determining the target factor graphical model is as follows: Figure 3 As shown, it includes the following steps: Step S301: Determine candidate images.
[0077] Optionally, candidate images are image frames selected from historical keyframe images that have the potential to perform loop closure matching with the current image. Specifically, based on the image acquisition sequence, scene spatial correlation, and image content similarity, several images to be compared can be initially screened from the historical mapping keyframe set as candidate images for subsequent loop closure detection, thereby narrowing the scope of subsequent feature retrieval and matching.
[0078] Step S302: Extract global feature vectors.
[0079] Optionally, the global feature vector is used to characterize the global semantic and appearance attributes of the overall scene in a single frame image. Global features are extracted from the current image and each candidate image to generate corresponding global feature vectors, which abstractly represent the overall scene features of the image, weaken the interference caused by local texture and lighting changes, and provide feature basis for subsequent image retrieval and similarity comparison.
[0080] Step S303: Index and obtain multiple adjacent images.
[0081] Optionally, based on the extracted global feature vector, a nearest neighbor search is performed in the historical keyframe library through a feature retrieval index structure. Multiple adjacent candidate images with high similarity to the global features of the current image are obtained in batch queries. The feature matching and screening of massive historical frames can be completed quickly through the index retrieval method, and the set of images suspected of loop closure association can be efficiently locked.
[0082] Step S304: Determine whether the interval exceeds 30 frames. If yes, proceed to step S303; otherwise, proceed to step S312.
[0083] Optionally, the frame interval between the current keyframe and historical candidate frames is used as the criterion. If the frame interval exceeds a preset threshold of 30 frames, it is considered to have loop closure detection value, and the subsequent adjacent image index retrieval process continues. If the frame interval does not reach 30 frames, it is determined to be a short temporal continuous frame with no obvious possibility of overlapping loop closure scenes. Step S305: Candidate target frame list.
[0084] Optionally, suspected loop closure images retrieved through global feature indexing and time-series frame interval filtering are uniformly collected and organized into a candidate target frame list. This list provides a well-organized set of objects to be processed for subsequent fine feature matching, geometric verification, and pose determination, enabling orderly management and batch processing of candidate frames.
[0085] Step S306: Feature matching.
[0086] Optionally, feature matching is a process of comparing local features of the current image with each candidate image in the candidate target frame list. Specifically, local feature points and corresponding descriptors of the current image and candidate images are extracted respectively. Initial feature point pair matching is completed through feature matching algorithm. Then, erroneous point pairs are eliminated by combining geometric constraints, and effective matching point pairs with stable correlation are selected, providing a reliable feature association basis for subsequent camera relative pose estimation.
[0087] Step S307: Estimate the relative pose of the camera.
[0088] Optionally, based on the obtained effective matching feature point pairs, combined with the camera intrinsic parameters, a perspective N-point pose solving algorithm or a five-point relative pose estimation algorithm is adopted, and a random sampling consensus algorithm is integrated for robust solving. The relative rotation matrix and translation vector between the camera corresponding to the current image and the camera corresponding to the candidate image are calculated to obtain the relative pose relationship between the cameras between the two frames, thereby realizing the cross-frame spatial position association solution.
[0089] Step S308: Determine whether the number of internal points exceeds 30 frames. If yes, proceed to step S309; otherwise, proceed to step S312.
[0090] Optionally, the number of inliers can be used as a quantitative indicator of the effectiveness of loop closure matching. The number of valid matching inliers remaining after geometric constraint verification and reprojection error verification is counted. If the number of inliers exceeds a preset threshold of 30, the current candidate frame is considered to have high credibility in matching the current image and has effective loop closure association features, and enters the loop closure status confirmation process. If the number of inliers does not reach the threshold, no matching is performed.
[0091] Step S309: Determine the processing loop status.
[0092] Optionally, based on the number of effectively matched inliers meeting a preset threshold, the spatial relationship between the current frame and candidate frames is comprehensively verified by combining feature matching consistency, relative pose physical rationality, and reprojection error constraints. Finally, it is determined that the current system is in an effective loop closure state, confirming that the historical position and the current position overlap in the scene, providing a basis for subsequent construction of loop closure constraints and backend optimization to correct pose drift.
[0093] Step S310: Generate closure constraints.
[0094] Optionally, after determining that the system is in a loop-closed state, a loop-closed relative pose constraint between the two frames is constructed based on the solved relative pose relationship between the current keyframe and the candidate keyframe. This constraint represents the global geometric association between the camera pose and the landmark points at different time sequences, and serves as a global loop-closed constraint condition to correct the pose drift error accumulated during the mapping process.
[0095] Step S311: Update and determine the target factor graph model.
[0096] Optionally, the initial factor graph model of the system is obtained, and the generated lap-loop relative pose constraints are transformed into constraint edges in the factor graph and connected to the initial factor graph in the form of new global constraint factors. Relying on the factor graph incremental optimization mechanism, only the variable nodes affected by the lap-loop constraints are locally updated and relinearized, without the need for full graph re-optimization. Finally, the target factor graph model incorporating the lap-loop constraints is updated to achieve global consistency optimization of robot pose and map structure.
[0097] Continue reading Figure 1 In one embodiment, when it is determined that the loop is closed, generating a target factor graph model based on the detection results further includes: The initial scaling factor of the image is determined by performing plane fitting based on the target depth image; The global depth factor is determined based on the depth value of at least one loop-matched point pair and the initial scaling factor. The global depth factor represents the relationship between the actual size of the object and the current map scale. Based on the global depth factor, determine the second objective constraint edge, and update the initial factor graph model based on the second objective constraint edge to determine the objective factor graph model.
[0098] Optionally, noise points and abnormal depth values are removed from the target depth image, and the extracted effective depth points are fitted using a plane fitting algorithm, such as the least squares method, to determine the dominant planes in the scene, such as the ground and walls. The initial scale relationship of the scene is then inferred through the plane equations, thereby determining the initial scaling factor. This initial scaling factor is not a fixed value but only serves as an initial reference for the absolute scale, allowing for continuous iterative optimization based on joint optimization of the factor graph.
[0099] Optionally, the global depth factor is a constraint factor that characterizes the correspondence between the actual size of an object and the current map scale. Its calculation depends on the depth information of the loop-matched point pairs and the initial scaling factor. The calculated global depth factor is transformed into the second target constraint edge in the target factor graph model, which can be added in parallel with the first target constraint edge obtained based on the loop-matched relative pose constraint and added to the initial factor graph model together. When updating the target factor graph model, the local update mechanism of the iSAM2 algorithm can be used to re-linearize and locally update only the variable nodes in the Bayesian tree affected by the global depth factor (such as camera pose and 3D map point scale), without optimizing the entire historical trajectory. This allows the reuse of the historical optimization results of unaffected nodes, thereby achieving pose drift correction and accurate optimization of map scale. By determining the second target constraint edge and optimizing the target factor graph model, the pose drift problem can be solved, ensuring that the target factor graph model can accurately characterize the constraint relationship between map variables in the loop-matched state, providing comprehensive constraint support for subsequent joint optimization.
[0100] In one embodiment, determining the global depth factor based on the depth values of at least one loop-matched point pair and an initial scaling factor includes: Obtain a first depth value for at least one loop-matching point in the target depth image, and a second depth value for at least one loop-matching point in the target candidate image; Perform pose transformation on at least one loop matching point corresponding to the second depth value to determine the third depth value; The depth error value is determined based on the first depth value and the third depth value; The initial scaling factor is adjusted based on the depth error value to determine the global depth factor.
[0101] Optionally, when a loop closure is determined, at least one set of actual observed depths of loop closure matching points are extracted from the target depth image and determined as the first depth value; simultaneously, the original depth data of the corresponding matching points are extracted from the target candidate image and determined as the second depth value. To ensure the accuracy of the calculation, inlier points in the loop closure matching can be preferentially selected, i.e., point pairs with high matching accuracy and no obvious deviation. Usually, multiple sets of matching point pairs are selected to avoid the influence of single data deviation on the results.
[0102] Next, the pose transformation of the second depth value is performed to determine the third depth value. Based on the relative pose parameters of the target depth image and the target candidate image obtained from the previous loop closure detection, the coordinate transformation and projection transformation of the second depth value of the matching point in the candidate image are performed, and mapped to the observation coordinate system of the target depth image to obtain the predicted depth after pose derivation, that is, the third depth value.
[0103] Optionally, the first depth value in the target depth image is used as the actual observed depth, and the third depth value obtained after pose transformation is used as the predicted depth. The error of each pair of matching points is calculated using the formula "depth error value = first depth value - third depth value". The larger the absolute value of the error, the more obvious the depth scale deviation between the two frames.
[0104] Optionally, using the initial scaling factor as the initial value, the depth error value is used as an optimization constraint term to construct a minimum error objective function. The initial scaling factor is continuously adjusted through an iterative optimization algorithm until the depth error value tends to converge. The scaling factor obtained at this time is the global depth factor. The global depth factor can unify the global depth scale and work in conjunction with the pose constraint to achieve global consistency of scene depth information.
[0105] Step S105: Generate a target map based on the target factor map model and the target depth image.
[0106] Optionally, the target map includes a two-dimensional map and a three-dimensional map. Optionally, after obtaining globally consistent and scale-accurate keyframe poses through factor map optimization, the local depth information is projected onto the global space and fused with the corresponding depth image to construct a target map that can be used for navigation and modeling, as shown in the following example. Figure 4 As shown, it includes the following steps: Step S401: Determine the target factor model and target depth image.
[0107] Optionally, the optimized target factor graph model is loaded, and the final pose of each keyframe in the world coordinate system is calculated from it, including position and orientation information. Simultaneously, the target depth image corresponding to the keyframe is obtained, containing the depth value of each pixel.
[0108] Step S402: Transform to the world coordinate system.
[0109] Optionally, the 3D points corresponding to each pixel in the depth image are transformed from the camera local coordinate system to the global world coordinate system using camera intrinsic parameters and optimized keyframe poses.
[0110] Step S403, Output type.
[0111] Optionally, the corresponding map output format can be selected according to the actual navigation task or environmental modeling application requirements: for navigation application scenarios, a two-dimensional occupancy grid map is preferred; for environmental modeling, path planning and obstacle avoidance application scenarios, a three-dimensional point cloud map or an octree map is preferred.
[0112] Step S404: Project onto the ground plane.
[0113] Optionally, three-dimensional spatial points are vertically projected onto a set ground plane, compressing the 3D spatial information into 2D planar information for constructing a planar grid for robot navigation.
[0114] Step S405: Update the 2D occupied grid map.
[0115] Optionally, based on the two-dimensional point cloud data obtained from projection, the occupancy status of each grid area is determined one by one: if there is an obstacle within the grid area, the grid is marked as occupied; if there is no obstacle within the grid area, the grid is marked as vacant. By continuously accumulating and iteratively updating multiple frames of sensor data, a globally consistent two-dimensional occupancy grid map is finally constructed and obtained.
[0116] Step S406: Save as PGM+YAML.
[0117] Optionally, the completed 2D raster map is stored in a common standard format. Portable Gray Map (PGM) image files are used to store the black and white raster texture image of the map. YAML is a readable configuration markup language for data serialization; the configuration file records key configuration information such as map resolution, coordinate origin, and scale parameters. This combined format is a common standard map storage format adapted to the robot navigation stack.
[0118] Step S407: Voxel filtering downsampling.
[0119] Optionally, voxel mesh downsampling can be performed on the global 3D point cloud to reduce the number of points, reduce memory usage, improve subsequent processing speed, and preserve spatial structure.
[0120] Step S408: Merge into the global point cloud / OctoMap.
[0121] Optionally, the current frame can be converted into 3D point cloud data in the world coordinate system and fused into the global map model. The global point cloud map can construct an intuitive and visual 3D environment model. OctoMap, or octree map, is a compact 3D occupancy probability map that is suitable for robot path planning and obstacle avoidance scenarios.
[0122] Step S409: Save as a PCD / BT file.
[0123] Optionally, the merged map data can be stored in separate formats: PCD files are in a standard 3D point cloud storage format; BT files are in an OctoMap-specific storage format. The saved files can be used for subsequent map loading, scene reproduction, offline mapping, and robot navigation applications.
[0124] In summary, the map generation method provided in the above embodiments can accurately determine the target factor map model through loop closure detection, thereby helping to generate the target map more accurately through the target factor map.
[0125] Based on the same inventive concept as the foregoing embodiments, this embodiment of the invention provides a computing device, such as... Figure 5 As shown, the computing device includes: a processor 510 and a memory 511 storing computer programs; wherein, Figure 5 The processor 510 shown in the diagram does not indicate that there is only one processor 510, but only indicates the positional relationship of the processor 510 relative to other devices. In practical applications, there can be one or more processors 510; similarly, Figure 5 The memory 511 shown in the diagram has the same meaning, that is, it is only used to indicate the positional relationship of memory 511 relative to other devices. In practical applications, there can be one or more memories 511. When the processor 510 runs the computer program, the above-described map generation method is implemented.
[0126] The computing device may also include at least one network interface 512. The various components of the computing device are coupled together via a bus system 513. It is understood that the bus system 513 is used to implement communication between these components. In addition to a data bus, the bus system 513 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 5 The general designated all buses as Bus System 513.
[0127] The memory 511 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 511 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0128] The memory 511 in this embodiment of the invention is used to store various types of data to support the operation of the computing device. Examples of this data include: any computer programs used to operate on the computing device, such as operating systems and applications; contact data; phonebook data; messages; pictures; videos, etc. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications, such as media players, browsers, etc., used to implement various application services. Here, the program implementing the method of this embodiment of the invention can be included in the application.
[0129] Based on the same inventive concept as the foregoing embodiments, this embodiment also provides a computer-readable storage medium storing a computer program. The computer-readable storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc. When the computer program stored in the computer-readable storage medium is executed by a processor, it implements the map generation method applied to the aforementioned computing device. For the specific steps implemented when the computer program is executed by the processor, please refer to [link to relevant documentation]. Figure 1 The description of the illustrated embodiments will not be repeated here.
[0130] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0131] In this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.
[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method of generating a map, characterized by, include: Determine the target depth image; Based on the target depth image, determine the target candidate image; Based on the target depth image and the target candidate image, loop closure detection is performed; When it is determined that the loop state is in effect, a target factor graph model is generated based on the detection results. The target factor graph model is used to characterize the constraint relationship between variables in the map under the loop state. A target map is generated based on the target factor graph model and the target depth image.
2. The method of claim 1, wherein, The acquisition of the target depth image includes: Acquire environmental images; The environmental image is converted into an initial depth image using a preset neural network model. The initial depth image is used to represent the distance information of at least one element in the environmental image. Based on the environmental image and the initial depth image, identify the target dynamic object; The region where the target dynamic object is located is processed to determine the target depth image.
3. The method of claim 1, wherein, The step of determining the target candidate image based on the target depth image includes: Determine at least one feature point in the target depth image; Based on at least one feature point in the target depth image, determine multiple candidate images of adjacent frames; The feature vectors of the target depth image and the plurality of candidate images are compared; Based on the comparison results, the target candidate image is determined from the multiple candidate images.
4. The method of claim 1, wherein, The step of performing loop closure detection based on the target depth image and the target candidate image includes: The target depth image is matched with the target candidate image to determine the matching loop closure point pairs that match between the target depth image and the target candidate image; Based on the determined loop closure matching point pairs, determine whether the current state is in a loop closure state.
5. The method of claim 4, wherein, The step of determining whether the current state is in a loop based on the determined loop closure matching point pair includes: The number of loop closure matching point pairs is obtained, and it is determined whether the number of loop closure matching point pairs is greater than a preset number threshold and whether the loop closure matching point pairs are in a preset distribution state. If so, then it is determined that it is in a loop state; If not, then it is determined that the system is not in a loop state.
6. The method of claim 4, wherein, When it is determined that the loop is closed, the step of generating a target factor graph model based on the detection results includes: Obtain the initial factor graph model; Based on at least one of the loop closure matching point pairs, a loop closure relative pose constraint between the target depth image and the target candidate image is determined, wherein the loop closure relative pose constraint characterizes the relative positional relationship between the target depth image and the target candidate image; The first target constraint edge is determined based on the relative pose constraint of the loop, and the initial factor graph model is updated based on the first target constraint edge to determine the target factor graph model.
7. The method of claim 6, wherein, The step of generating a target factor graph model based on the detection results when determining that the loop closure state is reached further includes: The initial scaling factor of the image is determined by performing plane fitting based on the target depth image; A global depth factor is determined based on the depth value of at least one of the loop-matching point pairs and the initial scaling factor, wherein the global depth factor characterizes the relationship between the actual size of the object and the current map scale. Based on the global depth factor, determine the second target constraint edge, and update the initial factor graph model based on the second target constraint edge to determine the target factor graph model.
8. The method of claim 7, wherein, The step of determining the global depth factor based on the depth value of at least one of the loop-matching point pairs and the initial scaling factor includes: Obtain a first depth value of at least one loop-matching point in the target depth image, and a second depth value of at least one loop-matching point in the target candidate image; Perform pose transformation on at least one loop matching point corresponding to the second depth value to determine the third depth value; Determine the depth error value based on the first depth value and the third depth value; The initial scaling factor is adjusted based on the depth error value to determine the global depth factor.
9. A computing device, comprising: include: A processor and a memory for storing executable instructions; wherein the processor is configured to execute the instructions to implement the map generation method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor, the method for generating a map as described in any one of claims 1-8 is implemented.