Three-dimensional mapping method

Through the mask prediction mechanism and the two-stage tracking mechanism, the SLAM technology is optimized, and the problems of graph construction accuracy and computing resources in dynamic environments are solved, and efficient three-dimensional graph construction on low-computing platforms are realized.

CN120298599APending Publication Date: 2025-07-11JILIN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510444230.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing SLAM technology is difficult to effectively distinguish dynamic and static characteristics in a dynamic environment, resulting in reduced graph construction accuracy and strong dependence on computing resources, making it difficult to achieve real-time operation on low-computing platforms.

Method used

The mask prediction mechanism is used to combine the lightweight YOLO-fastest model for dynamic object detection, design a two-stage tracking mechanism and buffer synchronization strategy, optimize computing resource allocation, and realize dynamic object filtering and high frame rate mapping.

Benefits of technology

Effective filtering of dynamic objects and high-precision mapping is realized on the low-computing platform, ensuring the real-time and efficiency of mapping construction, and is suitable for three-dimensional reconstruction of small computing units and large scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298599A_ABST
    Figure CN120298599A_ABST
Patent Text Reader

Abstract

The invention relates to the field of mapping, in particular to a three-dimensional mapping method, which comprises the following steps of: initializing a system; oRB features are extracted from the left camera and the right camera; constructing an initial key frame: sending image data of a first frame of a left camera to a semantic thread, and setting the frame as the initial key frame; performing a mask prediction mechanism, key frame tracking and dual-stage tracking on the image of the left camera, and alternately performing key frame tracking and static tracking on the right camera; and binocular stereo matching is carried out according to results of the left camera and the right camera, map points are updated, and an obtained result is a three-dimensional diagram. The method has practical significance when being practically applied to rapid three-dimensional mapping of a small computing unit or a large scene, dynamic objects are filtered out of the mapping in the mode, the real-time performance and the mapping efficiency of the mapping are guaranteed, three-dimensional mapping can be rapidly completed on a low-computing-power platform, and the method is suitable for large-scale popularization and application. And the method has good practical performance and application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mapping, and specifically to a method for three-dimensional mapping. Background Art

[0002] Currently, vision-based Simultaneous Localization and Mapping (SLAM) technology has been widely used in fields such as robotics, autonomous driving, and augmented reality. As a mainstream open-source solution, ORB-SLAM3 shows high positioning and mapping accuracy in static scenes through ORB feature extraction, key-frame matching, and Bundle Adjustment optimization.

[0003] However, it has significant deficiencies in actual dynamic environments. For example, dynamic object processing defects: traditional SLAM systems do not effectively distinguish between dynamic and static features, resulting in dynamic objects (such as pedestrians and vehicles) interfering with map construction and reducing the system accuracy; strong computing power dependence: existing dynamic SLAM solutions (such as DS-SLAM and DynaSLAM) rely on GPU-accelerated semantic segmentation networks (such as Mask R-CNN) and are difficult to run in real time on a pure CPU platform; low resource efficiency: binocular vision SLAM needs to process left and right camera data simultaneously, and the existing methods have insufficient heterogeneity in processing dual-camera features, resulting in computational redundancy. Summary of the Invention

[0004] Existing SLAM technologies have high computing power requirements for computing platforms, and pure CPU operations are relatively laborious. Moreover, the current application of identifying and filtering dynamic objects also relies on strong computing power support. The current binocular camera vision SLAM is a very promising direction, but it still has certain challenges if it is to be applied on low-computing-power platforms. To achieve the above objectives, the present invention proposes a visual SLAM solution based on a mask prediction mechanism that can filter dynamic objects, which can filter dynamic interference objects and obtain high mapping accuracy on the premise of effectively improving the system computing efficiency. Considering that the semantic thread has poor real-time performance and realizing frequency decoupling, allowing the semantic thread to run at a lower frequency, while using multi-thread parallelization to ensure a high frequency of the tracking thread, it also saves computing time, and designs two buffers to store data to avoid timing disorders between the detection thread and the tracking thread. At the same time, considering using a hybrid feature strategy of interleaving the LK optical flow method and key-frame extraction in the extraction of key frames to improve the tracking efficiency, which can ensure high-frame-rate mapping information and can be achieved only by CPU operation, greatly improving the application prospect of the SLAM solution. The present invention provides the following technical solutions:

[0005] A method for three-dimensional mapping, comprising the following steps:

[0006] The first step, system initialization: initialize both the left camera and the right camera;

[0007] Step 2: Extract the ORB features of the first frame: Extract the ORB features for both the left camera and the right camera;

[0008] Step 3: Construct the initial key frame: Send the image data of the first frame of the left camera to the semantic thread and set this frame as the initial key frame;

[0009] Step 4: For the images of the left camera, adopt the mask prediction mechanism, key frame tracking, and two-stage tracking, and for the right camera, perform key frame tracking and static tracking alternately;

[0010] Step 5: Based on the results of the left camera and the right camera, perform binocular stereo matching, update the map points, and the obtained result is the 3D map.

[0011] As a further solution of the present invention: Before adopting the mask prediction mechanism for the images of the left camera, the original images of the left camera are processed by the semantic thread. This processing is independent of the key frame tracking. The result of the semantic thread processing is placed in the buffer to retain information. The buffer only saves the result of the semantic thread processing of the latest frame. The tracking thread can directly obtain the corresponding mask and semantic information in the buffer without waiting for the output of the semantic thread, which can achieve good real-time performance and efficiency.

[0012] As a further solution of the present invention: The semantic thread processing adopts the YOLO-fastest model as the semantic module for detecting dynamic objects for processing, which can achieve good real-time data transmission.

[0013] As a further solution of the present invention: The mask prediction mechanism includes the following steps: detection; segmentation; motion saliency-driven dynamic point sampling; object-level UKF tracking; spatio-temporal joint clustering; prediction.

[0014] As a further solution of the present invention: The detection specifically uses the YOLO-fastest model alone as a thread to detect the dynamic objects in the images of the left camera and designs a double-buffer mechanism for it. When the detection thread writes to Buffer A, the tracking thread reads Buffer B, and the frame numbers of the two buffers are strictly aligned to avoid the time sequence misalignment between the detection results and the tracking thread.

[0015] As a further solution of the present invention: The segmentation specifically generates a depth map by binocular information disparity calculation within the YOLO detection box, performs DBSCAN clustering based on the depth values within the detection box to separate the foreground and the background, and then performs morphological optimization, performs dilation and closing operations on the clustering results, and fills the holes.

[0016] As a further solution of the present invention: Motion saliency-driven dynamic point sampling specifically performs hierarchical grid sampling on the dynamic region. The overall grid is divided into 10*10 grids, and the low-texture grids are further sampled with 5*5 sub-grids. The FAST corner point with the strongest response is taken for each grid (if there is no corner point, the center point is taken). The magnitude of the optical flow vector is calculated for the remaining points, and only the motion-salient points are retained.

[0017] As a further solution of the present invention: Object-level UKF tracking specifically performs overall modeling on each sampled clustering cluster, predicts the state through unscented Kalman filtering, and generates the observed state and prediction residuals.

[0018] As a further solution of the present invention: Spatiotemporal joint clustering specifically performs single-shot multi-feature clustering again, and fuses spatial, depth, motion information, and UKF prediction residuals for four-dimensional feature clustering to distinguish independent objects.

[0019] As a further solution of the present invention: Prediction specifically means that after clustering, the system will be aware of the position of the dynamic object in the current frame. Therefore, based on the precise mask of the dynamic object generated by clustering, the processed image is used for the subsequent tracking thread.

[0020] The two-stage tracking in the present invention includes dynamic tracking and static tracking. Dynamic tracking is only used for the images of the left camera and adopts a mask prediction mechanism to identify dynamic objects. However, when the mask prediction mechanism fails (although the mask prediction mechanism can achieve good results when identifying dynamic objects, the mask prediction mechanism depends on the output of the YOLO-fastest model, and in some cases, it may lead to detection failure or difficulty in keeping up with real-time performance), the LK optical flow method is used to predict the mask of the current frame using the mask of the previous frame. This design can provide a buffer time for the mask prediction mechanism in some emergencies and improve the robustness of tracking. Static tracking starts from the input frame with the latest ORB feature matching, and uses the LK optical flow method for fast tracking. The three-dimensional map points of the previous frame are linked to the two-dimensional key points tracked in the current frame to form geometric constraints, and then the initial pose of the current frame is calculated based on the motion model. Based on the above information, the RANSAC algorithm is applied to iteratively estimate the actual pose and filter out outliers.

[0021] Key frame tracking retains the ORB feature extraction scheme, establishes data associations between key frames, and provides initial feature points for static tracking. For the left camera, key frame tracking runs in parallel with two-stage tracking. However, due to the large computational resources required for key frame tracking, the processing frequency of key frame tracking is set lower than that of two-stage tracking, and two-stage tracking is used as the normal tracking method. For the right camera, only key frame tracking and static tracking are alternated. While maintaining the same frequency of key frame tracking as the left camera, the rest of the states maintain static tracking, that is, the dynamic tracking and mask prediction mechanisms are not adopted. Although some dynamic feature points will also be tracked in the right camera, since there are no matching feature points in the left camera, it can still filter dynamic objects. In this way, on the basis of ensuring the accuracy of mapping, the efficiency of computational resource allocation is effectively improved, the amount of computation is reduced, so that high-frame-rate mapping can be achieved only by the CPU.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0023] The present invention has practical significance in the actual application to small computing units or large-scale fast 3D mapping. In this way, dynamic objects are filtered out of the mapping, and the real-time performance and mapping efficiency of the mapping are ensured, so that 3D mapping can be quickly completed on a low-computing-power platform, with good practical performance and application potential. Description of the Drawings

[0024] Figure 1 It is a flowchart of the 3D mapping method in the embodiment of the present invention.

[0025] Figure 2 It is a flowchart of the left camera image tracking in the 3D mapping method in the embodiment of the present invention.

[0026] Figure 3 It is a flowchart of the mask prediction mechanism in the 3D mapping method in the embodiment of the present invention. Detailed Embodiments

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] In the binocular vision open-source library based on ORB-SLAM3 (an open-source library for real-time simultaneous localization and mapping), improvement measures for real-time performance and accuracy are made, which we name CSD-SLAM (CPU-efficient StereoDynamic SLAM). The following breakthroughs are achieved:

[0029] 1. Mask prediction mechanism: Combining the lightweight YOLO-fastest model with spatio-temporal clustering to achieve real-time detection and tracking of dynamic objects on the CPU.

[0030] 2. Two-stage heterogeneous tracking: Hierarchical processing of dynamic and static for the left camera, lightweight tracking for the right camera, optimizing the allocation of computing resources.

[0031] 3. Buffer synchronization strategy: Decoupling the semantic thread and the tracking thread through a double-buffer mechanism to ensure real-time performance and data consistency. This scheme can achieve high-frame-rate mapping processing relying only on the CPU based on accurate mapping after filtering out dynamic objects by using binocular pure vision, and has a good mapping effect on small computing units, especially suitable for mapping in the dynamic environment of indoor service robots and 3D reconstruction of large outdoor scenes for drones.

[0032] The following describes the specific implementation of the present invention in detail with specific embodiments.

[0033] Embodiment 1, see Figures 1 - 3 , The first step: System initialization

[0034] Start the binocular camera, load the calibration parameters (focal length, baseline distance, and distortion coefficient), verify the synchronization of binocular images, and complete the initialization of the binocular camera. The first frame of the left camera triggers the YOLO-fastest model, initializes the semantic thread and initializes the double buffer. Extract ORB feature points from the first frames of the left and right cameras, construct the first key frame and initialize the set of static map points.

[0035] The second step: Semantic processing thread (independent thread)

[0036] Input the left camera image, and detect the set of dynamic bounding boxes (confidence threshold 0.6) through the YOLO-fastest model. Calculate the disparity map (using the SAD algorithm, window size 9×9) for each detected bounding box respectively. The calculation method is as follows:

[0037]

[0038] Within the detected bounding box, project the pixel point into the three-dimensional space p i =(x, y, z), where After generating the depth information, apply DBSCAN clustering (eps = 0.3m, min_samples = 15) to segment the foreground and background, and then perform dilation (kernel size 3*3) and closing operations on the segmentation result to fill the holes.

[0039] Stratify and sample the dynamic region (main grid 10×10, sub-grid 5×5), extract FAST corner points (threshold R FAST = 20), and further divide the low-texture grid into 5×5 sub-grids. Calculate the LK optical flow magnitude ‖v‖ of the remaining points, and retain the significant motion points with ‖v‖ > 5 pixel / frame.

[0040] For each clustering cluster C k Define the state vector x k = [c x , c y , v x , v y T , and construct a prediction model

[0041]

[0042] The process noise covariance Q = diag(0.1, 0.1, 0.05, 0.05).

[0043] Construct a four-dimensional feature vector f = [αx, βy, γΔd, δ‖v‖] T , where the weights (α, β, γ, δ) = (0.4, 0.4, 0.1, 0.1). Use k-means clustering (k = 3) to merge regions that are spatially adjacent and have consistent motion. Finally, generate a binary mask M t for the clustering result, and perform dilation (kernel size 3*3, 2 iterations) to cover the object edges, generate the final mask, and write it to the output buffer.

[0044] Step 3: Left camera dynamic tracking

[0045] Read the current mask M from the buffer t , eliminate all ORB feature points in the region where M t (x, y) = 1, and then output them to the next-level operation module. If the semantic thread delay exceeds 33 ms, enable the LK optical flow method, and calculate the optical flow field v = (v t-1 , v x , v y ) based on the previous frame mask M t-1 , and perform an affine transformation on M

[0046] is the optical flow deformation function

[0047] Step 4: Left camera static tracking​

[0048] Both the left camera and the right camera use low-frame-rate ORB feature extraction and high-frame-rate LK optical flow method to obtain feature points. For static tracking, the left camera uses LK optical flow to track the static feature points of the previous frame, that is, starting from the ORB feature points of the nearest key frame KF ref to track to the current frame, and the search area is predicted by a constant velocity model:

[0049] u pred = u t-1 + Δt·v avg

[0050] where v avg is the historical average motion speed.

[0051] Next, the system constructs geometric constraints, projects the map points to the current frame and calculates the reprojection error. Abnormal points greater than 2.5 pixels are removed, and finally RANSAC iterative optimization is applied to solve the initial pose.

[0052] Step 5: Right camera tracking

[0053] With low-frame-rate ORB feature extraction and high-frame-rate LK optical flow method, the right camera also generates feature points for stereo matching of binocular cameras.

[0054] Step 6: Binocular stereo matching

[0055] Perform SGM (Semi-Global Matching) on static feature points to generate a disparity map D sparse . Triangulate the static map points X i :

[0056]

[0057] Perform depth filtering on the map points, and remove unstable points with σ z > 0.2 m based on chi-square test.

[0058] Step 7: Key frame tracking

[0059] Retain the original key frame tracking of ORB-SLAM3. When the static tracking effect is not good in some environments, it can be used as a second tracking scheme to ensure the overall robustness.

[0060] In addition, it should be understood that although this specification is described according to the implementation manners, not every implementation manner only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.

Claims

1. A method for three-dimensional mapping, characterized in that, It includes the following steps: The first step is system initialization: initialize both the left camera and the right camera; The second step is to extract the ORB features of the first frame: extract the ORB features for both the left camera and the right camera; The third step is to construct the initial key frame: send the image data of the first frame of the left camera to the semantic thread and set this frame as the initial key frame; The fourth step is to adopt a mask prediction mechanism, key frame tracking, and two-stage tracking for the images of the left camera, and perform key frame tracking and static tracking alternately for the right camera; The fifth step is to perform binocular stereo matching based on the results of the left camera and the right camera, update the map points, and the obtained result is the 3D map.

2. The method for three-dimensional mapping according to claim 1, characterized in that, Before adopting the mask prediction mechanism for the images of the left camera, the original images of the left camera are also processed by the semantic thread.

3. The method for three-dimensional mapping according to claim 2, characterized in that, The semantic thread processing is to use the YOLO-fastest model as the semantic module for detecting dynamic objects for processing.

4. The method for three-dimensional mapping according to claim 1, wherein The mask prediction mechanism includes the following steps: detection; segmentation; motion saliency-driven dynamic point sampling; object-level UKF tracking; spatio-temporal joint clustering; prediction.

5. The method for three-dimensional mapping according to claim 4, wherein The specific detection is to use the YOLO-fastest model alone as a thread to detect dynamic objects in the images of the left camera, and design a double-buffer mechanism for it. When the detection thread writes to Buffer A, the tracking thread reads Buffer B, and the frame numbers of the two buffers are strictly aligned.

6. The method for three-dimensional mapping according to claim 5, wherein The specific segmentation is to generate a depth map by binocular information disparity calculation within the YOLO detection box, perform DBSCAN clustering based on the depth values within the detection box to separate the foreground and the background, and then perform morphological optimization, perform dilation and closing operations on the clustering result, and fill the holes.

7. The method for three-dimensional mapping according to claim 6, wherein The specific motion saliency-driven dynamic point sampling is to perform hierarchical grid sampling on the dynamic region. The overall is divided into a 10*10 grid and the low-texture grids are sub-sampled with a 5*5 sub-grid. The FAST corner point with the strongest response is taken for each grid, and the optical flow vector amplitude is calculated for the remaining points, and only the motion-significant points are retained.

8. The method for three-dimensional mapping according to claim 7, wherein The specific object-level UKF tracking is to perform overall modeling on each sampled clustering cluster, predict the state through the unscented Kalman filter, and generate the observation state and the prediction residual.

9. The method for three-dimensional mapping according to claim 8, wherein The specific spatio-temporal joint clustering is to perform single multi-feature clustering again, and perform four-dimensional feature clustering by fusing spatial, depth, motion information, and the UKF prediction residual.

10. The method for three-dimensional mapping according to claim 9, characterized in that, The specific prediction is to generate an accurate mask of the dynamic object according to the clustering, and then use the processed image for the subsequent tracking thread.

Citation Information

Cited By

  • Dynamic scene real-time SLAM system fusing target detection and optical flow

    CN121053323A