Map updating method, system, device, medium and product based on visual slam system
Patent Information
- Application Number
- CN202411143076.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-08-20
AI Technical Summary
[0005]本申请的目的是提供一种基于视觉SLAM系统的地图更新方法、系统、设备、介质及产品,以解决现有VSLAM系统无法在未知的动态环境中高效完成定位以及建图任务,导致系统跟踪失败甚至崩溃的的问题
[0022]根据本申请提供的具体实施例,本申请公开了以下技术效果:本申请利用动态目标剔除线程,基于深度神经网络的目标检测算法识别场景图像中的动态目标,通过特征点运动概率传播方法有效剔除动态目标的特征点(即动态特征点),并设计了追踪线程、局部建图线程、回环检测线程以及全局光束法平差(BundleAdjustment,BA)线程,基于保留的静态特征点与局部地图中静态点之间的匹配关系,更新局部地图以及相机的相机位姿图,最终确定更新后的地图结构。本申请将SLAM与基于深度神经网络的目标检测算法集成在一起,在执行SLAM任务时,系统只会根据场景中的静态物体进行定位与建图,从而能够在未知的动态环境中可靠高效地完成定位以及建图的任务,避免了因为相机采集的图像出现动态物体的情况时,导致SLAM系统跟踪丢失的问题,进而提高了SLAM系统的鲁棒性。
Smart Images

Figure CN119027686B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot vision technology, and in particular to a map updating method, system, device, medium and product based on a visual SLAM system. Background Technology
[0002] Simultaneous localization and mapping (SLAM) systems utilize sensor input data to calculate the robot's relative pose (position and orientation) transformations in real time, achieving robot localization and simultaneously incrementally building a map of the observed scene during robot movement. With technological advancements, human demands on robots are constantly increasing, aiming to make robots intelligent and automated, thereby enabling them to complete complex tasks. Since its inception, SLAM has attracted widespread attention from researchers and is considered a crucial approach to solving the problem of autonomous robot navigation.
[0003] Based on the different sensors used, SLAM can be divided into Visual Simultaneous Localization and Mapping (VSLAM) and LiDAR SLAM. LiDAR sensors measure distance accurately with small errors, operate stably in direct light, and offer relatively simple point cloud processing. However, their high cost and energy consumption limit their widespread commercial application. Cameras, on the other hand, offer advantages such as low cost and rich image information, making them well-suited for environments with similar geometric structures and for handling loop closure detection. Therefore, Visual SLAM, using cameras as the primary sensor, has received widespread attention and made some progress, but it still faces significant challenges.
[0004] Most existing VSLAM algorithms assume a static environment and ignore dynamic objects. However, in real-world scenarios, moving objects such as vehicles, people, and animals can cause VSLAM systems to mismatch feature points, resulting in poor localization accuracy. This makes it difficult for the system to efficiently complete localization and mapping tasks in unknown dynamic environments, leading to tracking failures or even system crashes. Summary of the Invention
[0005] The purpose of this application is to provide a map update method, system, device, medium and product based on a visual SLAM system to solve the problem that existing VSLAM systems cannot efficiently complete localization and mapping tasks in unknown dynamic environments, leading to system tracking failure or even crash.
[0006] To achieve the above objectives, this application provides the following solution:
[0007] Firstly, this application provides a map updating method based on a visual SLAM system, including:
[0008] Using a dynamic target removal thread, a target detection algorithm based on a deep neural network is used to identify dynamic targets in a scene image, and a feature point motion probability propagation method is used to remove the feature points of the dynamic targets, while retaining the static feature points in the scene image.
[0009] Using a tracking thread, the static feature points are matched with static points in the local map, and each frame of the scene image is tracked to estimate the camera pose and the depth of the landmark points in the local map.
[0010] Using a local mapping thread, based on the landmark depth, optimize the landmarks and camera poses in the local map, and determine the optimized local map and camera pose map;
[0011] Using a loop closure detection thread, loop closures are detected based on the DBoW2 model, and the accumulated camera trajectory error is corrected by executing the camera pose graph to determine the optimized camera pose graph.
[0012] Using a global BA thread, the optimized local map is kept globally consistent with the optimized camera pose map, the optimized local map is updated, and the updated map structure is determined.
[0013] Secondly, this application provides a map update system based on a visual SLAM system, including:
[0014] The dynamic target removal thread is used to identify dynamic targets in scene images based on a deep neural network-based target detection algorithm, and to remove the feature points of the dynamic targets using a feature point motion probability propagation method, while retaining the static feature points in the scene image.
[0015] The tracking thread is used to match the static feature points with static points in the local map, track each frame of the scene image, and estimate the camera pose and the depth of the landmark points in the local map.
[0016] A local mapping thread is used to optimize the landmarks and camera poses in the local map based on the landmark depth, and to determine the optimized local map and camera pose map.
[0017] The loop closure detection thread is used to detect loop closures based on the DBoW2 model, and to correct the cumulative error of the camera trajectory by executing the camera pose graph, and to determine the optimized camera pose graph.
[0018] A global BA thread is used to keep the optimized local map and the optimized camera pose map globally consistent, update the optimized local map, and determine the updated map structure.
[0019] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the map update method based on the visual SLAM system described above.
[0020] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the map update method based on the visual SLAM system described above.
[0021] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the map update method based on the visual SLAM system described above.
[0022] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application utilizes a dynamic target culling thread to identify dynamic targets in scene images based on a deep neural network-based target detection algorithm. It effectively culls feature points (i.e., dynamic feature points) of dynamic targets through a feature point motion probability propagation method. Furthermore, it designs a tracking thread, a local mapping thread, a loop closure detection thread, and a global bundle adjustment (BA) thread. Based on the matching relationship between the retained static feature points and static points in the local map, it updates the local map and the camera pose map, ultimately determining the updated map structure. This application integrates SLAM with a deep neural network-based target detection algorithm. When performing SLAM tasks, the system only performs localization and mapping based on static objects in the scene, thus reliably and efficiently completing localization and mapping tasks in unknown dynamic environments. This avoids the problem of SLAM system tracking loss when dynamic objects appear in the images captured by the camera, thereby improving the robustness of the SLAM system. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1This is a flowchart of a map update method based on a visual SLAM system in one embodiment of this application;
[0025] Figure 2 A detailed illustration of motion probability propagation provided for an embodiment of this application;
[0026] Figure 3 This is a schematic diagram illustrating the minimization of reprojection error between two adjacent camera frames, provided in an embodiment of this application.
[0027] Figure 4 This is a schematic diagram of a partial BA optimization provided in an embodiment of this application;
[0028] Figure 5 This is a diagram illustrating an improved ORB-SLAM2 algorithm framework provided in one embodiment of this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] In recent years, the rapid development of deep learning has also greatly accelerated the innovation of object detection algorithms. Since VSLAM systems are often prone to failure in dynamic environments, in order to avoid the problem of camera tracking loss in dynamic scenes, using object detection algorithms to obtain the label information of the surrounding environment for subsequent removal of dynamic objects has become a research direction.
[0032] This application provides a map update method based on a visual SLAM system, such as... Figure 1 As shown, the map update method based on the visual SLAM system includes:
[0033] Step 101: Using a dynamic target removal thread, a target detection algorithm based on a deep neural network is used to identify dynamic targets in the scene image, and the feature points of the dynamic targets are removed by using a feature point motion probability propagation method, while retaining the static feature points in the scene image.
[0034] In an exemplary embodiment, step 101 can be replaced by steps 201-203, which specifically include the following steps.
[0035] Step 201: Use the YOLO v8 network to preprocess and propagate the scene image forward. At the same time, propagate the motion probability frame by frame in the tracking thread. When a dynamic target is detected, insert the image frame corresponding to the dynamic target as a key frame into the local map, and update the motion probability of the landmark points in the local map corresponding to the key points in the key frame.
[0036] Step 202: Estimate the motion probability of each key point in the local map by feature matching and matching point expansion.
[0037] Step 203: When a key point matches a key point in the previous frame, determine that the motion probability of the key point has been propagated, remove the key point whose motion probability has been propagated, and retain the static feature points in the scene image; wherein, the key point is a feature point of a dynamic target.
[0038] In one exemplary embodiment, a deep neural network-based object detection algorithm is used to detect dynamic objects. The YOLO v8 network is employed to identify pedestrians, animals, and other dynamic or potentially dynamic objects that could negatively impact map reuse in images captured by the camera. For example, once a person is detected, whether walking or standing still, it is considered a potential dynamic object, and feature points within the detected area will be removed from the image.
[0039] Furthermore, Oriented Fast and Rotated Brief (ORB) feature points are extracted from the input camera images. These ORB feature points consist of Features from Accelerated Segment Test (FAST) keypoints and Binary Robust Independent Elementary Features (BRIEF) descriptors. After the images acquired by the camera are input into the SLAM algorithm framework, ORB feature points from both images are extracted in the preprocessing module. The original input image data is then discarded, and all subsequent algorithmic operations are based on the extracted ORB feature points. The image representation is transformed from a set of pixels to a set of ORB feature points, reducing the amount of data cached during runtime. Specifically, ORB feature point extraction can be implemented using a cross-platform computer vision library (OpenBomputerVision Library, OpenCV). First, the positions of the FAST keypoints are detected, and then the BRIEF descriptors are calculated based on these positions.
[0040] Further, the probability that a feature point belongs to a dynamic object is defined as the motion probability of the feature point. The feature points in the image are divided into four states according to the motion probability P. When the motion probability P of a feature point is greater than 0.75, it is called a high-confidence dynamic point; when 0.5 < P < 0.75, it is called a low-confidence dynamic point; when 0.25 < P < 0.5, it is called a low-confidence static point; when P < 0.25, it is called a high-confidence static point.
[0041] Dynamic targets are eliminated by a feature point motion probability propagation method, and two strategies are proposed: 1) detecting dynamic objects in a detection frame, and updating the motion probability of landmarks in a local map by using the detection results; 2) propagating the motion probability through feature matching and extension of matching points, and efficiently removing features extracted on dynamic objects before camera pose estimation.
[0042] First, the input camera image is preprocessed and forward propagated through the YOLO v8 network, and the motion probability is propagated frame by frame in the tracking thread. Once the detection result is obtained, the frame is inserted into the local map as a key frame, and the motion probability in the local map is updated. The feature points in the key frame are called key points, and the motion probability of the 3D landmark corresponding to the matching key point in the key frame is updated according to the following formula:
[0043] P t (X i )=(1-α)P t-1 (X i )+αS t (x i )
[0044] wherein, P t-1 (X i )is the updated motion probability of the 3D landmark X t-1 in the previous key frame I i . If it is a new key point, P t-1 (X i )=P init =0.5 will be set. The state of the matching key point x t in the key frame I i is S t (x i ), which depends on the dynamic target detection area output by the YOLO v8 network. If the key point x i is within the bounding box of a dynamic object, it is regarded as a definite dynamic point, and its state value S t (x i )=1. Other points are regarded as definite static points, and the state value is S t (x iα = 0. α is an influence factor used to smooth immediate detection results. A higher value means greater sensitivity to immediate detection results, while a lower value means considering more historical results from multiple perspectives.
[0045] Furthermore, after step 202, the method further includes: based on the propagated key points, extending the motion probability of the key points from high-confidence points to neighboring key points that did not find a corresponding matching point in the previous frame during feature matching; the high-confidence points include high-confidence dynamic points and high-confidence static points; the high-confidence dynamic points are feature points with a motion probability greater than 0.75, and the high-confidence static points are feature points with a motion probability less than 0.25; extending the influence area of high confidence to a circular area of a set radius, and determining the motion probability of neighboring key points within the circular area that did not find a corresponding matching point in the previous frame.
[0046] Reference Figure 2 The diagram illustrates detailed information about motion probability propagation. The motion probability of each keypoint is estimated through feature matching and matching point expansion, i.e., feature point motion probability propagation. This is because the motion probability in the current frame is only propagated from the keypoints of the previous frame.
[0047] During feature matching, when a key point Key points from the previous frame During matching, motion probability It was spread.
[0048] Once a key point is matched with any 3D landmark on the local map The match is also assigned a motion probability, which is equal to the 3D landmark of the match. The value of . Note that if a point not only has a matching point in the previous frame but also a matching 3D landmark in the current map, the motion probability corresponding to the 3D landmark should be given priority. At the same time, assign an initial probability P to other unmatched points in this frame. init =0.5, because there are no prior assumptions about which state these points belong to.
[0049] In summary, the operation of using feature point matching to propagate motion probabilities is summarized in the following formula:
[0050]
[0051] in, and They represent The ORB feature points are θ, where θ is the threshold for feature matching.
[0052] Furthermore, since the states of keypoints in a certain neighborhood are consistent in most cases, this is taken as a strong assumption. Based on this strong assumption, the motion probability is extended from high-confidence points (including high-confidence dynamic points and high-confidence static points) to other neighboring keypoints that do not have corresponding matching points in the feature matching operation.
[0053] After feature matching propagation, a high-confidence point χt is selected. The influence area of the high-confidence point is expanded to a circular region with radius r, and unmatched keypoints are searched within this region. The unmatched keypoints found (i.e., neighboring keypoints within the circular region) are then identified. The probability of movement is updated according to the following rules:
[0054]
[0055] in, For road signs The probability of movement; P init This is the initial motion probability, set to 0.5. If a point is influenced by multiple high-confidence points, the sum of the influences of all neighboring high-confidence points will be calculated. When considering the influence of high-confidence points, the difference in motion probabilities needs to be taken into account. And the distance factor λ(d). If a point is within the influence region of a high-confidence point (d≤r), the distance factor λ(d) = Ce -d / r , where C is a constant value; otherwise (d>r), λ(d)=0.
[0056] Step 102: Using the tracking thread, perform feature point matching between the static feature points and static points in the local map, and track each frame of scene image to estimate the camera pose and the depth of the landmark points in the local map.
[0057] In an exemplary embodiment, the camera pose is estimated by minimizing the reprojection error in two adjacent frames using the Perspective-n-Point (PnP) algorithm, and the depth of landmark points in the local map is calculated using triangulation. The method for determining whether the current frame is a keyframe remains the same as the rules of ORB-SLAM2.
[0058] In an exemplary embodiment, the camera pose is initialized and estimated. A constant-velocity camera model is used to initialize the camera pose. It is assumed that the camera motion is uniform, and the relative motion between frames remains constant. Based on the relative pose between frames t-2 and t-1, the pose at the current time t frame is initialized from frame t-1. Since the initialization of the pose assumes uniform camera motion, which is not the case in reality, this assumption is rather coarse and requires further optimization by tracking a local map to refine the pose estimation.
[0059] Furthermore, refer to Figure 3 As shown, based on ORB feature point matching between two adjacent frames, and by retrieving ORB feature points in the local map that match the current frame of the camera, a PnP problem is constructed to minimize the reprojection error of feature points in the current frame, thereby optimizing the current camera pose and obtaining a more accurate optimized camera pose T.
[0060] Furthermore, by utilizing the optimized camera pose T and the matched ORB feature points from the two frames, triangulation is performed to calculate a large number of new landmark depth values, thereby obtaining more high-quality landmarks. These new landmarks, as a supplement to the local map, greatly improve the accuracy and robustness of the tracking thread.
[0061] Finally, set the keyframe selection criteria to determine whether the current frame should be set as a keyframe. If the current frame meets any of the following conditions, then set the current frame as a keyframe.
[0062] 1. The current frame is the first frame after the bionic eye gaze control reaches the target gaze area.
[0063] 2. More than 13 frames have passed since the last global relocation.
[0064] 3. The local mapping thread is in an idle state.
[0065] 4. More than 15 frames have passed since the last keyframe was set.
[0066] 5. The translation distance between the current frame and the previously set keyframe exceeds the threshold t. th .
[0067] 6. The number of successfully tracked feature points in the current frame reaches 70 or more.
[0068] 7. The number of successfully tracked feature points in the current frame is less than 85% of that in the reference keyframe.
[0069] Step 103: Using the local mapping thread, optimize the landmarks and camera poses in the local map based on the landmark depth, and determine the optimized local map and camera pose map.
[0070] In an exemplary instance, step 103 specifically includes: using a local mapping thread, updating the keyframes and landmarks in the local map based on the landmark depth, and using the PnP algorithm to perform joint nonlinear optimization on the landmarks and camera poses in the updated local map by minimizing the reprojection error, thereby determining the optimized local map and camera pose map.
[0071] In an exemplary embodiment, a local mapping thread is designed to manage a local map and perform local BA optimization, with the aim of updating and maintaining landmarks and keyframes in the local map, as well as optimizing the poses of keyframes and the coordinates of landmarks through local BA.
[0072] First, the keyframes in the local map are updated. After the tracking thread determines that the current frame is set as a keyframe, the new keyframe is associated with the landmark co-view relationship between it and the previous keyframes, and the bag-of-words representation of the new keyframe is calculated based on the Dispersed Bag of Words (DBoW) model.
[0073] Furthermore, landmarks in the local map are updated and maintained. For a landmark to remain in the local map, it must meet two conditions within the first three keyframes after its creation: 1. More than 25% of the keyframes in which the landmark is visible based on its pose prediction must be successfully tracked; 2. If more than one new keyframe is added after the landmark's creation, the landmark must be observed in at least three keyframes. Once a landmark meets these two conditions, it will only be removed if the number of subsequent keyframes observing it is less than three.
[0074] Furthermore, local BA optimization is essentially a PnP problem, incorporating the coordinates of landmarks as parameters in the optimization process. (See reference...) Figure 4 As shown, BA is explained from the perspective of graph optimization. Figure 4 In the diagram, circles C1, C2, and C3 represent camera pose nodes, indicating the camera pose parameters to be optimized; circles P1, P2...P7 represent landmark nodes, indicating the 3D coordinate parameters of the landmarks to be optimized. The lines connecting the nodes represent error terms defined in the nonlinear optimization process. The defined error terms are:
[0075]
[0076] Where u2 represents the observed coordinates of 3D landmark point P, e represents the error between the projected coordinates and observed coordinates of 3D landmark point P, i.e., reprojection error, K represents the camera intrinsic parameters, and T represents the camera extrinsic parameters.
[0077] During iterative optimization, the derivative of the reprojection error e with respect to the camera pose is used:
[0078]
[0079] Where δξ represents the left perturbation of the camera pose T, X',Y',Z' represent the coordinates of the 3D landmark transformed into the camera coordinate system, fx represents the focal length of the camera in the x-direction, and fy represents the focal length of the camera in the y-direction.
[0080] Then, take the derivative of the reprojection error e with respect to the 3D landmark point P:
[0081]
[0082] Where R is the rotation matrix in the camera pose.
[0083] After obtaining the derivative of the reprojection error e with respect to the camera pose T and the coordinates of the 3D landmark P, the objective function (i.e., the reprojection error e) is optimized using the Gauss-Newton method or the Levenberg-Marquardt method. The gradient direction guides the update of the optimization variables, camera pose T and the coordinates of the 3D landmark P in the map. After iterating until the objective function error converges, the optimal keyframe camera pose and landmark coordinates in the local map can be obtained.
[0084] Finally, discard local keyframes. A keyframe is discarded if 90% of its landmarks can be observed in at least three other keyframes, as this is redundant. Failure to limit the number of keyframes will lead to an ever-increasing data size in the local map building process, slowing down optimization and impacting the real-time performance of local mapping.
[0085] Step 104: Utilize the loop closure detection thread to detect loop closures based on the DBoW2 model, and correct the cumulative error of the camera trajectory by executing the camera pose graph to determine the optimized camera pose graph.
[0086] In an exemplary embodiment, step 104 specifically includes: detecting loop closures based on DPoW2 to find candidate keyframes for loop closures, and calculating the relative pose between the current keyframe and the candidate keyframes for loop closures; closing the loop closures according to the co-view relationship, and correcting the loop closures based on the relative poses between the current keyframe and the candidate keyframes for loop closures; constructing a nonlinear optimization problem using the camera poses before and after loop closure correction, calculating the optimization gradient, and then performing graph optimization on the camera pose to correct the cumulative error of the camera trajectory and determine the optimized camera pose graph.
[0087] In an exemplary embodiment, a loop closure thread is implemented based on DPoW2 to detect large loop closures and correct the accumulated camera trajectory error by performing camera pose graph optimization. First, loop closures are detected using DPoW2 to find candidate keyframes, and the relative pose between the current keyframe and the candidate keyframes is calculated. Then, the loop is closed based on the co-view relationship. Finally, the loop is corrected based on the relative pose between the current keyframe and the candidate keyframes calculated in the previous step, and camera pose graph optimization is performed.
[0088] Camera pose graph optimization only optimizes the camera pose. The nodes in the camera pose graph are simply camera pose nodes, while the edges connecting the nodes represent estimates of the relative poses between two camera pose nodes. The initial values of the pose nodes are the camera poses of each keyframe before loop closure correction, while the edges represent the relative poses between the camera poses of each keyframe calculated after loop closure correction. Assume there are K... i K j The two keyframes, with camera poses T before loop closure correction, are respectively... Wi and T Wj After correcting the camera poses of keyframes in the loop fusion, the obtained K... i and K j The relative pose between them is T ij Then the error e ij for:
[0089]
[0090] Based on this error, a nonlinear optimization is constructed, and the error terms are solved with respect to T. Wi and T Wj The derivative of the equation can be used for optimization, which can then be used to optimize the camera pose graph.
[0091] Step 105: Using a global BA thread, keep the optimized local map and the optimized camera pose map globally consistent, update the optimized local map, and determine the updated map structure.
[0092] In an exemplary embodiment, step 105 specifically includes: merging the optimized camera pose with the unoptimized camera pose using a spanning tree, and correcting the coordinates of landmarks in the optimized local map based on the update of the reference camera pose, thereby determining the updated map structure.
[0093] In an exemplary embodiment, a global BA thread is initiated to obtain globally consistent camera poses and map structure. After the loop closure thread completes camera pose graph optimization, global BA optimization is performed in a separate thread to obtain the globally optimal solution. If a new loop is detected during global BA optimization, the global BA thread is terminated, and it is restarted after the loop closure camera pose graph optimization is completed. During global BA thread optimization, the updated keyframe camera poses are merged with the unupdated keyframe camera poses using a spanning tree, while the coordinates of landmark points are corrected based on the update of their reference keyframe camera poses.
[0094] Based on the same inventive concept, embodiments of this application also provide a map update system based on a visual SLAM system, including:
[0095] The dynamic target removal thread is used to identify dynamic targets in scene images based on a deep neural network-based target detection algorithm, and to remove the feature points of the dynamic targets using a feature point motion probability propagation method, while retaining the static feature points in the scene image.
[0096] The tracking thread is used to match the static feature points with static points in the local map, track each frame of the scene image, and estimate the camera pose and the depth of the landmark points in the local map.
[0097] A local mapping thread is used to optimize the landmarks and camera poses in the local map based on the landmark depth, and to determine the optimized local map and camera pose map.
[0098] The loop closure detection thread is used to detect loop closures based on the DBoW2 model, and to correct the cumulative error of the camera trajectory by executing the camera pose graph, thereby determining the optimized camera pose graph.
[0099] A global BA thread is used to keep the optimized local map and the optimized camera pose map globally consistent, update the optimized local map, and determine the updated map structure.
[0100] like Figure 5 As shown, this invention employs an improved ORB-SLAM2 algorithm framework, adding a dynamic target culling thread and integrating SLAM with the YOLO v8 object detection network. It utilizes a feature point motion probability propagation method to cullate dynamic targets in the scene, thereby constructing a visual SLAM system for dynamic target culling, aiming to improve the robot's ability to localize and map in unknown dynamic environments. Specifically, this invention proposes two strategies for dynamic target culling: 1) detecting dynamic objects in keyframes and using the detection results to update the motion probability of landmarks in the local map; 2) propagating motion probabilities through feature matching and the expansion of matching points, efficiently removing feature points extracted from dynamic objects before camera pose estimation. These innovations avoid the problem of tracking loss caused by a large number of dynamic objects in the images acquired by the camera, improving the stability of the system when performing SLAM tasks. The method of this invention has significant practical application value and can be widely applied in fields such as robot navigation and autonomous driving, providing strong support for technological innovation and progress in related industries.
[0101] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores map update data based on a visual SLAM system. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a map update method based on a visual SLAM system.
[0102] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0103] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0104] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0105] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0106] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0107] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0109] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A map update method based on a visual SLAM system, characterized in that, The map update method based on the visual SLAM system includes: A dynamic target removal thread is used to identify dynamic targets in a scene image based on a deep neural network-based target detection algorithm. The feature points of the dynamic targets are then removed using a feature point motion probability propagation method, while retaining static feature points in the scene image. Specifically, this includes: The scene image is preprocessed and propagated forward using a YOLO v8 network. Simultaneously, motion probabilities are propagated frame by frame in the tracking thread. When a dynamic target is detected, the image frame corresponding to the dynamic target is inserted into the local map as a key frame, and the motion probabilities of the landmark points in the local map corresponding to the key points in the key frame are updated. The motion probability of each key point is estimated in the local map by feature matching and matching point expansion; When a key point matches a key point in the previous frame, the motion probability of the key point is determined to be propagated, the key point with the propagated motion probability is removed, and the static feature points in the scene image are retained; wherein, the key point is the feature point of the dynamic target. Using a tracking thread, the static feature points are matched with static points in the local map, and each frame of the scene image is tracked to estimate the camera pose and the depth of the landmark points in the local map. Using a local mapping thread, based on the landmark depth, optimize the landmarks and camera poses in the local map, and determine the optimized local map and camera pose map; Using a loop closure detection thread, loop closures are detected based on the DBoW2 model, and the accumulated camera trajectory error is corrected by executing the camera pose graph to determine the optimized camera pose graph. Using a global BA thread, the optimized local map is kept globally consistent with the optimized camera pose map, the optimized local map is updated, and the updated map structure is determined.
2. The map update method based on a visual SLAM system according to claim 1, characterized in that, Updating the motion probability of the landmark points in the local map corresponding to the key points in the keyframe specifically includes: According to the formula Update the motion probability of the landmark points in the local map corresponding to the key points in the keyframe; wherein, The probability of movement of the landmarks in the local map; For the previous keyframe Central Road Marker Updated motion probabilities; Impact factor; Key point The state.
3. The map update method based on a visual SLAM system according to claim 1, characterized in that, The motion probability of each keypoint in the local map is estimated through feature matching and matching point expansion, specifically including: According to the formula Extend the estimation of the motion probability of each key point in the local map; wherein, for t Time of the first i Key points The probability of movement; for t -1 moment i Key points The probability of movement; for t Time of the first i Road signs corresponding to each key point The probability of movement; The initial probability; for t Time of the first i Key points ORB feature points; for t -1 moment i Key points ORB feature points; for t Time of the first i Road signs corresponding to each key point ORB feature points; The threshold for feature matching.
4. The map update method based on a visual SLAM system according to claim 1, characterized in that, The motion probability of each keypoint in the local map is estimated through feature matching and matching point expansion, followed by: Based on the propagated key points, the motion probability of the key points is extended from high-confidence points to neighboring key points that did not find a corresponding matching point in the previous frame during feature matching; the high-confidence points include high-confidence dynamic points and high-confidence static points; the high-confidence dynamic points are feature points with a motion probability greater than 0.75, and the high-confidence static points are feature points with a motion probability less than 0.25; The high-confidence influence area is extended to a circular area of a set radius, and the motion probability of neighboring key points within the circular area that did not find a corresponding matching point in the previous frame is determined.
5. The map update method based on a visual SLAM system according to claim 4, characterized in that, The motion probability of adjacent key points within the circular area for: in, The initial probability; for t Time of the first j One key point; Distance factor; For road signs The probability of movement; This is a high confidence point.
6. The map update method based on a visual SLAM system according to claim 1, characterized in that, Estimating the camera pose and the depth of landmarks in the local map specifically includes: In two adjacent frames, the camera pose is estimated by minimizing the reprojection error based on the PnP algorithm, and the depth of landmarks in the local map is calculated by triangulation.
7. The map update method based on a visual SLAM system according to claim 1, characterized in that, Using a local mapping thread, based on the landmark depth, optimize the landmarks and camera poses in the local map to determine the optimized local map and camera pose map, specifically including: Using a local mapping thread, keyframes and landmarks in the local map are updated based on the landmark depth. Then, based on the PnP algorithm, the landmarks and camera poses in the updated local map are jointly nonlinearly optimized by minimizing the reprojection error to determine the optimized local map and camera pose map.
8. The map update method based on a visual SLAM system according to claim 1, characterized in that, Using a loop closure detection thread, loop closures are detected based on the DBoW2 model. The accumulated camera trajectory error is corrected by executing the camera pose graph, and an optimized camera pose graph is determined. Specifically, this includes: Loop closure detection is performed based on DBoW2 to find candidate keyframes for loop closure, and the relative pose between the current keyframe and the candidate keyframes for loop closure is calculated. The loop is closed based on the common-view relationship, and the loop is corrected based on the relative pose of the current keyframe and the loop keyframe. Using the camera pose before and after loop closure correction, a nonlinear optimization problem is constructed. After calculating the optimization gradient, graph optimization is performed on the camera pose to correct the cumulative error of the camera trajectory and determine the optimized camera pose graph.
9. The map update method based on a visual SLAM system according to claim 1, characterized in that, Using a global BA thread, the optimized local map is kept globally consistent with the optimized camera pose map. The optimized local map is then updated, and the updated map structure is determined. This process specifically includes: The optimized camera pose is merged with the unoptimized camera pose by generating a tree, and the coordinates of landmarks in the optimized local map are corrected based on the updated reference camera pose to determine the updated map structure.
10. A map update system based on a visual SLAM system, characterized in that, The map update method based on a visual SLAM system according to any one of claims 1-9 includes: The dynamic target removal thread is used to identify dynamic targets in scene images based on a deep neural network-based target detection algorithm, and to remove the feature points of the dynamic targets using a feature point motion probability propagation method, while retaining the static feature points in the scene image. The tracking thread is used to match the static feature points with static points in the local map, track each frame of the scene image, and estimate the camera pose and the depth of the landmark points in the local map. A local mapping thread is used to optimize the landmarks and camera poses in the local map based on the landmark depth, and to determine the optimized local map and camera pose map. The loop closure detection thread is used to detect loop closures based on the DBoW2 model, and to correct the cumulative error of the camera trajectory by executing the camera pose graph, and to determine the optimized camera pose graph. A global BA thread is used to keep the optimized local map and the optimized camera pose map globally consistent, update the optimized local map, and determine the updated map structure.
11. A computer device, comprising: The memory and processor contain a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the map update method based on the visual SLAM system as described in any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the map update method based on the visual SLAM system as described in any one of claims 1-9.
13. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the map update method based on the visual SLAM system as described in any one of claims 1-9.
Citation Information
Patent Citations
A method and apparatus for setting enhanced interactive content
CN109656363A
Image processing method of augmented reality device and augmented reality device
CN110310373A