Visual SLAM method and system based on saliency prediction and storage medium
By introducing saliency prediction technology to guide feature extraction and matching in visual SLAM methods, select key frames and optimize poses, the problem of ignoring effective structured areas in traditional methods is solved, and the stability and performance of the system are improved.
Patent Information
- Application Number
- CN202511063273.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Traditional visual SLAM methods ignore the key role of effective structured areas such as edges and corners in feature extraction, matching and pose estimation, resulting in waste of computing resources and limited system performance improvement.
A visual SLAM method based on saliency prediction is adopted. Through multimodal fusion saliency prediction technology, feature extraction and matching are guided, key frames are selected, and pose optimization is performed to improve the focus on effective structured areas.
It achieves accurate perception and efficient matching of effective structured areas, improves the stability and performance of the system, and improves the efficiency of key frame usage and the accuracy of pose estimation.
Smart Images

Figure CN120599233A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of simultaneous positioning and mapping, and particularly relates to a visual SLAM method, system and storage medium based on saliency prediction. Background Art
[0002] Simultaneous localization and mapping (SLAM) technology is at the core of autonomous navigation and environmental perception, enabling robots to achieve real-time localization and mapping in unknown environments. As a key branch of SLAM, visual SLAM utilizes camera images for localization and mapping, offering advantages such as low cost and high performance. However, traditional visual SLAM methods still have shortcomings.
[0003] Current mainstream visual SLAM methods often ignore the critical role of effective structured regions, such as edges and corners, in feature extraction, matching, and pose estimation, and instead treat all image regions equally. This not only fails to fully exploit the rich information contained in effective structured regions, but is also susceptible to the negative effects of textureless and repetitive textured regions, leading to wasted computational resources and hindering overall system performance. Therefore, there is an urgent need for a visual SLAM method that can consistently focus on effective structured regions in the environment to improve system performance and efficiency. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a visual SLAM method and system based on saliency prediction. By introducing saliency prediction technology using multimodal fusion to guide feature extraction and matching, keyframe selection and pose optimization processes, the focus on effective structured areas in the environment is improved, effectively addressing the problems of insufficient performance and low efficiency of traditional methods.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: In one aspect, the present invention provides a visual SLAM method based on saliency prediction, comprising the following steps: Use the depth camera on the mobile platform to obtain image frames and depth frames of the current surrounding environment; Perform saliency prediction based on multimodal fusion according to the current image frame and depth frame to obtain the grayscale image of the current image frame and the saliency mask containing the effective structured area; Combined with the saliency mask, the saliency mask is used to perform feature extraction and matching on the current image frame to obtain matching feature points with saliency values; Calculate the saliency entropy of the current image frame based on the matching feature points with saliency values, and determine whether the current image frame is the current key frame based on the saliency entropy and save it to the key frame set KeyFrames; Use the current keyframe to create the current map point and perform real-time grading to obtain a graded local map; Perform global BA weighted optimization on the camera pose based on the current map point level and the saliency mean of the hierarchical local map; The global map is constructed by continuously processing new image frames and depth frames captured by the depth camera in real time, augmenting and maintaining the hierarchical local map.
[0006] As a preferred technical solution, a saliency mask containing a valid structured area of the current image frame is obtained, specifically: Grayscale processing is performed on the current image frame to obtain a corresponding grayscale image; the grayscale processing adopts a weighted average method; The grayscale image is processed using a Sobel operator to obtain an edge map of the current image frame; the Sobel operator performs a two-dimensional convolution operation on the grayscale image in the horizontal gradient direction and the vertical gradient direction, and calculates the gradient amplitude of the pixel point using the Pythagorean theorem to obtain the edge map; The edge map is smoothed using a Gaussian filter, and then the depth frame is used to perform depth adjustment on the smoothed edge map to obtain a saliency mask containing valid structured areas in the current image frame.
[0007] As a preferred technical solution, the obtaining of matching feature points with significant values is specifically as follows: Construct the Gaussian image pyramids IPyramid and Mpyramid of the current image frame grayscale image and saliency mask respectively; The IPyramid and MPyramid layers are rasterized and divided. The side length of each grid after rasterization is calculated based on the grid parameters and the global gradient mean of the current image frame. Traverse each grid in each layer of IPyramid and calculate the ORB feature point extraction threshold of each raster image in each layer of IPyramid based on the raster image and the corresponding raster image in MPyramid; According to the ORB feature point extraction threshold, the ORB feature point extraction algorithm is used to obtain the feature points and corresponding descriptors in the raster images of each layer of IPyramid. The extracted feature points and corresponding descriptors are saved in the feature point set KeyPoints and the descriptor set KeyPointsDes respectively with the layer number in IPyramid as the label; Perform quadtree homogenization on the feature points in the feature point set KeyPionts, remove some unevenly distributed feature points, and save the homogenized feature points and corresponding descriptors back to the feature point set KeyPionts and descriptor set KeyPointsDes with the layer number in IPyramid as the label as the feature points of the current image frame; Traverse the feature points in the feature point set KeyPoints, obtain the grayscale value of the corresponding position in MPyramid according to the coordinates of the feature point as the saliency value of the feature point in the current image frame, and save it to the saliency value set KeyPointsSal in the same storage structure; Perform feature point matching on the current image frame and the previous image frame to obtain a matching feature point set KeyPointsMatch.
[0008] As a preferred technical solution, the calculation method of the ORB feature point extraction threshold is: Based on the number of pixels and pixel gradient values of each raster image in each layer of IPyramid and the global gradient mean of the current image frame, the global gradient variance is calculated to describe the prominence of each raster image in each layer of IPyramid in the overall image frame; Based on the number of pixels, pixel gradient value and gradient mean of each raster image in each layer of IPyramid, the local gradient variance is calculated to describe the structural richness of each raster image in each layer of IPyramid. Calculate the grayscale mean of each raster image in each layer of IPyramid and the corresponding position raster image in each layer of MPyramid, and quantify the significance value of the raster image in the corresponding position in each layer of MPyramid by normalizing the grayscale mean; According to the ratio of the global gradient variance to the local gradient variance of each raster image in each layer of IPyramid, combined with the saliency value of the raster image at the corresponding position in each layer of MPyramid, the ORB feature point extraction threshold of each raster image in each layer of IPyramid is calculated; The feature point matching is specifically as follows: Calculate the significance weight of each feature point in the current image frame based on the significance value of each feature point in the significance value set KeyPointsSal; For each feature point in the current image frame, traverse its descriptors with all feature points in the previous image frame, calculate and record the significance weighted Hamming distance, select the descriptor with the smallest distance as the best matching candidate, and if the minimum distance is less than the preset distance threshold, the feature point corresponding to the descriptor is determined to be the matching feature point, and the indexes of the matched feature points of the current image frame and the previous image frame are saved in the matching feature point set KeyPointsMatch.
[0009] As a preferred technical solution, the method of judging whether the current image frame is the current key frame according to the saliency entropy is specifically as follows: Calculate the intra-frame entropy based on the number of matching feature points in the current image frame and the probability distribution of the saliency values of the matching feature points; Calculating the inter-frame entropy between the current image frame and adjacent image frames according to the intra-frame entropy of the image frames within the time window; Combining the intra-frame entropy and inter-frame entropy of the current image frame to obtain the saliency entropy of the current image frame; The saliency entropy is judged. If it is greater than the preset entropy threshold, the current image frame is the current key frame, and its frame index is saved to the key frame set KeyFrames; otherwise, the current image frame is not the current key frame.
[0010] As a preferred technical solution, the current map point is created using the current key frame and graded in real time, specifically: Triangulate the matching feature points of the current keyframe, calculate the 3D map points corresponding to the 2D feature points in the current keyframe, and add the current map points to the hierarchical local map; Assign the saliency value of the current keyframe matching feature point to the current map point as the saliency value of the map point; The level of the current map point is divided according to the comparison between the saliency value of the current map point and the saliency mean value of all map points in the graded local map.
[0011] As a preferred technical solution, the camera pose is globally optimized by weighted BA based on the current map point level and the saliency mean of the hierarchical local map, specifically: Calculate the location optimization weight corresponding to the current map point based on the saliency value and level of the current map point and the saliency mean of the graded local map; The information matrix and reprojection error in the global BA optimization process are weighted by optimizing the weighted weight corresponding to the current map point.
[0012] As a preferred technical solution, the information matrix and reprojection error in the global BA optimization process are weighted, specifically: The objective function of global BA optimization is constructed using the position of 3D map points, the rotation matrix and translation vector of the camera as optimization variables; The reprojection error and information matrix in the objective function are weighted using the position optimization weight corresponding to the current map point.
[0013] On the other hand, the present invention provides a visual SLAM system based on saliency prediction, which is applied to the visual SLAM method based on saliency prediction, including a data acquisition unit, a saliency prediction unit, a feature extraction and matching unit, a key frame determination unit, a map point creation and classification unit, a global BA weighted optimization unit and a global map construction unit; The data acquisition unit is used to obtain image frames and depth frames of the current surrounding environment using a depth camera mounted on the mobile platform; The saliency prediction unit is used to perform saliency prediction based on multimodal fusion according to the current image frame and the depth frame, and obtain a grayscale image of the current image frame and a saliency mask containing a valid structured area; The feature extraction and matching unit is used to perform saliency mask-driven feature extraction and matching on the current image frame in combination with the saliency mask to obtain matching feature points with saliency values; The key frame determination unit is used to calculate the significance entropy of the current image frame based on the matching feature points with significance values, and determine whether the current image frame is the current key frame based on the significance entropy and save it to the key frame set KeyFrames; The map point creation and grading unit is used to create the current map point using the current key frame and perform real-time grading to obtain a graded local map; The global BA weighted optimization unit is used to perform global BA weighted optimization on the camera pose according to the current map point level and the saliency mean of the hierarchical local map; The global map construction unit is used to expand and maintain the hierarchical local map by continuously processing new image frames and depth frames captured in real time by the depth camera, thereby constructing a global map.
[0014] In another aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the visual SLAM method based on saliency prediction.
[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. This paper introduces a saliency prediction technology based on multimodal fusion, which can accurately perceive the effective structured areas in the image and provide guidance for subsequent visual processing tasks. By fusing geometric information and depth information, the spatial distribution characteristics of the effective structured areas in the image are characterized from multiple levels.
[0016] 2. The present invention utilizes a saliency mask-driven feature extraction and matching method to achieve efficient capture and precise matching of salient features in effective structured areas, significantly enhancing the system's perception and association capabilities. At the same time, it effectively alleviates the low matching efficiency of traditional methods in areas with sparse or repetitive textures, and comprehensively improves the stability and overall performance of the system.
[0017] 3. The present invention adopts a key frame selection method based on saliency entropy, which effectively realizes the priority retention of high-saliency information frames and avoids the frequent introduction of redundant frames; by fully expressing key visual information, the efficiency of key frame utilization is significantly improved.
[0018] 4. The present invention implements a pose optimization method based on map point classification, which greatly enhances the contribution of map points in effective structured areas in the pose optimization process; by prioritizing the optimization of high-quality map points, the accuracy of pose estimation is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 This is an overall flow chart of a visual SLAM method based on saliency prediction in an embodiment of the present invention.
[0021] Figure 2 1 is an overall block diagram of a visual SLAM system based on saliency prediction in an embodiment of the present invention.
[0022] Figure 3 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0024] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0025] like Figure 1 As shown, this embodiment provides a visual SLAM method based on saliency prediction, comprising the following steps: S1. Use the depth camera mounted on the mobile platform to obtain image frames and depth frames of the current surrounding environment.
[0026] In this embodiment, the mobile platform is a robot equipped with a depth camera to capture image data and distance data of the current surrounding environment; however, the mobile platform is not limited to this, and other mobile platforms that need to achieve autonomous positioning and environmental perception also fall within the scope of protection of this application.
[0027] S2. Perform saliency prediction based on multimodal fusion according to the current image frame and depth frame to obtain the grayscale image of the current image frame and the saliency mask containing the effective structured area.
[0028] Furthermore, in step S2, a saliency prediction technique based on multimodal fusion is first performed on the current image frame and the depth frame to obtain a saliency mask that can describe the effective structured area of the image frame, specifically: S201, grayscale the current image frame to obtain the corresponding grayscale image I gray Grayscale processing uses a weighted average method, and its calculation formula is: Gray = 0.299·R + 0.587·G + 0.114·B, where Gray represents the pixel value of the image frame after grayscale processing, and R, G, and B represent the pixel values of the red, green, and blue channels of the image frame respectively; S202, use the Sobel operator to process the grayscale image to obtain the edge map I of the image frame edge ; The Sobel operator performs two-dimensional convolution operations on the grayscale image in the horizontal gradient direction and the vertical gradient direction, and uses the Pythagorean theorem to calculate the gradient amplitude of the pixel point to obtain the edge map.
[0029] In this embodiment, the horizontal direction G of the Sobel operator x and vertical direction G y The convolution kernel is defined as follows: , Then the gradient of the image frame in the horizontal and vertical directions is represented by the grayscale image I gray Respectively with G x and G y Convolution gives: S x (x,y) = G x *I gray (x,y), S y (x,y) = G y *I gray (x,y), Among them, S x (x,y) is the horizontal gradient, S y (x,y) is the vertical gradient, * represents the two-dimensional convolution operation; The Pythagorean theorem is used to calculate the gradient amplitude at the pixel point, and then the edge map I of the current image frame is obtained. edge: .
[0030] S203, edge graph I edge Use Gaussian filter for smoothing to enhance spatial connectivity; the Gaussian filter formula is: , Where (x, y) is the coordinate of the pixel in the edge map, G(·) represents the Gaussian function, and σ represents the standard deviation of the Gaussian filter.
[0031] S204. Finally, the depth frame is used to perform depth adjustment on the smoothed edge map to obtain the final saliency mask I containing the effective structured area. mask In this embodiment, when adjusting the depth, the grayscale value of the near edge is increased accordingly, and the grayscale value of the far edge is decreased, so that the final saliency mask containing the effective structured area can give priority to the near clear area; the depth adjustment formula is: , Among them, I mask (x,y) represents the pixel value at position (x,y) in the saliency mask, I gaussian (x, y) represents the pixel value at the (x, y) position in the edge map after smoothing. depth (x, y) is the pixel value at the (x, y) position in the depth frame. S3. Combined with the saliency mask, perform saliency mask-driven feature extraction and matching on the current image frame to obtain matching feature points with saliency values.
[0032] Furthermore, in step S3, the saliency mask is combined with the saliency mask to perform saliency mask-driven feature extraction and matching on the current image frame to obtain matching feature points with saliency values, specifically: S301 , constructing Gaussian image pyramids IPyramid and MPyramid of the grayscale image of the current image frame and the saliency mask respectively, so as to extract stable and scale-invariant feature points at different scales.
[0033] S302, rasterize each layer of IPyramid and MPyramid; the side length W of each grid after rasterization is calculated based on the grid parameters and the global gradient mean of the current image frame, and the calculation formula is: , Where μ is the grid parameter, is the global gradient mean of the current image frame.
[0034] S303 , then traverse each grid of each IPyramid layer, and calculate the threshold for ORB feature point extraction of each grid image in each IPyramid layer according to the grid image and the corresponding grid image in each MPyramid layer.
[0035] Furthermore, the calculation process of ORB feature point extraction threshold is: First, calculate the global gradient variance Var global To describe the prominence of the raster image in the overall image frame, the formula is as follows: , Where N is the total number of pixels in the raster image, G(x i , y i ) is the gradient value of the i-th pixel in the raster image, is the global gradient mean of the current image frame; Then calculate the local gradient variance Var local To describe the structural richness within the raster image, the formula is as follows: , Where N is the total number of pixels in the raster image, is the mean gradient of the raster image; Then, the grayscale mean of the corresponding position of each raster image in each layer of IPyramid is calculated, and the significance value of the raster image in the corresponding position in each layer of MPyramid is quantified by the normalized grayscale mean. i , calculated as: , Among them, σ i is the grayscale mean of the raster image at the corresponding position in each layer of MPyramid for the i-th raster image in each layer of IPyramid, is the grayscale mean of all raster images in each layer of MPyramid; Finally, the ORB feature point extraction threshold of each raster image in each layer of IPyramid is calculated using the formula: , Among them, φ(i) is the ORB feature point extraction threshold of the i-th raster image in a certain layer of IPyramid, λ is the threshold parameter, s i is the normalized significance value of the i-th raster image in a certain layer of IPyramid in the corresponding raster image in the corresponding layer of MPyramid, where the global gradient variance of the raster image i in a certain layer of IPyramid is calculated. and local gradient variance The ratio of is used to characterize the structural characteristics of the raster image.
[0036] S304. According to the ORB feature point extraction threshold φ, the ORB feature point extraction algorithm is used to obtain the feature points and corresponding descriptors in the raster images of each layer of IPyramid. The extracted feature points and corresponding descriptors are saved in the feature point set KeyPoints and the descriptor set KeyPointsDes respectively with the layer number in IPyramid as the label.
[0037] S305. Perform quadtree homogenization on the feature points in the feature point set KeyPionts, remove some unevenly distributed feature points, and save the homogenized feature points and corresponding descriptors back to KeyPionts and KeyPointsDes with the layer number in IPyramid as the label as the feature points of the current image frame.
[0038] S306 , traverse the feature points in the feature point set KeyPoints, obtain the grayscale value of the corresponding position in MPyramid according to the coordinates of the feature point as the saliency value of the feature point in the current image frame, and save it to the saliency value set KeyPointsSal in the same storage structure.
[0039] S307 : Perform feature point matching on the current image frame and the previous image frame to obtain a matching feature point set KeyPointsMatch.
[0040] Furthermore, feature point matching is specifically as follows: First, calculate the saliency weight of each feature point in the current image frame: , Among them, ε ij represents the saliency weight of the jth feature point in the i-th layer image in the current image frame IPyramid, KeyPointsSal[i][j] represents the saliency value corresponding to the jth feature point in the i-th layer image in the current image frame IPyramid, and b is the offset.
[0041] Next, for each feature point in the current image frame, traverse its descriptors with all feature points in the previous image frame, calculate and record the significance weighted Hamming distance between it and all feature points in the previous image frame, select the descriptor with the smallest distance as the best matching candidate, and if the minimum distance is less than the preset distance threshold, the feature point corresponding to the descriptor is determined to be the matching feature point, and the indexes of the matched feature points of the current image frame and the feature points of the previous image frame are saved in the matching feature point set KeyPointsMatch.
[0042] In this embodiment, for each feature point KeyPointsSal[i][j] in the current image frame, the saliency weighted Hamming distance calculation formula between it and the feature point in the previous image frame is: , in, is the significance-weighted Hamming distance between the jth feature point of the i-th layer image in the current image frame IPyramid and the yth feature point of the x-th layer image in the previous image frame IPyramid; It is an indicator function that returns 1 if the condition is met, otherwise it returns 0. Indicates the kth bit of the descriptor of the jth feature point of the i-th layer image in the current image frame IPyramid, Represents the kth bit of the descriptor of the yth feature point in the xth layer image of the previous image frame IPyramid.
[0043] S4. Calculate the saliency entropy of the current image frame based on the matching feature points with saliency values, and determine whether the current image frame is a current key frame based on the saliency entropy and save it to the key frame set KeyFrames.
[0044] Furthermore, the step S4 of determining whether the current image frame is the current key frame is as follows: S401, calculating intra-frame entropy E in To evaluate the distribution of the saliency values of the matching feature points in the current image frame, the calculation formula is: , Where N is the total number of matching feature points in the current image frame, Represents the saliency value s of the matching feature point i i The probability distribution of E in When E is larger, the distribution of the saliency values of the matching feature points is more dispersed, and the image frame contains more different effective structured areas; on the contrary, if E in A smaller value indicates that the saliency values of most feature points in the image frame are close, and there are fewer effective structured areas in the image frame.
[0045] S402, calculating inter-frame entropy E out To measure the change in the saliency value of feature points between adjacent image frames within the time window, the calculation formula is: , Where M is the number of frames in the time window, t is the frame index in the time window, and the value range is [kM, k], including the current image frame k, represents the intra-frame entropy of image frame t Probability distribution of inter-frame entropy E outThe larger the value, the greater the difference in the effective structured regions of each frame in the time window, and the richer the scene information. out The smaller it is, the more similar the effective structured areas of each frame are, and the higher the repetition of scene information.
[0046] S403: Combine the intra-frame entropy E of the current image frame in and inter-frame entropy E out , to calculate the saliency entropy ξ of the current image frame: ξ = E in + E out If the saliency entropy ξ of the current image frame is greater than the preset entropy threshold, the current image frame is determined to be the current key frame, and its frame index is saved to the key frame set KeyFrames; otherwise, it is not the current key frame.
[0047] S5. Use the current key frame to create the current map point and perform real-time grading to obtain a graded local map.
[0048] Furthermore, the method of obtaining the hierarchical local map in step S5 is as follows: S501 : triangulate the matching feature points of the current key frame, calculate and obtain the 3D map points corresponding to the 2D feature points in the current key frame, and add the current map points to the hierarchical local map.
[0049] S502: Assign the saliency value of the current key frame matching feature point to the current map point as the saliency value of the current map point.
[0050] S503: Perform grading based on the saliency value of the current map point and the average saliency value of all map points in the graded local map.
[0051] The specific mechanism of level division in this embodiment is as follows: , Among them, Level(m p ) is the current map point m p The level, S(m p ) represents the current map point m p The significance value of Represents the mean significance of all map points in the hierarchical local map; this division mechanism divides the current map point into three levels, where the third-level map point is the most important map point and the first-level map point is the least important map point, based on which a hierarchical local map is obtained.
[0052] S6. Perform global weighted BA optimization on the camera pose based on the current map point level and the saliency mean of the hierarchical local map.
[0053] Furthermore, in step S6, the camera pose is globally weighted BA optimization based on the map point level and the saliency mean of the graded local map, specifically: S601, according to the level of the current map point Level (m p ), significance value S(m p ) and the significance mean of the graded local map , calculate the current map point m p The corresponding position optimization weighted weight ω: , Among them, λ Level Represents the weight coefficient related to the current map point level, which is used to balance the impact of map points of different levels on the optimization results.
[0054] S602, using the current map point m p The corresponding position optimization weight ω is used to weight the information matrix and reprojection error in the global BA optimization process.
[0055] Furthermore, the global BA optimization process is as follows: First, the position of the 3D map point, the camera's rotation matrix and translation vector are used as optimization variables to construct the objective function of global BA optimization: , Among them, {X i ,R l ,t l} is the optimization variable, including the position X of the 3D map point i and the camera's rotation matrix R l and translation vector t l , P L Refers to the three-dimensional coordinates of all map points that can be observed in the local map keyframe, K L Indicates the pose of the local map keyframe that has a common view relationship with the current keyframe, K F Indicates the pose of the local map keyframe that has no co-viewing relationship with the current keyframe, X k Defined as P L The set of matches between the map points in and the feature points in key frame k, ρ() is the robust kernel function, E k,j Represents the reprojection error between the estimated pose and the true value.
[0056] Then use the position corresponding to the current map point to optimize the weighted weight ω for the reprojection error E in the objective function k,j and information matrix The weighting is as follows: , , Among them, ω jOptimize the weighted weight for the current map point j, is the observation position of the current map point j, π (·) is the projection function of the pinhole camera, It is an information matrix used to describe the uncertainty between the observed value and the true value; the reprojection error and information matrix of the map points are reasonably controlled through adaptive weights; for important map points, the reprojection error is appropriately amplified to increase the system's attention during the optimization process; at the same time, its information matrix is appropriately increased to improve the credibility of its observation value during the optimization process; for unimportant map points, the opposite strategy is adopted.
[0057] S7. By continuously processing new image frames and depth frames captured by the depth camera in real time, the hierarchical local map is expanded and maintained to build a global map.
[0058] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0059] Based on the same idea as the visual SLAM method based on saliency prediction in the above-mentioned embodiment, the present invention also provides a visual SLAM system based on saliency prediction, which can be used to execute the above-mentioned visual SLAM method based on saliency prediction. For ease of explanation, the structural diagram of an embodiment of a visual SLAM system based on saliency prediction only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0060] like Figure 2 As shown, another embodiment of the present invention provides a visual SLAM system based on saliency prediction, including a data acquisition unit, a saliency prediction unit, a feature extraction and matching unit, a key frame determination unit, a map point creation and classification unit, a global BA weighted optimization unit, and a global map construction unit; The data acquisition unit is used to obtain image frames and depth frames of the current surrounding environment using a depth camera mounted on the mobile platform; The saliency prediction unit is used to perform saliency prediction based on multimodal fusion according to the current image frame and the depth frame, and obtain the grayscale image of the current image frame and the saliency mask containing the effective structured area; The feature extraction and matching unit is used to perform saliency mask-driven feature extraction and matching on the current image frame in combination with the saliency mask to obtain matching feature points with saliency values; The key frame determination unit is used to calculate the saliency entropy of the current image frame based on the matching feature points with saliency values, and determine whether the current image frame is the current key frame based on the saliency entropy and save it to the key frame set KeyFrames; The map point creation and grading unit is used to create the current map point using the current key frame and perform real-time grading to obtain a graded local map; The global BA weighted optimization unit is used to perform global BA weighted optimization on the camera pose according to the current map point level and the saliency mean of the hierarchical local map; The global map construction unit is used to build a global map by continuously processing new image frames and depth frames captured by the depth camera in real time, expanding and maintaining the hierarchical local map.
[0061] It should be noted that a visual SLAM system based on saliency prediction of the present invention corresponds one-to-one to a visual SLAM method based on saliency prediction of the present invention. The technical features and beneficial effects described in the embodiment of the above-mentioned visual SLAM method based on saliency prediction are applicable to the embodiment of a visual SLAM system based on saliency prediction. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.
[0062] In addition, in the implementation of a visual SLAM system based on saliency prediction in the above-mentioned embodiment, the logical division of each program module is only an example. In actual applications, the above-mentioned functions can be assigned to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the visual SLAM system based on saliency prediction is divided into different program modules to complete all or part of the functions described above.
[0063] like Figure 3 As shown, in one embodiment, a computer-readable storage medium is provided, which stores a program in a memory. When the program is executed by a processor, the visual SLAM method based on saliency prediction is implemented, specifically: Use the depth camera on the mobile platform to obtain image frames and depth frames of the current surrounding environment; Perform saliency prediction based on multimodal fusion according to the current image frame and depth frame to obtain the grayscale image of the current image frame and the saliency mask containing the effective structured area; Combined with the saliency mask, the saliency mask is used to perform feature extraction and matching on the current image frame to obtain matching feature points with saliency values; Calculate the saliency entropy of the current image frame based on the matching feature points with saliency values, and determine whether the current image frame is the current key frame based on the saliency entropy and save it to the key frame set KeyFrames; Use the current keyframe to create the current map point and perform real-time grading to obtain a graded local map; Perform global BA weighted optimization on the camera pose based on the current map point level and the saliency mean of the hierarchical local map; The global map is constructed by continuously processing new image frames and depth frames captured by the depth camera in real time, augmenting and maintaining the hierarchical local map.
[0064] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0065] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0066] The above embodiments are preferred implementations of the present invention, but the implementations of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A visual SLAM method based on saliency prediction, characterized in that: The steps include: Use the depth camera on the mobile platform to obtain image frames and depth frames of the current surrounding environment; Perform saliency prediction based on multimodal fusion according to the current image frame and depth frame to obtain the grayscale image of the current image frame and the saliency mask containing the effective structured area; Combined with the saliency mask, the saliency mask is used to perform feature extraction and matching on the current image frame to obtain matching feature points with saliency values; Calculate the saliency entropy of the current image frame based on the matching feature points with saliency values, and determine whether the current image frame is the current key frame based on the saliency entropy and save it to the key frame set KeyFrames; Use the current keyframe to create the current map point and perform real-time grading to obtain a graded local map; Perform global BA weighted optimization on the camera pose based on the current map point level and the saliency mean of the hierarchical local map; The global map is constructed by continuously processing new image frames and depth frames captured by the depth camera in real time, augmenting and maintaining the hierarchical local map.
2. The visual SLAM method based on saliency prediction according to claim 1, wherein Obtain the saliency mask containing the valid structured area of the current image frame, specifically: Grayscale processing is performed on the current image frame to obtain a corresponding grayscale image; the grayscale processing adopts a weighted average method; The grayscale image is processed using a Sobel operator to obtain an edge map of the current image frame; the Sobel operator performs a two-dimensional convolution operation on the grayscale image in the horizontal gradient direction and the vertical gradient direction, and calculates the gradient amplitude of the pixel point using the Pythagorean theorem to obtain the edge map; The edge map is smoothed using a Gaussian filter, and then the depth frame is used to perform depth adjustment on the smoothed edge map to obtain a saliency mask containing valid structured areas in the current image frame.
3. The visual SLAM method based on saliency prediction according to claim 1, wherein The obtaining of matching feature points with significant values is specifically as follows: Construct the Gaussian image pyramids IPyramid and Mpyramid of the current image frame grayscale image and saliency mask respectively; The IPyramid and MPyramid layers are rasterized and divided. The side length of each grid after rasterization is calculated based on the grid parameters and the global gradient mean of the current image frame. Traverse each grid in each layer of IPyramid and calculate the ORB feature point extraction threshold of each raster image in each layer of IPyramid based on the raster image and the corresponding raster image in MPyramid; According to the ORB feature point extraction threshold, the ORB feature point extraction algorithm is used to obtain the feature points and corresponding descriptors in the raster images of each layer of IPyramid. The extracted feature points and corresponding descriptors are saved in the feature point set KeyPoints and the descriptor set KeyPointsDes respectively with the layer number in IPyramid as the label; Perform quadtree homogenization on the feature points in the feature point set KeyPionts, remove some unevenly distributed feature points, and save the homogenized feature points and corresponding descriptors back to the feature point set KeyPionts and descriptor set KeyPointsDes with the layer number in IPyramid as the label as the feature points of the current image frame; Traverse the feature points in the feature point set KeyPoints, obtain the grayscale value of the corresponding position in MPyramid according to the coordinates of the feature point as the saliency value of the feature point in the current image frame, and save it to the saliency value set KeyPointsSal in the same storage structure; Perform feature point matching on the current image frame and the previous image frame to obtain a matching feature point set KeyPointsMatch.
4. The visual SLAM method based on saliency prediction according to claim 3, wherein The calculation method of the ORB feature point extraction threshold is: Based on the number of pixels and pixel gradient values of each raster image in each layer of IPyramid and the global gradient mean of the current image frame, the global gradient variance is calculated to describe the prominence of each raster image in each layer of IPyramid in the overall image frame; Based on the number of pixels, pixel gradient value and gradient mean of each raster image in each layer of IPyramid, the local gradient variance is calculated to describe the structural richness of each raster image in each layer of IPyramid. Calculate the grayscale mean of each raster image in each layer of IPyramid and the corresponding position raster image in each layer of MPyramid, and quantify the significance value of the raster image in the corresponding position in each layer of MPyramid by normalizing the grayscale mean; According to the ratio of the global gradient variance to the local gradient variance of each raster image in each layer of IPyramid, combined with the saliency value of the raster image at the corresponding position in each layer of MPyramid, the ORB feature point extraction threshold of each raster image in each layer of IPyramid is calculated; The feature point matching is specifically as follows: Calculate the significance weight of each feature point in the current image frame based on the significance value of each feature point in the significance value set KeyPointsSal; For each feature point in the current image frame, traverse its descriptors with all feature points in the previous image frame, calculate and record the significance weighted Hamming distance, select the descriptor with the smallest distance as the best matching candidate, and if the minimum distance is less than the preset distance threshold, the feature point corresponding to the descriptor is determined to be the matching feature point, and the indexes of the matched feature points of the current image frame and the previous image frame are saved in the matching feature point set KeyPointsMatch.
5. The visual SLAM method based on saliency prediction according to claim 1, wherein The method of judging whether the current image frame is the current key frame according to the saliency entropy is specifically as follows: Calculate the intra-frame entropy based on the number of matching feature points in the current image frame and the probability distribution of the saliency values of the matching feature points; Calculating the inter-frame entropy between the current image frame and adjacent image frames according to the intra-frame entropy of the image frames within the time window; Combining the intra-frame entropy and inter-frame entropy of the current image frame to obtain the saliency entropy of the current image frame; The saliency entropy is judged. If it is greater than the preset entropy threshold, the current image frame is the current key frame, and its frame index is saved to the key frame set KeyFrames; otherwise, the current image frame is not the current key frame.
6. The visual SLAM method based on saliency prediction according to claim 1, wherein The method of using the current key frame to create the current map point and perform real-time grading is as follows: Triangulate the matching feature points of the current keyframe, calculate the 3D map points corresponding to the 2D feature points in the current keyframe, and add the current map points to the hierarchical local map; Assign the saliency value of the current keyframe matching feature point to the current map point as the saliency value of the map point; The level of the current map point is divided according to the comparison between the saliency value of the current map point and the saliency mean value of all map points in the graded local map.
7. The visual SLAM method based on saliency prediction according to claim 1, wherein The camera pose is optimized globally by weighted BA based on the current map point level and the saliency mean of the hierarchical local map, specifically: Calculate the location optimization weight corresponding to the current map point based on the saliency value and level of the current map point and the saliency mean of the graded local map; The information matrix and reprojection error in the global BA optimization process are weighted by optimizing the weighted weight corresponding to the current map point.
8. The visual SLAM method based on saliency prediction according to claim 7, wherein The information matrix and reprojection error in the global BA optimization process are weighted, specifically: The objective function of global BA optimization is constructed using the position of 3D map points, the rotation matrix and translation vector of the camera as optimization variables; The reprojection error and information matrix in the objective function are weighted using the position optimization weight corresponding to the current map point.
9. A visual SLAM system based on saliency prediction, characterized in that: The visual SLAM method based on saliency prediction as described in any one of claims 1 to 8 comprises a data acquisition unit, a saliency prediction unit, a feature extraction and matching unit, a key frame determination unit, a map point creation and classification unit, a global BA weighted optimization unit, and a global map construction unit; The data acquisition unit is used to obtain image frames and depth frames of the current surrounding environment using a depth camera mounted on the mobile platform; The saliency prediction unit is used to perform saliency prediction based on multimodal fusion according to the current image frame and the depth frame, and obtain a grayscale image of the current image frame and a saliency mask containing a valid structured area; The feature extraction and matching unit is used to perform saliency mask-driven feature extraction and matching on the current image frame in combination with the saliency mask to obtain matching feature points with saliency values; The key frame determination unit is used to calculate the significance entropy of the current image frame based on the matching feature points with significance values, and determine whether the current image frame is the current key frame based on the significance entropy and save it to the key frame set KeyFrames; The map point creation and grading unit is used to create the current map point using the current key frame and perform real-time grading to obtain a graded local map; The global BA weighted optimization unit is used to perform global BA weighted optimization on the camera pose according to the current map point level and the saliency mean of the hierarchical local map; The global map construction unit is used to expand and maintain the hierarchical local map by continuously processing new image frames and depth frames captured in real time by the depth camera, thereby constructing a global map.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the visual SLAM method based on saliency prediction according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multi-mode SLAM method, system and device, medium and program product
CN119131754A