Visual positioning and mapping method based on feature tracking and matching
By quantifying the image illumination and texture, combining MobileNetV4 to evaluate the credibility and using feature descriptors for closed-loop detection, optimizing the fusion of direct method and feature point method, the limited computing resources and error accumulation problems of visual/inertial navigation SLAM system in complex scenarios are solved, and high-precision positioning and mapping are achieved.
Patent Information
- Application Number
- CN202510766605.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing visual/inertial navigation SLAM system cannot fully utilize the fusion of environmental factors to guide the direct method and feature point method in complex scenarios, and the cumulative error correction effect is poor.
By quantifying the illumination and texture of the images collected by the camera, the credibility of the direct method and feature point method is evaluated using MobileNetV4 neural network, and closed-loop detection, optimized state estimation, and fusion of the results weight ratio of the two methods is realized through feature descriptors.
It improves the positioning accuracy and robustness of the system in complex environments, reduces cumulative errors, and improves the overall performance of the system.
Smart Images

Figure CN120279102A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of simultaneous localization and mapping, and particularly relates to a visual localization and mapping method based on feature tracking and matching. Background Art
[0002] The key to a mobile robot in autonomous navigation lies in accurate and robust state estimation. In traditional visual / inertial SLAM systems, the direct method and the feature point method each have their own characteristics but also have inherent limitations. How to effectively integrate the two methods, give full play to their complementary advantages in different environments, and improve the performance of the system in complex scenarios is an important research topic.
[0003] In terms of existing research, SVO realizes feature tracking and pose estimation by minimizing photometric error, but lacks a loop detection mechanism. LDSO introduces feature descriptors to achieve loop detection, but the feature point information is not used for state estimation. Other studies have tried to integrate the two methods through different fusion mechanisms, such as a switching mechanism based on the number of features, error weighted fusion, etc., but there is still room for improvement in terms of accuracy and computational efficiency. However, currently, the system mainly has two problems. One is how to make full use of environmental factors to guide the fusion of the direct method and the feature point method under limited computing resources, and the other is how to effectively correct the cumulative error. Therefore, a new technical solution is needed to solve these problems. Summary of the Invention
[0004] Object of the Invention: In order to overcome the deficiencies in the prior art, a visual localization and mapping method based on feature tracking and matching is provided.
[0005] Technical Solution: To achieve the above object, the present invention provides a visual localization and mapping method based on feature tracking and matching, including the following steps:
[0006] S1: Quantify the images collected by the camera in terms of both illumination and texture intensity to obtain the illumination quantization result and the texture quantization result;
[0007] S2: Obtain the direct method credibility and the feature point method credibility through MobileNetV4 using the obtained illumination quantization result and texture quantization result;
[0008] S3: Input the direct method credibility and the feature point method credibility into the fusion module to obtain the result weight ratios of the direct method and the feature point method;
[0009] S4: Input the result weight ratios of the direct method and the feature point method into the state estimation module for state estimation.
[0010] Further, the step S1 includes:
[0011] A1: obtain a data set through camera acquisition, and process the data set acquired by the camera to obtain a processed data set;
[0012] A2: Processing the data set processed in step A1 by using an illumination quantization algorithm to obtain an illumination quantization result;
[0013] A3: Processing the data set processed in step A1 by a texture quantization algorithm to obtain a texture quantization result.
[0014] Furthermore, the step A1 specifically includes: performing feature tracking and matching on the images in the data set using the feature point method and the direct method, obtaining the pixel position of the feature point by the feature point method, and obtaining the gradient of the tracking feature by the direct method; processing one frame of the image, manually adjusting the pixel brightness or dynamic blur, performing feature tracking and matching on the processed image again, recording the pixel position changes of the matching features by the feature point method and the gradient changes of the tracking features by the direct method, acquiring all the images in the data set with the camera, and repeating the above steps to form a data set.
[0015] Furthermore, the step A2 specifically includes:
[0016] A2-1: Establish the current frame With reference frame The corresponding relationship of feature points is obtained by the feature point method between Grayscale features of feature points Matching is performed, and the similarity between the descriptors of the feature points is calculated by the Hamming distance calculation method in the feature point method, and the most similar feature point is found as Corresponding feature points Grayscale features , They are the grayscale features of the 1st, 2nd, ..., nth feature points respectively;
[0017] A2-2: Initial illumination quantification results , Used to accumulate and calculate the pixel grayscale difference of grayscale features; traverse through the for loop Grayscale features of feature points and Each corresponding feature point Grayscale features , execute the formula:
[0018]
[0019] That is, calculate the absolute value of the pixel grayscale difference between the matching feature points in two adjacent frames of images and add it to the degree of illumination change. ,in ( ) represents the pixel gray value corresponding to the gray feature in the image ; ( ) represents the pixel gray value of the feature point in the image ;
[0020] A2-3: After all feature points in the image are processed, the loop ends; divide the accumulated by the total number of feature points , that is, execute , to obtain the illumination quantization result and output it.
[0021] Furthermore, the specific steps of step A3 include:
[0022] A3-1: For the image , set the pixel gray level k, step size , and direction of the image;
[0023] Process the image , and initialize the contrast parameter value at the same time; ;
[0024] A3-2: Traverse the direction through a for loop to calculate the gray-level co-occurrence matrix in each direction. Under the conditions of step size and direction , the probability that the starting pixel gray level is and the ending pixel gray level is is recorded as the corresponding occurrence frequency . When the gray values of the image are divided into levels, and the spatial relationship between the step size and the direction is determined; using as the row index and as the column index, each row of the matrix corresponds to a starting pixel gray level, and each column corresponds to an ending pixel gray level; use the probability value as the matrix element to construct a matrix:
[0025]
[0026] A3-3: Calculate the contrast parameter value in this direction based on the matrix; The calculation formula of is as follows:
[0027]
[0028] The contrast parameter value calculated in the current direction is accumulated into the contrast parameter accumulation value ; when all specified directions are traversed, the loop ends; the contrast parameter accumulation value is divided by the number of directions to obtain the final texture quantization result .
[0029] Furthermore, the step S2 specifically includes: using the MobileNetV4 neural network for confidence regression to form a confidence prediction model with a three-layer fully connected network; the dimension of the network input layer is 2, and the activation function is Relu; the number of hidden layer nodes is 256, and the activation function is Relu; the number of output layer nodes is 2, and there is no activation function; the network takes the light change L and the texture intensity T as inputs and outputs the direct method confidence and the feature point method confidence .
[0030] Furthermore, the step S3 specifically includes:
[0031] B1: Analyze the direct method confidence:
[0032] In normal adjacent frames and , there are features of feature point matching obtained according to the direct method and , change the brightness or blur degree of the frame to obtain the corresponding frame , use to match the features in to obtain new matching features , where and are not exactly the same due to changes in environmental factors, calculate and the gradient difference of the corresponding features in and find the arithmetic mean ;
[0033] B2: Analyze the confidence of the feature point method:
[0034] In normal adjacent frames and , there are matching features and , change the brightness or blur degree of the frame to obtain the corresponding frame , use to Match the features in it to obtain new matching features , where and are not exactly the same due to changes in environmental factors. Calculate and the pixel position difference of the corresponding features in it and calculate the arithmetic mean ;
[0035] B3: The weight ratio of the operation results of the direct method and the feature point method and are defined as follows:
[0036] .
[0037] Furthermore, the step S4 specifically includes:
[0038] C1: For the relative pose between adjacent frames of nodes, it is calculated by the following formula:
[0039]
[0040] where is the position of the th frame relative to the th frame, in the coordinate system of the th frame, and is obtained by transforming the positions of the two frames in the world coordinate system , and the rotation matrix of the th frame; and are respectively the pose-related quantities of the two frames, and their difference gives the relative pose ;
[0041] C2: After a loop occurs, the following constraints are established for the poses of the key frames located between the loop frames:
[0042]
[0043] where is the error term, including two parts: position error and pose error, and represent different frame numbers; is the inverse matrix of the rotation matrix of the th frame, used to transform the positions , in the world coordinate system to the coordinate system of the th frame and calculate the position error with the relative position ; The pose error is calculated from the pose difference between the loop frames.
[0044] Further, the objective of pose graph optimization in step C2 is as follows:
[0045]
[0046] Wherein, is the set composed of all consecutive frames, is the set composed of all loop-closed frames, represents the set of variables to be optimized, is the kernel function, is the error term The square of the two-norm. The above loop-closure detection and optimization further play the role of feature point descriptors, reduce the cumulative error of historical key frames, and fully improve the positioning accuracy of the system.
[0047] Beneficial effects: Compared with the prior art, the present invention proposes a novel adaptive fusion algorithm that fuses direct method and feature point method for visual / inertial SLAM. Different from the existing switching mechanism, the present invention fully considers the influence of weak light and weak texture environment on the algorithm performance, establishes the relationship between environmental factors and algorithm performance through MobileNetV4, and uses feature descriptors to achieve loop-closure detection, thereby improving the overall performance of the system, having the advantages of high trajectory accuracy and good robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the flow framework diagram of the method of the present invention;
[0049] Figure 2 is the production flow chart of the data set in the present invention;
[0050] Figure 3 is the simulated illumination change image in the production process of the data set of the present invention;
[0051] Figure 4 is the simulated motion blur image in the production process of the data set of the present invention;
[0052] Figure 5 is the simplified structure diagram of MobileNetV4 of the present invention;
[0053] Figure 6 is the schematic diagram of the situation where environmental factors change feature matching in the present invention;
[0054] Figure 7 is the schematic diagram of loop-closure constraint in the system loop of the present invention;
[0055] Figure 8 is the trajectory visualization result of three algorithms in some test sequences. DETAILED DESCRIPTION OF THE INVENTION
[0056] The present invention will be further illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.
[0057] Embodiment 1:
[0058] As Figure 1 shown, this embodiment provides a visual positioning and mapping method based on feature tracking and matching, including the following steps:
[0059] S1: Quantify the images collected by the camera in terms of both illumination and texture intensity to obtain an illumination quantization result and a texture quantization result;
[0060] Step S1 includes:
[0061] A1: Obtain a data set through camera acquisition, process the data set collected by the camera to obtain a processed data set;
[0062] Figure 2 is the data set production process provided by this embodiment. Feature tracking and matching are performed on the selected image pairs, and the pixel positions of the features matched by the feature point method and the gradients of the features tracked by the direct method are recorded.
[0063] Process one of the frames of the image, adjust the pixel brightness or apply dynamic blur, Figure 3 is a partial data set after brightness adjustment. Brightness adjustment is performed on the selected image pairs in the data set to simulate the light changes in the real scene. Figure 3 In (a) is the original data set image, Figure 3 In (b) is the image with 40% increased brightness, Figure 3 In (c) is the image with 40% decreased brightness. The weak texture scene in the actual scene is simulated by adding dynamic blur.
[0064] Figure 4 In (a) is the original data set image, Figure 4 In (b) is the image with 4 pixels of dynamic blur, Figure 4 In (c) is the image with 8 pixels of dynamic blur. Feature tracking and matching are performed again on the processed image pairs, and the pixel positions of the features matched by the feature point method and the gradients of the features tracked by the direct method are recorded to verify whether the initial screening conditions are exceeded. Using the normal frame as the middle frame, the same features in the three frames are connected, and the credibility of the direct method and the feature point method is calculated. Repeat the above steps to record all data and form a data set.
[0065] A2: Process the processed data set in step A1 through an illumination quantization algorithm to obtain an illumination quantization result;
[0066] Step A2 specifically includes:
[0067] A2-1: Establish the correspondence of feature points between the current frame and the reference frame by the feature point method, and match the gray-scale features of the feature points in . Calculate the similarity between the descriptors of the feature points through the Hamming distance calculation method in the feature point method, and find the most similar feature points as the corresponding feature points in , which are the gray-scale features of the 1st, 2nd,..., nth feature points respectively;
[0068] A2-2: The initial light quantization result , is used to accumulate the pixel gray-scale differences for calculating the gray-scale features; traverse the gray-scale features of the feature points in by a for loop and each corresponding feature point in , and execute the formula:
[0069]
[0070] That is, calculate the absolute value of the pixel gray-scale difference between the matching feature points in two adjacent frames of images and accumulate it to the degree of light change , where ( ) represents the pixel gray-scale value of the corresponding gray-scale feature in image , and ( ) represents the pixel gray-scale value of the feature point in image ;
[0071] A2-3: When the operations on all feature points in image are completed, the loop ends; divide the accumulated by the total number of feature points , that is, execute to obtain the light quantization result and output it.
[0072] A3: Process the data set processed in step A1 through a texture quantization algorithm to obtain a texture quantization result;
[0073] Step A3 specifically includes:
[0074] A3-1: For the image , in this embodiment, the gray level of the image pixels is set , step size , direction ;
[0075] The step size defines the spatial distance between pixels when calculating the gray-level co-occurrence matrix; the direction is the set of directions considered when calculating the gray-level co-occurrence matrix;
[0076] Process the image , change the gray value range of the pixels in the original image to the set new range; at the same time, initialize the contrast parameter value ; ;
[0077] A3-2: Traverse the directions through a for loop , calculate the gray-level co-occurrence matrix for each direction , under the conditions of the step size and the direction , the probability that the starting pixel gray level is and the ending pixel gray level is is recorded as the corresponding occurrence frequency . When the gray values of the image are divided into levels, and the step size and the spatial relationship of the direction are determined; with as the row index and as the column index, each row of the matrix corresponds to a starting pixel gray level (represented by , ranging from 1 to ), and each column corresponds to an ending pixel gray level (represented by , ranging from 1 to ); the probability value is used as the matrix element to construct a matrix:
[0078] ;
[0079] A3-3: Calculate the contrast parameter value in this direction based on the matrix ; The calculation formula of
[0080]
[0081] The contrast parameter value calculated in the current direction is accumulated to the contrast parameter accumulation value In; when all specified directions have been traversed, the loop ends; the cumulative value of the contrast parameter is divided by the number of directions to obtain the final texture quantization result .
[0082] S2: The obtained illumination quantization result and texture quantization result are used to obtain the direct method credibility and the feature point method credibility through MobileNetV4;
[0083] In this embodiment, the MobileNetV4 neural network is used for credibility regression. As Figure 5 shown, a credibility prediction model is constructed with a three-layer fully connected network, namely the input layer, the hidden layer, and the output layer. The dimension of the network input layer is 2, and the activation function is Relu; the number of nodes in the hidden layer is 256, and the activation function is Relu; the number of nodes in the output layer is 2, and there is no activation function. The network takes the illumination change L and the texture intensity T as inputs and outputs the direct method credibility and the feature point method credibility .
[0084] In this embodiment, the training set is denoted as , where represents the input data, and represents the training label. With the network input and the mean square error of the training label, a loss function is constructed:
[0085] ,
[0086] After obtaining the loss function, the network parameters are updated based on MobileNetV4 by performing gradient descent.
[0087] S3: Output the direct method credibility and the feature point method credibility and input them into the fusion module to obtain the result weight ratio of the direct method and the feature point method;
[0088] Because of the change of environmental factors, errors occur in feature tracking or matching of the direct method and the feature point method. For example, under normal circumstances, the feature matching result is as shown in Figure 6 (a) in, and when the environmental factors change, the matching result is as shown in Figure 6 (b) in. Because of the change of environmental factors, mis-matching of features appears in the image. Therefore, it is necessary to quantitatively evaluate the performance of the direct method and the feature point method in different environments.
[0089] Step S3 specifically includes:
[0090] B1: Analyze the credibility of the direct method:
[0091] In normal adjacent frames and respectively have features with feature point matches obtained by the direct method and changing the brightness or blur degree of the frame to obtain the corresponding frame using to match the features in to obtain new matching features where and are not exactly the same due to changes in environmental factors, calculate and the gradient difference of the corresponding features in and calculate the arithmetic mean ;
[0092] B2: Analyze the credibility of the feature point method:
[0093] In normal adjacent frames and respectively have matching features and changing the brightness or blur degree of the frame to obtain the corresponding frame using to match the features in to obtain new matching features where and are not exactly the same due to changes in environmental factors, calculate and the pixel position difference of the corresponding features in and calculate the arithmetic mean ;
[0094] B3: The credibility of the direct method and the credibility of the feature point method cannot be directly used as weights in the fusion algorithm. This is because these two values, as measures for evaluating the performance of the direct method and the feature point method, participating as weights in the system state estimation may cause the visual information to account for too large a proportion in the overall optimization, restricting the influence of other quantities. Both the direct method and the feature point method are ways of using visual information to provide visual constraints for the system. Therefore, in state estimation, their importance should be consistent with other residual terms. Therefore, the weight ratios of the results of the direct method and the feature point method and are defined as follows:
[0095] ;
[0096] S4: Input the weight ratios of the results obtained from the direct method and the feature point method into the state estimation module for state estimation.
[0097] In this embodiment, Figure 1 the pre-integration part in
[0098] can be implemented with reference to the following literature:
[0099]
C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. M. Montiel, and J.D. Tardós, “ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual–Inertial, and Multimap SLAM,” IEEE Transactions on Robotics, vol. 37, no. 6,pp. 1874–1890, Dec. 2021, doi: 10.1109 / TRO.2021.3075644.
[0100] In the odometer, the feature point method can use the calculated descriptors to achieve loop closure detection. Usually, it is judged whether a loop occurs by comparing the similarity between the current scene and the map. When a loop is detected, the system updates the positioning according to this result, thereby correcting the accumulated errors. In visual / inertial SLAM, the inertial measurement contains gravity information. After the system is initialized, the magnitude and direction of gravity are known to the system. When a loop occurs, the six-degree-of-freedom pose of the system is simplified to a four-degree-of-freedom correction, optimizing the position and yaw angle while keeping the roll angle and pitch angle unchanged. The establishment of a loop will introduce loop constraints. As
[0101] shown, when a loop occurs in the system, loop constraints are established. C1: In the odometer, the feature point method can use the calculated descriptors to achieve loop closure detection. Usually, it is judged whether a loop occurs by comparing the similarity between the current scene and the map. When a loop is detected, the system updates the positioning according to this result, thereby correcting the accumulated errors. For the relative pose between adjacent frames of nodes, it is calculated by the following formula:
[0102]
[0103] where, is the position of the th frame relative to the th frame in the coordinate system of the th frame, obtained by the positions of the two frames in the world coordinate system , and the rotation matrix of the nth frame is obtained through transformation; and are respectively the pose-related quantities of two frames, and the relative pose is obtained by taking the difference between the two. ;
[0104] C2: After loop closure occurs, the following constraints are established for the poses of key frames located between loop closure frames:
[0105]
[0106] where is the error term, which includes two parts: position error and attitude error. and represent different frame numbers; is the inverse matrix of the rotation matrix of the nth frame, which is used to transform the position , in the world coordinate system to the relative position in the coordinate system of the nth frame for calculating the position error; the attitude error is calculated from the attitude difference between loop closure frames.
[0107] The objective of pose graph optimization is:
[0108]
[0109] where is the set composed of all consecutive frames, is the set composed of all loop closed frames, represents the set of variables to be optimized; is the kernel function. Introducing the kernel function in the optimization is mainly to handle outliers, downweight data points with large errors, and avoid these outliers having too much negative impact on the optimization result; when the error is small, the kernel function acts approximately like an identity function; when the error is large, the kernel function reduces the weight of the error, making the optimization process more robust. is the square of the two-norm of the error term , which is used to quantify the magnitude of the error. In the optimization objective, this value is minimized to reduce the pose error. The above loop closure detection and optimization further play the role of feature point descriptors, reduce the cumulative error of historical key frames, and fully improve the positioning accuracy of the system.
[0110] Example 2:
[0111] To verify the effectiveness and performance of the present invention, the algorithm provided by the present invention was run on the publicly available EuRoc dataset in this embodiment, and the good performance of the present invention in low-light and low-texture environments was verified by comparing the experimental results with those of other algorithms, as follows:
[0112] In the experimental part, ORB-SLAM3 and VI-DSO were selected as comparison algorithms. Both of these algorithms use sliding window optimization, but geometric error and photometric error optimization methods are applied respectively in state estimation. To ensure the objectivity and comparability of the experiment, the sliding window parameters of all algorithms were uniformly set to 10 frames.
[0113] The experimental evaluation was carried out based on 11 sequences of the EuRoc dataset. Each algorithm was tested 5 times repeatedly on each sequence, and the cumulative test time was more than 300 minutes, and the total test time of a single algorithm exceeded 100 minutes. To quantitatively evaluate the system performance, the root mean square error (RMSE) of trajectory translation was used as the evaluation criterion for positioning accuracy. The performances of three algorithms, namely ODF-SLAM (the algorithm of the present invention), ORB-SLAM3 and VI-DSO, on each sequence were listed in detail, and Figure 8 The trajectory visualization results of the three algorithms in some test sequences were given. Among them, the dashed line represents the true trajectory, which is the accurate trajectory for reference; "VI-DSO" and "ORB-SLAM3" are comparison algorithms, and "ODF-SLAM" is the algorithm of the present invention. Figure 8 In (a), it is the test result in the dataset scenario named MH_01. It can be seen that the trajectory (red line) generated by the ODF-SLAM algorithm closely fits the true trajectory (dashed line). Compared with the two comparison algorithms, VI-DSO (green line) and ORB-SLAM3 (blue line), ODF-SLAM can more accurately follow the twists and turns of the true trajectory and shows better positioning accuracy in complex paths, indicating that it is more effective in extracting and utilizing environmental features and more accurate in pose estimation in this scenario. Figure 8 In (b), (c), and (d), they are the test results in the dataset scenarios named MH_03, V1_01, and V1_03 respectively. It can be seen that the test results also show a similar situation to Figure 8 in (a), indicating that in these test scenarios, the ODF-SLAM algorithm is superior to other comparison algorithms in the accuracy of trajectory estimation and has better performance.
Claims
1. A visual positioning and mapping method based on feature tracking and matching, characterized in that It includes the following steps: S1: Perform quantization calculations on the images collected by the camera in terms of both illumination and texture intensity to obtain the illumination quantization result and the texture quantization result; S2: Obtain the direct method credibility and the feature point method credibility through MobileNetV4 using the obtained illumination quantization result and texture quantization result; S3: Input the direct method credibility and the feature point method credibility into the fusion module to obtain the result weight ratios of the direct method and the feature point method; S4: Input the result weight ratios of the direct method and the feature point method into the state estimation module for state estimation.
2. The visual positioning and mapping method based on feature tracking and matching according to claim 1, wherein The step S1 includes: A1: Obtain a data set through camera collection, and process the data set collected by the camera to obtain a processed data set; A2: Process the processed data set in step A1 through an illumination quantization algorithm to obtain an illumination quantization result; A3: Process the processed data set in step A1 through a texture quantization algorithm to obtain a texture quantization result.
3. A visual localization and mapping method based on feature tracking and matching according to claim 2, characterized in that, The step A1 specifically includes: performing feature tracking and matching on the images in the data set using the feature point method and the direct method, obtaining the pixel positions of the feature points through the feature point method, and obtaining the gradients of the tracked features through the direct method; processing one of the frames of images, artificially adjusting the pixel brightness or dynamic blur, performing feature tracking and matching on the processed images again, recording the changes in the pixel positions of the features matched by the feature point method and the changes in the gradients of the features tracked by the direct method, collecting all the images in the data set by the camera, and repeating the above steps to form a data set.
4. A visual positioning and mapping method based on feature tracking and matching according to claim 2, characterized in that, The step A2 specifically includes: A2-1: Establish the current frame and the reference frame to obtain the corresponding relationship of feature points by the feature point method, and match the gray-scale features of the feature points in . Calculate the similarity between the descriptors of the feature points by the Hamming distance calculation method in the feature point method, and find the most similar feature points as the corresponding feature points in gray-scale features , which are the gray-scale features of the 1st, 2nd, …, nth feature points respectively; A2-2: Initial Light Quantization Result , The pixel gray difference used for cumulative calculation of gray features; traverse through the for loop the gray features of the feature points in and each corresponding feature point in the gray features of , execute the formula: ; That is, calculate the absolute value of the pixel gray - level difference of the matching feature points in two adjacent frames of images, and accumulate it to the degree of illumination change , where ( ) represents the pixel gray - level value corresponding to the gray - level feature in the image , and ( ) represents the pixel gray - level value of the feature point in the image ; A2-3: After operations on all feature points in the image are completed, the loop ends; divide the accumulated by the total number of feature points , that is, execute to obtain and output the illumination quantization result.
5. A visual localization and mapping method based on feature tracking and matching according to claim 2, characterized in that The step A3 specifically includes: A3-1: For an image , set the gray level k of the image pixels, the step size , and the direction ; Process the image and initialize the contrast parameter value simultaneously ; A3-2: Traverse directions through a for loop , calculate the gray-level co-occurrence matrix for each direction , at the step size and direction conditions, the probability that the starting pixel gray level is , and the ending pixel gray level is , and the corresponding occurrence frequency is recorded as , when the gray values of the image are divided into levels, and the step size and the direction spatial relationship is determined; using as the row index and as the column index, each row of the matrix corresponds to a starting pixel gray level, and each column corresponds to an ending pixel gray level; taking the probability value as the matrix element, construct a matrix: ; A3-3: Calculate the contrast parameter value in this direction based on the matrix ; The calculation formula is as follows: ; The contrast parameter value calculated for the current direction is accumulated into the contrast parameter accumulation value ; when all specified directions have been traversed, the loop ends; divide the contrast parameter accumulation value by the number of directions to obtain the final texture quantization result .
6. A visual positioning and mapping method based on feature tracking and matching according to claim 1, characterized in that The specific steps of step S2 include: performing confidence regression using the MobileNetV4 neural network to construct a confidence prediction model with a three-layer fully connected network; the dimension of the network input layer is 2, and the activation function is Relu; the number of nodes in the hidden layer is 256, and the activation function is Relu; the number of nodes in the output layer is 2, and there is no activation function; the network takes the illumination change L and the texture intensity T as inputs and outputs the confidence of the direct method and the confidence of the feature point method .
7. A visual positioning and mapping method based on feature tracking and matching according to claim 1, characterized in that The step S3 specifically includes: B1: Analyze the direct method credibility: In normal adjacent frames and respectively have features of feature point matching obtained by the direct method and change the brightness or blur degree of the frame to obtain the corresponding frame use to match the features in to obtain new matching features wherein and are not completely the same due to changes in environmental factors, calculate the gradient difference of the corresponding features in and find the arithmetic mean B2: Analyze the credibility of the feature point method: In normal adjacent frames and there are matching features respectively and By changing the brightness or blurriness of the frame, the corresponding frame is obtained. Using to match the features in new matching features are obtained, where and are not exactly the same due to changes in environmental factors. Calculate and the pixel position differences of the corresponding features in and calculate the arithmetic mean ; B3: Weight ratio of the results of the direct method and the feature point method and are defined as follows: 。 8. A visual localization and mapping method based on feature tracking and matching according to claim 1, characterized in that The step S4 specifically includes: C1: For the relative pose between adjacent frames of adjacent nodes, it is calculated through the following formula: ; Among them, is the position of the frame relative to the frame, which is obtained by transforming the positions of the two frames in the world coordinate system , and the rotation matrix of the frame; and are respectively the attitude-related quantities of the two frames, and the relative attitude is obtained by taking the difference between them; C2: After a loop occurs, establish the following constraints on the poses of the key frames located between the loop frames: ; Among them, is the error term, which includes two parts: position error and attitude error. and represent different frame numbers; is the inverse rotation matrix of the th frame, which is used to transform the position , in the world coordinate system to the relative position in the th frame coordinate system for calculating the position error; the attitude error is calculated from the attitude difference between loopback frames.
9. A visual positioning and mapping method based on feature tracking and matching according to claim 8, characterized in that The objective of the pose graph optimization in the step C2 is: ; Among them, is the set composed of all consecutive frames, is the set composed of all closed-loop frames, represents the set of variables to be optimized, is the kernel function, is the error term the square of the two-norm.
Citation Information
Patent Citations
A fast monocular vision odometer navigation and positioning method combining a feature point method and a direct method
CN109544636A
Self-adaptive visual inertial navigation odometer output method
CN117576218A
Monocular vision inertial odometer positioning method based on self-adaptive mixed vision residual error
CN119756351A