Visual SLAM method suitable for low-light dynamic environment

By using a hybrid attention mechanism in the visual SLAM algorithm to improve the self-calibrated lighting framework, adaptive ORB feature point extraction and instance segmentation network, the problem of insufficient positioning accuracy and robustness of the visual SLAM algorithm in low-light and dynamic environments is solved, and more efficient feature extraction and pose estimation are achieved.

CN120031971APending Publication Date: 2025-05-23FUZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510114399.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing visual SLAM algorithms are difficult to effectively resist the interference caused by insufficient lighting and dynamic objects in low light and dynamic environments, resulting in insufficient positioning accuracy and robustness.

Method used

The hybrid attention mechanism is used to improve the self-calibrated lighting framework to enhance the image of low-light images; the adaptive threshold is constructed through the standard deviation of image grayscale value, and the ORB feature point extraction method is improved; the example segmentation network YOLOv8-seg is used to segment the prior dynamic object mask, eliminate dynamic feature points, and only static feature points are used for inter-frame matching and pose estimation.

Benefits of technology

In low light and dynamic environments, the positioning accuracy and robustness of the visual SLAM algorithm are significantly improved, the visibility and adaptability of feature extraction are enhanced, and interference caused by dynamic objects is effectively handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031971A_ABST
    Figure CN120031971A_ABST
Patent Text Reader

Abstract

The invention provides a visual SLAM method suitable for a low-illumination dynamic environment, and the method comprises the following steps: S1, obtaining an RGB image of a scene, converting the RGB image into an HSV color space, carrying out the brightness detection of a V channel of the image, improving a self-calibration illumination frame through a mixed attention mechanism, and carrying out the image enhancement of a low-illumination image; s2, utilizing an improved ORB feature point extraction method to perform feature point extraction on the images with different illumination intensities by applying an adaptive threshold value; s3, dynamic feature points in the image are extracted and filtered in combination with an instance segmentation network YOLOv8-seg and a motion consistency check method; and S4, performing inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain an optimal matching feature point, and performing camera pose estimation based on an RANSAC algorithm and the optimal matching feature point to obtain a camera motion result. According to the method, the problem that the existing visual SLAM navigation positioning cannot effectively resist the interference caused by the low-illumination scene and the existence of the moving object in the low-illumination and dynamic environment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a visual SLAM method suitable for low-light dynamic environments. Background Art

[0002] Simultaneous localization and mapping (SLAM) is one of the key technologies for intelligent mobile cars to achieve autonomous motion. Due to the advantages of visual sensors such as easy integration, low cost, and low energy consumption, the method of pose estimation and map construction using visual sensors has become a mainstream technology and has been widely used. However, common problems faced by visual SLAM, such as lighting changes, insufficient lighting, object motion, and lack of texture, still have a significant impact on the accuracy and robustness of the algorithm.

[0003] In low-light or dark environments, images captured by cameras often appear too dark due to insufficient lighting, resulting in loss of details, and visual SLAM algorithms are unable to extract effective image information to complete positioning tasks. Currently, most visual SLAM algorithms for low-light environments rely on traditional image enhancement techniques, but these methods have obvious defects. For example, they may lose details, amplify noise, produce oversaturation after enhancement, and perform poorly under complex lighting conditions. In addition, deep learning-based methods often find it difficult to adaptively enhance images when faced with local lighting changes, resulting in color distortion and image overexposure, and are unable to provide reliable positioning results. Therefore, in low-light environments, the positioning accuracy and robustness of these algorithms often cannot meet the needs of practical applications.

[0004] In traditional visual SLAM algorithms, they are usually designed based on the assumption that objects in the environment are static or slowly moving. However, in the real world, the existence of dynamic objects such as vehicles and pedestrians is a common phenomenon. The existence of these dynamic objects poses a challenge to SLAM algorithms because they introduce data association errors, which in turn affects the positioning accuracy and robustness of the system. In order to improve the performance of visual SLAM algorithms in dynamic environments, it is necessary to detect moving objects and exclude or reduce the influence of feature points in dynamic areas to reduce their negative impact on the positioning accuracy of the system.

[0005] In view of this, the present invention proposes a visual SLAM method suitable for low-light dynamic environments. Summary of the invention

[0006] The purpose of the present invention is to propose a visual SLAM method suitable for low-light dynamic environments to solve the problem that the existing visual SLAM navigation and positioning in low-light and dynamic environments cannot effectively resist the interference caused by low-light scenes and the presence of moving objects.

[0007] To achieve the above object, the technical solution of the present invention is: a visual SLAM method suitable for low-light dynamic environment, comprising the following steps:

[0008] S1: Obtain an RGB image of the scene, convert the image from the RGB color space to the HSV color space, perform brightness detection on the V channel of the image to divide the image with normal brightness and the image with low light intensity, use the hybrid attention mechanism to improve the self-calibration illumination framework and perform image enhancement on the low light intensity image;

[0009] S2: using the improved ORB feature point extraction method, constructing an adaptive threshold through the standard deviation of the image gray value, and applying the adaptive threshold to extract feature points from the images with different light intensities obtained in step S1;

[0010] S3: Use the instance segmentation network YOLOv8-seg to segment the normal illumination image and the enhanced image, obtain the mask of the prior dynamic object in the environment, and remove the feature points that fall on the mask of the prior dynamic object, and filter out the non-prior dynamic feature points in the environment through the motion consistency principle;

[0011] S4: Perform inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain the best matching feature points, perform camera pose estimation based on the RANSAC algorithm and the best matching feature points to obtain the camera motion results.

[0012] Preferably, the RGB image of the scene is acquired using an RGB-D camera.

[0013] Preferably, the brightness detection is performed on the V channel of the image to divide the image into the image with normal brightness and the image with low light intensity, specifically:

[0014] Calculate the root mean square of the brightness value of the image V channel:

[0015]

[0016] Where V RMS is the RMS value of image brightness; V i,j is the brightness value of the pixel in the i-th row and j-th column in the image, 1≤i≤h, 1≤j≤b; h is the height of the original image; b is the width of the original image;

[0017] Preset the RMS brightness threshold V A , the brightness value RMS V RMS Below the threshold V A The image with the same brightness is classified as a low-light image, otherwise it is classified as an image with normal brightness;

[0018] For the image that is initially judged to be of normal brightness, a second brightness detection is performed after random cropping, and the second brightness detection result is used as the final brightness detection result.

[0019] Preferably, the self-calibration lighting framework is improved by using a hybrid attention mechanism, specifically:

[0020] In the cascaded multi-stage illumination optimization process, a hybrid attention mechanism is introduced to optimize the illumination component. The initial t=0 stage process is expressed as:

[0021]

[0022] The process from t=1 to T-1 is expressed as:

[0023]

[0024] In the formula, y is the initial image; u t and x t Represent the residual term and illumination of the t-th stage (t=0,1,...,T-1) respectively; H represents the illumination estimation network and is independent of the number of stages; A(·) represents the attention function; F(·) represents the illumination estimation function; att represents the weight map generated by the hybrid attention mechanism; G(·) represents the self-calibration function;

[0025] Among them, the self-calibration module in the self-calibration illumination framework is expressed as:

[0026]

[0027] Where s is the self-correction mapping; K φ is the introduced parameterized operator with learnable parameters φ; v t is the input for illumination estimation at stage t after calibration at each stage; z t is the clear image at stage t.

[0028] Preferably, the hybrid attention mechanism is specifically:

[0029] For the input feature map, global average pooling and maximum pooling are concatenated in the V channel dimension. The resulting comprehensive illumination feature map is refined and fused with brightness features through convolution operations, and then the illumination attention weight map is generated through the Sigmoid activation function.

[0030] Preferably, the improved ORB feature point extraction method is used to construct an adaptive threshold through the standard deviation of the image grayscale value, and the adaptive threshold is applied to the image with different illumination intensities to extract the feature points, specifically: the ORB feature points include FAST key points and BRIEF descriptors, the FAST key points are identified by detecting the grayscale changes in the image, and the BRIEF descriptors are used to generate binary features of the key points;

[0031] The method of identifying FAST key points by detecting grayscale changes in an image is specifically as follows: using the standard deviation of the image grayscale value to set two adaptive FAST thresholds θ and θ min , and the threshold θ>θ min ; Use threshold θ to extract feature points from the image. If the feature points cannot be extracted from the image using threshold θ, then use threshold θ min Re-extract feature points from the image.

[0032] Preferably, the two adaptive FAST thresholds θ and θ are set. min , specifically:

[0033] For a pixel point p in the image, the grayscale value standard deviation δ of the image is used to dynamically adjust the comparison threshold θ between the brightness values ​​of point p and the surrounding 16 pixels on a circle with a radius of 3:

[0034]

[0035] In the formula, η is a parameter that controls the threshold value of feature point extraction; δ is the standard deviation of the image grayscale value, expressed as:

[0036]

[0037] Where, μ is the mean gray value of the image; I(i,j) is the gray value of the image at position (i,j); i and j represent the position of the pixel 1≤i≤h, 1≤j≤b; h and b represent the height and width of the image, respectively;

[0038] Threshold θ min It is expressed as:

[0039]

[0040] Preferably, the method uses the instance segmentation network YOLOv8-seg to segment the normal illumination image and the enhanced image, obtains the mask of the prior dynamic object in the environment, removes the feature points falling on the mask of the prior dynamic object, and filters out the non-prior dynamic feature points in the environment by the motion consistency principle, which specifically includes the following steps:

[0041] The normal illumination image and the enhanced image are input into the instance segmentation network YOLOv8-seg, segmented based on the prior dynamic objects, and the masks of the prior dynamic objects in the image are obtained, and the ORB feature points falling on the masks of the prior dynamic objects are removed;

[0042] The optical flow pyramid is used to calculate the matching relationship between feature points in two adjacent frames. The basic matrix F is obtained through the RANSAC algorithm, the epipolar line I is calculated, and the distance D from the matching point to the epipolar line is calculated to determine whether the point is a dynamic feature point. Only static feature points in the image are retained.

[0043] Preferably, the optical flow pyramid is used to calculate the matching relationship between feature points in two adjacent frames, the basic matrix F is obtained by the RANSAC algorithm, the epipolar line I is calculated, and the distance D from the matching point to the epipolar line is calculated to determine whether the point is a dynamic feature point, and only the static feature points in the image are retained, specifically:

[0044] Let p 1 and p 2 is the coordinates of the matching points in the previous frame and the current frame, P 1 and P 2 Yes 1 and p 2 The homogeneous coordinate form of is:

[0045] p 1 =[u 1 ,v 1 ], p 2 =[u 2 ,v 2 ], P 1 =[u 1 ,v 1 ,1],P 2 =[u 2 ,v 2 ,1]

[0046] Where u and v represent pixel coordinate values, then the epipolar equation I 1 for:

[0047]

[0048] In the formula, X, Y, Z are line vectors, F is the basic matrix, and the distance from each matching point to the epipolar line is:

[0049]

[0050] If the distance D from the matching point to the epipolar line is greater than the preset threshold λ, the point is judged as a dynamic feature point and filtered out, and the remaining static feature points are retained.

[0051] Preferably, performing inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain the best matching feature points specifically includes:

[0052] Four pairs of non-collinear static feature points are randomly selected in the previous frame and the current frame to calculate the homography matrix H, which is recorded as model M. Among them, facing a pair of 2D points, the relationship between the image coordinates can be directly obtained through the homography matrix:

[0053]

[0054] In the formula, p 1 and p 2 are the coordinates of the matching points in the previous frame and the current frame; H is the homography matrix of the matching of two feature points; K is the camera intrinsic parameter matrix; P is the plane where the feature points are located; R, t represent the movement of the camera in the rotation and translation coordinate systems respectively;

[0055] Use model M to transform all ORB feature points in the RGB image frame, and calculate the projection error between each ORB feature point and its corresponding position in the model; if the error of an ORB feature point is less than a predetermined threshold, it is regarded as an inlier and added to the inlier set; check whether the inlier set of the current iteration contains more inliers than the previously recorded optimal inlier set, if so, update the optimal inlier set and record the current iteration number k; if the current iteration number reaches the preset maximum iteration number, stop the iteration, and the final model obtained is the optimal model, achieving the best inter-frame feature matching.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] 1) Under different lighting conditions, the features of the image may become less obvious, affecting the effectiveness of feature detection and descriptors, and further affecting the performance of the SLAM algorithm. In addition, the uneven brightness of the image will affect the image enhancement effect, causing overexposure, noise amplification, or loss of details. In order to improve this, the present invention combines the self-calibration lighting network framework with the hybrid attention mechanism to assign different weight parameters to different areas of the image, adjust the contrast and brightness of the input image, and enhance the visibility of the features.

[0058] 2) To address the problem of decreased ORB feature extraction performance under complex lighting conditions, the present invention utilizes the grayscale value standard deviation of the image to optimize the ORB feature extraction process, and dynamically adjusts the judgment threshold of the FAST key points in combination with the image grayscale value standard deviation to achieve adaptation to complex lighting conditions.

[0059] 3) When there are occlusions or dynamic objects in the environment, these factors will interfere with the accurate extraction and tracking of feature points, reducing the accuracy of positioning and map construction. In order to overcome this challenge, the present invention introduces an instance segmentation network, which can effectively segment the prior dynamic objects in the image. By eliminating the feature points on the prior dynamic objects and combining the motion consistency check method to eliminate the feature points on the non-prior dynamic objects in the image, ORB-SLAM3 can better handle the interference caused by occlusion and dynamic objects, and improve the robustness of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram of the brightness detection process of a specific embodiment of the present invention;

[0061] Figure 2 A structural diagram of a hybrid attention mechanism according to a specific embodiment of the present invention;

[0062] Figure 3 It is a diagram of an improved self-calibration lighting framework according to a specific embodiment of the present invention;

[0063] Figure 4 This is a rendering of the present invention in a dark dynamic environment. DETAILED DESCRIPTION

[0064] The following is combined with Figure 1-4 , the technical solution of the present invention is specifically described.

[0065] The present invention proposes a visual SLAM method suitable for low-light dynamic environments. The method firstly adopts an average brightness threshold to detect the brightness of an image, and combines a hybrid attention mechanism with a self-calibrated illumination (SCI) framework to enhance the low-light image; secondly, an ORB-SLAM3 feature point extraction strategy improved based on a FAST key point extraction threshold method is proposed to improve the number of feature point extractions in a complex lighting environment; at the same time, instance segmentation is performed on a normal lighting image and an enhanced image to extract a mask on a priori dynamic objects in the environment; finally, feature points falling on the mask of a priori dynamic objects are eliminated, and the remaining dynamic feature points in the environment are filtered out using the motion consistency principle, and only static feature points are used for feature matching and camera pose estimation; the method specifically comprises the following steps:

[0066] S1: Get the RGB image of the scene, convert the image from the RGB color space to the HSV color space, perform brightness detection on the V channel of the image to divide the image with normal brightness and the image with low light intensity, use the hybrid attention mechanism combined with the Retinex theory to improve the self-calibration illumination framework and perform image enhancement on the low light image, gradually correct the input of each stage to affect the output of each stage, and realize adaptive enhancement of the image;

[0067] S2: Using the improved ORB feature point extraction method, an adaptive threshold is constructed through the standard deviation of the image gray value, and the adaptive threshold is applied to the image with different light intensities obtained in step S1 to extract feature points; using the adaptive FAST threshold calculation method to optimize the extraction process of ORB feature points, the extraction threshold of FAST key points is adjusted through the standard deviation of the image gray value, so as to realize the adaptability of ORB feature points to different lighting conditions;

[0068] S3: Combine YOLOv8-seg and motion consistency check methods to extract and filter dynamic feature points in the image, and obtain all static feature points in the ORB feature points; specifically, use the instance segmentation network YOLOv8-seg to segment the normal illumination image and the enhanced image, obtain the mask of the prior dynamic object in the environment, and remove the feature points that fall on the mask of the prior dynamic object, and filter out the non-prior dynamic feature points in the environment through the motion consistency principle;

[0069] S4: Perform inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain the best matching feature points, perform camera pose estimation based on the RANSAC algorithm and the best matching feature points to obtain the camera motion results.

[0070] In this embodiment, the RGB image of the scene is acquired using an RGB-D camera.

[0071] In this embodiment, reference Figure 1 , the brightness detection is performed on the V channel of the image to divide the image with normal brightness and the image with low light intensity, specifically:

[0072] Calculate the root mean square V of the brightness value of the image V channel RMS , used to evaluate the overall distribution and variation of image brightness:

[0073]

[0074] Where V RMS is the RMS value of image brightness; V i,j is the brightness value of the pixel in the i-th row and j-th column in the image, 1≤i≤h, 1≤j≤b; h is the height of the original image; b is the width of the original image;

[0075] Preset the RMS brightness threshold V A , the brightness value RMS V RMS Below the threshold V A The image with the same brightness is classified as a low-light image, otherwise it is classified as an image with normal brightness;

[0076] Optionally, the RMS threshold V of the brightness valueA The results of the normal illumination sequence in the LSRW dataset show that within different RMS intervals of brightness, 96.71% of the images have a brightness RMS value exceeding 0.3. A Set to 0.3, any image with a RMS brightness value below 0.3 will be classified as a low-light image.

[0077] For images that are initially judged to have normal brightness, a second brightness detection is performed after random cropping, and the second brightness detection result is used as the final brightness detection result to reduce the impact of uneven image brightness on brightness detection.

[0078] In this embodiment, the hybrid attention mechanism is used to improve the self-calibration lighting framework, specifically:

[0079] For low-light images, according to the Retinex theory, the relationship between the desired clear image z and the low-light image y is y = z x, where x represents the lighting component of the image. In the cascaded multi-stage lighting optimization process, a hybrid attention mechanism is introduced to optimize the lighting component. The initial t = 0 stage process is expressed as:

[0080]

[0081] The process from t=1 to T is expressed as:

[0082]

[0083] In the formula, y is the initial image; u t and x t Represent the residual term and illumination of the t-th stage (t=0,1,...,T-1) respectively; H represents the illumination estimation network and is independent of the number of stages; A(·) represents the attention function; F(·) represents the illumination estimation function; att represents the weight map generated by the hybrid attention mechanism; G(·) represents the self-calibration function;

[0084] Among them, the self-calibration module ensures that the output can converge to the same state at each stage of training. The self-calibration module in the self-calibration lighting framework is expressed as:

[0085]

[0086] Where s is the self-correction mapping; K φ is the introduced parameterized operator with learnable parameters φ; v t is the input for illumination estimation at stage t after calibration at each stage; z tis the clear image of the tth stage. The improved self-calibration illumination framework extracts the illumination attention weight map of the initial image, and performs adaptive illumination estimation and enhancement stage by stage based on the attention weight map and Retinex theory. The self-calibration module is used to determine whether the outputs of different stages can converge to the same state, and finally the enhanced image is obtained. The overall framework of the improved network is shown in Figure 1. Figure 3 shown.

[0087] In this embodiment, reference Figure 2 , the hybrid attention mechanism is specifically:

[0088] For the input feature map, global average pooling and maximum pooling are spliced ​​in the V channel dimension. The resulting comprehensive illumination feature map containing rich brightness information is refined and fused with brightness features through convolution operations to enhance the model's sensitivity and recognition ability to brightness changes. The illumination attention weight map is then generated through the Sigmoid activation function to identify the importance of each position in the input feature map. At the same time, the self-calibration module is combined to ensure that the output can converge to the same state, thereby achieving convergence between stages and the overall network.

[0089] In this embodiment, the improved ORB feature point extraction method is used to construct an adaptive threshold through the standard deviation of the image grayscale value, and the adaptive threshold is applied to the image with different illumination intensities to extract the feature points. Specifically, the ORB feature points include FAST key points and BRIEF descriptors, and the FAST key points are identified by detecting the grayscale changes in the image, and the BRIEF descriptors are used to generate binary features of the key points;

[0090] The method of identifying FAST key points by detecting grayscale changes in an image is specifically as follows: using the standard deviation of the image grayscale value to set two adaptive FAST thresholds θ and θ min , and the threshold θ>θ min ; Use threshold θ to extract feature points from the image. If the threshold θ fails to extract feature points from the image, a relatively small threshold θ is used to ensure that enough feature points are obtained. min Re-extract feature points from the image.

[0091] In this embodiment, the two adaptive FAST thresholds θ and θ are set. min , specifically:

[0092] For a pixel point p in the image, the grayscale value standard deviation δ of the image is used to dynamically adjust the comparison threshold θ between the brightness values ​​of point p and the surrounding 16 pixels on a circle with a radius of 3:

[0093]

[0094] In the formula, η is a parameter that controls the threshold value of feature point extraction; δ is the standard deviation of the image grayscale value, expressed as:

[0095]

[0096] Where, μ is the mean gray value of the image; I(i,j) is the gray value of the image at position (i,j); i and j represent the position of the pixel 1≤i≤h, 1≤j≤b; h and b represent the height and width of the image respectively;

[0097] Threshold θ min It is expressed as:

[0098]

[0099] After extracting the ORB feature points that meet the requirements, proceed to the next step of calculation.

[0100] In this embodiment, the instance segmentation network YOLOv8-seg is used to segment the normal illumination image and the enhanced image, obtain the mask of the prior dynamic object in the environment, and remove the feature points falling on the mask of the prior dynamic object, and filter out the non-prior dynamic feature points in the environment by the motion consistency principle, which specifically includes the following steps:

[0101] Input the normal illumination image and the enhanced image into the instance segmentation network YOLOv8-seg, and segment them based on the prior dynamic objects. For example, segment the human as the prior dynamic object, obtain the mask of the prior dynamic object in the image, and remove the ORB feature points that fall on the mask of the prior dynamic object.

[0102] The optical flow pyramid is used to calculate the matching relationship between feature points in two adjacent frames. The basic matrix F is obtained through the RANSAC algorithm, the epipolar line I is calculated, and the distance D from the matching point to the epipolar line is calculated to determine whether the point is a dynamic feature point. Only static feature points in the image are retained.

[0103] In this embodiment, the optical flow pyramid is used to calculate the matching relationship between feature points in two adjacent frames, the basic matrix F is obtained by the RANSAC algorithm, the epipolar line I is calculated, and the distance D from the matching point to the epipolar line is calculated to determine whether the point is a dynamic feature point, and only the static feature points in the image are retained, specifically:

[0104] Let p 1 and p 2 is the coordinates of the matching points in the previous frame and the current frame, P 1 and P 2 Yes 1 and p 2 The homogeneous coordinate form of is:

[0105] p 1=[u 1 ,v 1 ], p 2 =[u 2 ,v 2 ], P 1 =[u 1 ,v 1 ,1],P 2 =[u 2 ,v 2 ,1]

[0106] Where u and v represent pixel coordinate values, then the epipolar equation I 1 for:

[0107]

[0108] In the formula, X, Y, Z are line vectors, F is the basic matrix, and the distance from each matching point to the epipolar line is:

[0109]

[0110] If the distance D from the matching point to the epipolar line is greater than the preset threshold λ, the point is judged as a dynamic feature point and filtered out, and the remaining static feature points are retained.

[0111] In this embodiment, performing inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain the best matching feature points specifically includes:

[0112] Four pairs of non-collinear static feature points are randomly selected in the previous frame and the current frame to calculate the homography matrix H, which is recorded as model M. Among them, facing a pair of 2D points, the relationship between the image coordinates can be directly obtained through the homography matrix:

[0113]

[0114] In the formula, p 1 and p 2 are the coordinates of the matching points in the previous frame and the current frame; H is the homography matrix of the matching of two feature points; K is the camera intrinsic parameter matrix; P is the plane where the feature points are located; R, t represent the movement of the camera in the rotation and translation coordinate systems respectively;

[0115] Use model M to transform all ORB feature points in the RGB image frame, and calculate the projection error between each ORB feature point and its corresponding position in the model through the homography matrix H; if the error of an ORB feature point is less than a predetermined threshold, it is regarded as an inlier and added to the inlier set; check whether the inlier set of the current iteration contains more inliers than the previously recorded optimal inlier set, if so, update the optimal inlier set and record the current iteration number k; if the current iteration number reaches the preset maximum iteration number, stop the iteration, and the final model obtained is the optimal model, achieving the best inter-frame feature matching.

[0116] A person skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, the steps of the method can be implemented. The storage medium, such as ROM / RAM, a disk, an optical disk, etc.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A visual SLAM method suitable for low-light dynamic environments, characterized in that: The steps include: S1: Obtain an RGB image of the scene, convert the image from the RGB color space to the HSV color space, perform brightness detection on the V channel of the image to divide the image with normal brightness and the image with low light intensity, use the hybrid attention mechanism to improve the self-calibration illumination framework and perform image enhancement on the low light intensity image; S2: using the improved ORB feature point extraction method, constructing an adaptive threshold through the standard deviation of the image gray value, and applying the adaptive threshold to extract feature points from the images with different light intensities obtained in step S1; S3: Use the instance segmentation network YOLOv8-seg to segment the normal illumination image and the enhanced image, obtain the mask of the prior dynamic object in the environment, and remove the feature points that fall on the mask of the prior dynamic object, and filter out the non-prior dynamic feature points in the environment through the motion consistency principle; S4: Perform inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain the best matching feature points, perform camera pose estimation based on the RANSAC algorithm and the best matching feature points to obtain the camera motion results.

2. The visual SLAM method suitable for low-light dynamic environments according to claim 1, characterized in that: The RGB image of the scene is acquired using an RGB-D camera.

3. The visual SLAM method suitable for low-light dynamic environments according to claim 1, characterized in that: The brightness detection is performed on the V channel of the image to divide the image with normal brightness into the image with low light intensity, specifically: Calculate the root mean square of the brightness value of the image V channel: Where V RMS is the RMS value of image brightness; V i,j is the brightness value of the pixel in the i-th row and j-th column in the image, 1≤i≤h, 1≤j≤b; h is the height of the original image; b is the width of the original image; Preset the RMS brightness threshold V A , the brightness value RMS V RMS Below the threshold V A The image with the same brightness is classified as a low-light image, otherwise it is classified as an image with normal brightness; For the image that is initially judged to be of normal brightness, a second brightness detection is performed after random cropping, and the second brightness detection result is used as the final brightness detection result.

4. The visual SLAM method suitable for low-light dynamic environments according to claim 1, characterized in that: The hybrid attention mechanism is used to improve the self-calibration lighting framework, specifically: In the cascaded multi-stage illumination optimization process, a hybrid attention mechanism is introduced to optimize the illumination component. The initial t=0 stage process is expressed as: The process from t=1 to T-1 is expressed as: Where y is the initial image; u t and x t Represent the residual term and illumination of the t-th stage (t=0,1,...,T-1) respectively; H represents the illumination estimation network and is independent of the number of stages; A(·) represents the attention function; F(·) represents the illumination estimation function; att represents the weight map generated by the hybrid attention mechanism; G(·) represents the self-calibration function; Among them, the self-calibration module in the self-calibration illumination framework is expressed as: Where s is the self-correction mapping; K φ is the introduced parameterized operator with learnable parameters φ; v t is the input for illumination estimation at stage t after calibration at each stage; z t is the clear image at stage t.

5. A visual SLAM method suitable for low-light dynamic environments according to claim 4, characterized in that: The hybrid attention mechanism is specifically: For the input feature map, global average pooling and maximum pooling are concatenated in the V channel dimension. The resulting comprehensive illumination feature map is refined and fused with brightness features through convolution operations, and then the illumination attention weight map is generated through the Sigmoid activation function.

6. The visual SLAM method suitable for low-light dynamic environment according to claim 1, characterized in that: The improved ORB feature point extraction method is used to construct an adaptive threshold through the standard deviation of the image grayscale value, and the adaptive threshold is applied to the images with different light intensities to extract feature points. Specifically, the ORB feature points include FAST key points and BRIEF descriptors, and the FAST key points are identified by detecting the grayscale changes in the image, and the BRIEF descriptors are used to generate binary features of the key points; The method of identifying FAST key points by detecting grayscale changes in an image is specifically as follows: using the standard deviation of the image grayscale value to set two adaptive FAST thresholds θ and θ min , and the threshold θ>θ min ; Use threshold θ to extract feature points from the image. If the feature points cannot be extracted from the image using threshold θ, then use threshold θ min Re-extract feature points from the image.

7. A visual SLAM method suitable for low-light dynamic environments according to claim 6, characterized in that: The two adaptive FAST thresholds θ and θ are set min , specifically: For a pixel point p in the image, the grayscale value standard deviation δ of the image is used to dynamically adjust the comparison threshold θ between the brightness values ​​of point p and the surrounding 16 pixels on a circle with a radius of 3: In the formula, η is a parameter that controls the threshold value of feature point extraction; δ is the standard deviation of the image grayscale value, expressed as: Where, μ is the mean gray value of the image; I(i,j) is the gray value of the image at position (i,j); i and j represent the position of the pixel 1≤i≤h, 1≤j≤b; h and b represent the height and width of the image, respectively; Threshold θ min It is expressed as:

8. The visual SLAM method suitable for low-light dynamic environment according to claim 1, characterized in that: The example segmentation network YOLOv8-seg is used to segment the normal illumination image and the enhanced image, obtain the mask of the prior dynamic object in the environment, and remove the feature points falling on the mask of the prior dynamic object, and filter out the non-prior dynamic feature points in the environment by the motion consistency principle, which specifically includes the following steps: The normal illumination image and the enhanced image are input into the instance segmentation network YOLOv8-seg, segmented based on the prior dynamic objects, and the masks of the prior dynamic objects in the image are obtained, and the ORB feature points falling on the masks of the prior dynamic objects are removed; The optical flow pyramid is used to calculate the matching relationship between feature points in two adjacent frames. The basic matrix F is obtained through the RANSAC algorithm, the epipolar line I is calculated, and the distance D from the matching point to the epipolar line is calculated to determine whether the point is a dynamic feature point. Only static feature points in the image are retained.

9. The visual SLAM method suitable for low-light dynamic environment according to claim 8, characterized in that: The optical flow pyramid is used to calculate the matching relationship between feature points in two adjacent frames, the basic matrix F is obtained through the RANSAC algorithm, the epipolar line I is calculated, and the distance D from the matching point to the epipolar line is calculated to determine whether the point is a dynamic feature point, and only the static feature points in the image are retained. Specifically, Let p1 and p2 be the coordinates of the matching points in the previous frame and the current frame, and P1 and P2 are the homogeneous coordinate forms of p1 and p2: p1=[u1,v1], p2=[u2,v2], P1=[u1,v1,1], P2=[u2,v2,1] Where u and v represent pixel coordinate values, then the epipolar equation I1 is: In the formula, X, Y, Z are line vectors, F is the basic matrix, and the distance from each matching point to the epipolar line is: If the distance D from the matching point to the epipolar line is greater than the preset threshold λ, the point is judged as a dynamic feature point and filtered out, and the remaining static feature points are retained.

10. The visual SLAM method suitable for low-light dynamic environment according to claim 1, characterized in that: The step of performing inter-frame feature point matching on all static feature points in the obtained ORB feature points to obtain the best matching feature points specifically includes: Four pairs of non-collinear static feature points are randomly selected in the previous frame and the current frame to calculate the homography matrix H, which is recorded as model M. Among them, facing a pair of 2D points, the relationship between the image coordinates can be directly obtained through the homography matrix: Where p1 and p2 are the coordinates of the matching points in the previous frame and the current frame; H is the homography matrix of the matching of two feature points; K is the camera intrinsic parameter matrix; P is the plane where the feature points are located; R, t represent the movement of the camera in the rotation and translation coordinate systems respectively; Use model M to transform all ORB feature points in the RGB image frame, and calculate the projection error between each ORB feature point and its corresponding position in the model; if the error of an ORB feature point is less than a predetermined threshold, it is regarded as an inlier and added to the inlier set; check whether the inlier set of the current iteration contains more inliers than the previously recorded optimal inlier set, if so, update the optimal inlier set and record the current iteration number k; if the current iteration number reaches the preset maximum iteration number, stop the iteration, and the final model obtained is the optimal model, achieving the best inter-frame feature matching.

Citation Information

Cited By

  • Deep learning feature matching method for low-light environment

    CN120726355A