A visual SLAM construction method suitable for low-light environments
By using event cameras and inertial measurement units in the visual SLAM system, combined with Arc* and Harris algorithms for feature detection and refinement screening, the problem of feature loss and tracking failure of visual SLAM system in low-light environments is solved, the positioning accuracy and robustness are improved, and the precise positioning of mobile robots in low-light environments is supported.
Patent Information
- Application Number
- CN202310742418.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-31
- Filing Date
- 2023-06-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-06-21
AI Technical Summary
In an environment with insufficient lighting, the visual SLAM system will have features loss and feature tracking failure, resulting in reduced positioning accuracy and poor robustness.
The event camera is used to collect event streams in a low-light environment, and generate TS and PTS through SAE updates, filter candidate event characteristics, use Arc* and Harris algorithms for feature detection and refinement screening, and combine the pre-integration constraints of the inertial measurement unit to construct the state variables of the SLAM system, and perform back-end optimization to improve positioning accuracy.
It improves the positioning accuracy and robustness of the visual SLAM system in low-light environments, solves the problem of degradation of positioning accuracy caused by feature loss and tracking failure, and supports the precise positioning of mobile robots in low-light environments.
Smart Images

Figure CN116758311B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer vision positioning, and in particular relates to a visual SLAM construction method suitable for a weak light environment. Background Art
[0002] With the rapid development of artificial intelligence technology, research in the fields of unmanned driving, virtual reality, and face recognition has become a current hot topic. Among them, unmanned driving has achieved lane-level positioning effects in some urban environments with known maps, but when driving on unstructured roads with unknown environments and signal source positioning sensors such as GPS, Beidou, and Galileo cannot be used, how to achieve autonomous map construction and precise positioning is one of the research difficulties in this field. SLAM (Simultaneous Localization and Mapping) refers to the method in which an unmanned platform uses signal source sensors such as cameras, Lidar, odometers, inertial sensors and signal source sensors such as GPS and Beidou in an unknown environment to realize the perception of the environment by the unmanned platform and simultaneously build a high-precision map and locate the body posture. SLAM is the premise and foundation for realizing autonomous navigation and environmental exploration. However, in an environment with insufficient lighting, visual SLAM will have problems such as feature loss and feature tracking failure, resulting in a decrease in the positioning accuracy and poor robustness of the system. Therefore, it is urgent to study a visual SLAM construction method suitable for low-light environments. Summary of the invention
[0003] In view of the problems and shortcomings in the prior art, the object of the present invention is to provide a visual SLAM construction method suitable for low-light environments.
[0004] To achieve the purpose of the invention, the technical solution adopted by the present invention is as follows:
[0005] The present invention provides a visual SLAM construction method suitable for a weak light environment, and the visual SLAM construction method suitable for a weak light environment comprises the following steps:
[0006] S1: Collect the event stream output by the event camera and the measurement value output by the inertial measurement unit mounted on the same motion platform;
[0007] S2: Update SAE (Surface of Active Events, SAE) according to the fixed frequency of the event stream, and use the updated SAE to generate TS (Time of Surface, TS) and PTS (Polarity Time of Surface, PTS) according to the frequency of the event stream;
[0008] S3: using the TS as a mask, traversing the event stream, finding the pixel value corresponding to each event in the event stream on the TS, screening out events whose pixel values are greater than a set threshold, and obtaining candidate events;
[0009] S4: performing candidate feature detection on the candidate event on the SAE to obtain candidate event features; performing refinement screening on the candidate event features to screen out event corner points from the candidate event features;
[0010] S5: obtaining speed information and depth information of the event corner point, and using the position information, speed information, depth information and TS of the event corner point as visual information;
[0011] S6: constructing pre-integration constraints between adjacent PTSs using the measurement values output by the inertial measurement unit to obtain pre-integration information;
[0012] S7: Initialize the SLAM system, input the visual information and the pre-integrated information with the same timestamp into the sliding window model of the back end of the SLAM system that has been initialized for processing, and obtain the state variables of the SLAM system at the current moment;
[0013] S8: Using the state variables of the SLAM system at the current moment obtained in step S7, through the position difference of the PTS at adjacent moments and the parallax of the corresponding event corner points, determine whether the TS generated in step S2 is a key frame; if it is a key frame, execute step S9; if it is not a key frame, return to step S2 to process the event stream at the next moment;
[0014] S9: Perform loop detection on the key frame. If a historical loop frame of the key frame is detected, execute step S10; if no historical loop frame of the key frame is detected, add the key frame to the key frame database, and then return to step S2 to process the event stream at the next moment;
[0015] S10: finding a feature matching relationship between the key frame and the historical loop frame, then verifying the feature matching relationship, removing wrong matches, and obtaining a posture change between the key frame and the historical loop frame;
[0016] S11: Input the posture change between the key frame and the historical loop frame as the loop state quantity into the sliding window model of the SLAM system backend for optimization to obtain the optimized loop state quantity; perform posture graph optimization on the optimized loop state quantity to eliminate the trajectory error between the key frame and the historical loop frame, and use the trajectory after eliminating the error to update the key frame database.
[0017] According to the above-mentioned visual SLAM construction method, preferably, in step S2, the formula for generating TS according to the frequency of the event stream using the updated SAE is:
[0018]
[0019] The formula for generating PTS according to the frequency of the event stream using the updated SAE in step S2 is:
[0020]
[0021] Where (x, y) is the pixel position on the SAE, t is the time when TS is generated, and t last is the timestamp of the previous event at pixel (x, y), δ is the constant decay parameter, exp is the exponential function; p is the polarity of the event at pixel (x, y).
[0022] According to the above-mentioned visual SLAM construction method, preferably, in step S4, the Arc* algorithm is used to perform candidate feature detection on the candidate event on the SAE.
[0023] According to the above-mentioned visual SLAM construction method, preferably, in step S4, Harris algorithm is used to refine and screen the candidate event features.
[0024] According to the above-mentioned visual SLAM construction method, preferably, the method for obtaining the speed information of the event corner point in step S5 is: for the event corner point, using LK optical flow on PTS to track the event corner point at the previous moment, and obtaining the speed information of the event corner point on the pixel plane.
[0025] According to the above-mentioned visual SLAM construction method, preferably, the method for obtaining the depth information of the event corner point in step S5 is: performing dedistortion and triangulation processing on the event corner point to obtain the depth information of the event corner point.
[0026] According to the above-mentioned visual SLAM construction method, preferably, in step S7, the specific operation of inputting the visual information and the pre-integration information with consistent timestamps into the sliding window model of the back end of the SLAM system that has been initialized for processing is: using the visual information and the IMU pre-integration information in the sliding window model to construct visual residuals and pre-integration residuals, and obtaining the state variables of the SLAM system at the current moment by performing nonlinear optimization on the visual residuals and pre-integration residuals.
[0027] According to the above-mentioned visual SLAM construction method, preferably, the nonlinear optimization method in step S7 is the Gauss-Newton method.
[0028] According to the above-mentioned visual SLAM construction method, preferably, the specific operation of step S10 is: extract FAST feature points from the key frame, use the Brief descriptor algorithm to calculate the descriptors of the FAST feature points and event corner points of the key frame, find the feature matching relationship between the key frame and the historical loop frame, and then use the PnP method to verify the feature matching relationship, remove the wrong matches, and obtain the posture changes of the key frame and the historical loop frame. More preferably, the number of FAST feature points extracted from the key frame is related to the scene, and the number of feature points is different for different scenes; for the scene of the present invention, the number of FAST feature points extracted from the key frame is 400.
[0029] According to the above-mentioned visual SLAM construction method, preferably, the measurement values output by the inertial measurement unit in step S1 are linear acceleration and angular velocity; and the pre-integration constraints in step S6 are position constraints, attitude constraints and velocity constraints.
[0030] According to the above-mentioned visual SLAM construction method, preferably, in step S7, pure visual 3D reconstruction and visual inertial alignment are used to initialize the SLAM system.
[0031] According to the above visual SLAM construction method, preferably, in step S9, a bag-of-words model is used for loop detection. By judging the similarity between the TS generated in step S2 and the key frames in the key frame database, it is judged whether there is a loop relationship.
[0032] According to the above-mentioned visual SLAM construction method, preferably, the trajectory error between the key frame and the historical loop frame in step S11 is the error of all PTS postures between the current TS (TS and PTS have the same timestamp because they are both created according to the frequency of the event stream, so the TS at this moment corresponds to one frame of PTS) and the historical loop frame.
[0033] According to the above-mentioned visual SLAM construction method, preferably, the motion platform in step S1 is a mobile robot.
[0034] Compared with the prior art, the present invention has the following positive and beneficial effects:
[0035] (1) Event camera is a biologically inspired visual sensor that outputs an event when the brightness change of a pixel exceeds a threshold. Compared with a standard camera, event camera has the advantages of high dynamic range, high temporal resolution (low latency), and no motion blur. The present invention adopts event camera to collect event streams in low-light environment, and performs event feature screening and detection on the event stream, so as to improve the positioning accuracy and robustness of the visual SLAM system in low-light environment, and better support the positioning of mobile robots in low-light environment. At the same time, it also solves the technical problem that the existing visual SLAM construction method suffers from feature loss and feature tracking failure in an environment with insufficient lighting, resulting in reduced positioning accuracy.
[0036] (2) The present invention uses TS as a mask to traverse the event stream used to update SAE, find the pixel value corresponding to each event in the event stream on TS, filter out events with pixel values greater than a set threshold, and then perform feature screening. This operation can greatly reduce the number of events used for feature detection, ensure that only the latest events are subsequently processed, and ensure the real-time nature of subsequent processing.
[0037] (3) After using the Arc* algorithm to detect event features, the present invention further uses the Harris algorithm to refine and screen the candidate event features. After the refinement and screening by the Harris algorithm, the event features can be extracted more accurately. At the same time, it also solves the technical problem that the Arc* algorithm is inaccurate in detecting event features and will erroneously detect event features.
[0038] (4) The present invention inputs the visual information and the pre-integration information with consistent timestamps into the sliding window model of the back-end of the SLAM system that has completed the initialization for optimization processing, wherein the visual information is the reprojection error of the event feature, and the IMU information is the pre-integration residual. In the back-end optimization, by reducing the reprojection error and the pre-integration residual, the position and posture of the mobile robot are updated, and the position and posture of the mobile robot approaching the optimal can be obtained.
[0039] (5) The present invention can eliminate the cumulative error of the robot's moving trajectory and improve the accuracy of detection by optimizing the posture graph of the optimized loop state quantity. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 The figure is a flow chart of the visual SLAM construction method suitable for low-light environment of the present invention, which is divided into two parts: the front end and the back end;
[0041] Figure 2 It is a schematic diagram of the flow of front-end calculation in the visual SLAM construction method suitable for low-light environment of the present invention;
[0042] Figure 3It is a flow chart of front-end event feature detection in the visual SLAM construction method suitable for weak light environment of the present invention;
[0043] Figure 4 The schematic diagram is a principle diagram of the present invention using the Arc* algorithm to perform candidate feature detection on the candidate event on SAE;
[0044] Figure 5 The figure is a schematic diagram of the optimization process in the visual SLAM construction method suitable for low-light environments of the present invention. DETAILED DESCRIPTION
[0045] The technical solutions in the implementation process of the present invention will be clearly and detailedly introduced below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described in the present invention are only part of the embodiments of the present invention, not all the embodiments.
[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. When the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, components and / or combinations thereof. The experimental methods in the following examples where specific conditions are not specified are all based on conventional techniques in the art.
[0048] Embodiment 1:
[0049] A visual SLAM construction method suitable for low-light environments, such as Figure 1 As shown, it includes front-end calculation and back-end optimization.
[0050] Among them, Figure 2 , Figure 3 As shown, the steps of front-end calculation are as follows:
[0051] S1: An event camera and an inertial measurement unit (IMU) are mounted on the mobile robot. After the mounting is completed, the internal and external parameters of the event camera and IMU are calibrated. After completing the above operations, the mobile robot equipped with the sensor moves in a scene with insufficient lighting, and collects the event stream output by the event camera and the measurement values output by the inertial measurement unit. The measurement values output by the inertial measurement unit are linear acceleration and angular velocity. Among them, the SLAM system receives the event stream output by the event camera, and the event is expressed as (x, y, t, p), where x and y represent the location where the event occurs, t represents the timestamp of the event, and p represents the polarity of the event. p=1 means that the brightness of the pixel corresponding to the event is enhanced, and p=-1 means that the brightness of the pixel position corresponding to the event is reduced.
[0052] S2: Update SAE (Surface of Active Events, SAE) according to the fixed frequency of the event stream, and use the updated SAE to generate TS (Time of Surface, TS) and PTS (Polarity Time of Surface, PTS) according to the frequency of the event stream.
[0053] SAE is a two-dimensional image, and each pixel stores the timestamp of the latest event at that pixel position. The formula for generating TS using the updated SAE according to the frequency of the event stream is:
[0054]
[0055] The formula for generating PTS according to the frequency of the event stream using the updated SAE in step S2 is:
[0056]
[0057] Where (x, y) is the pixel position on the SAE, t is the time when TS is generated, and t last is the timestamp of the previous event at pixel (x, y), δ is the constant decay parameter, exp is the exponential function; p is the polarity of the event at pixel (x, y).
[0058] S3: using the TS as a mask, traversing the event stream, finding the pixel value corresponding to each event in the event stream on the TS, screening out events whose pixel values are greater than a set threshold, and obtaining candidate events;
[0059] S4: Using the Arc* algorithm to perform candidate feature detection on the candidate event on the SAE to obtain candidate event features; then using the Harris algorithm to refine and screen the candidate event features to screen out event corner points from the candidate event features.
[0060] The principle of using the Arc* algorithm to detect candidate features for the candidate events on SAE is as follows (e.g. Figure 4 As shown in Figure 3, the corner point is the intersection of two edges. When the event camera and the scene move relative to each other, the event streams triggered by the two edges associated with the corner point show significant differences in the distribution of SAE, with one part having a significantly higher value than the other. Figure 4 The white area on the SAE indicates an area without event triggering, and the blue area indicates an area with events. The darker the color, the greater the pixel intensity (timestamp) on the SAE, and the newer the triggered event.
[0061] The specific operation of using the Harris algorithm to refine and screen the candidate event features is as follows: first, for each candidate event feature, take a 9×9 template on the SAE with the candidate event feature as the center; then construct a binary pixel block P of the same size of 9×9, and set the pixel coordinates on this pixel block P to be (i, j), where i and j represent the vertical coordinate and horizontal coordinate of the pixel on P respectively; arrange the timestamps stored in all pixels in the template in ascending order, and select the first 25 timestamps from them, and record the pixel positions (p, q) of the selected 25 timestamps, where p and q represent the vertical coordinate and horizontal coordinate of the pixel on the patch respectively, and the pixels at the corresponding positions on the pixel block P are set to 1, and the other positions are defaulted to 0, that is:
[0062]
[0063] Then, the Harris score of each candidate event feature is calculated using the binary pixel block. If the Harris score is greater than a threshold, the candidate event feature is considered to be an event corner point.
[0064] S5: For the event corner point, LK optical flow is used on PTS to track the event corner point at the previous moment to obtain the velocity information of the event corner point on the pixel plane; then the event corner point is dedistorted and triangulated to restore the depth information of the event corner point; the position information, velocity information, depth information and TS of the event corner point are used as visual information;
[0065] S6: constructing pre-integration constraints between adjacent PTSs using the measurement values output by the inertial measurement unit to obtain pre-integration information; the pre-integration constraints are position constraints, attitude constraints and velocity constraints.
[0066] like Figure 5 As shown, the specific steps of backend optimization are as follows:
[0067] S7: Determine whether the SLAM system is initialized. If the SLAM system is not initialized, use pure visual three-dimensional reconstruction and visual inertial alignment to initialize the SLAM system until the SLAM system is successfully initialized. Input the visual information and the pre-integration information with consistent timestamps into the sliding window model of the back end of the SLAM system that has completed the initialization. Use the visual information and the IMU pre-integration information in the sliding window model to construct visual residuals and pre-integration residuals. Use the Gauss-Newton method to perform nonlinear optimization on the visual residuals and pre-integration residuals to obtain the state variables of the SLAM system at the current time (i.e., the current PTS). The state variables include position, posture, speed, and IMU zero bias. Using the sliding window model optimization can limit the scale of optimization and ensure the real-time performance of the SLAM system.
[0068] S8: Using the state variables of the SLAM system at the current moment obtained in step S7, determine whether the TS generated in step S2 is a key frame through the position difference of the PTS at adjacent moments and the parallax of the corresponding event corner points; if it is a key frame, execute step S9; if it is not a key frame, return to step S2 to process the event stream at the next moment.
[0069] S9: Use the bag-of-words model to perform loop detection on the key frame (use the bag-of-words model to determine the similarity between the key frame and the key frames stored in the key frame database). If a historical loop frame of the key frame is detected, execute step S10; if no historical loop frame of the key frame is detected, add the key frame to the key frame database, and then return to step S2 to process the event stream at the next moment.
[0070] S10: Extract FAST feature points from the key frame, use the Brief descriptor algorithm to calculate the descriptors of the FAST feature points and event corner points of the key frame, find the feature matching relationship between the key frame and the historical loop frame (the feature matching relationship is the matching event corner points), and then use the PnP method to verify the feature matching relationship, remove false matches, and obtain the posture change between the key frame and the historical loop frame.
[0071] S11: Input the posture change between the key frame and the historical loop frame as the loop state quantity into the sliding window model of the SLAM system backend for optimization to obtain the optimized loop state quantity; perform posture graph optimization on the optimized loop state quantity to eliminate the trajectory error between the key frame and the historical loop frame, and use the trajectory after eliminating the error to update the key frame database.
[0072] The above description is only a preferred embodiment of the present invention, but is not limited to the above examples. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A visual SLAM construction method suitable for low-light environments, It is characterized in that The following steps are involved: S1: Collect the event stream output by the event camera and the measurement value output by the inertial measurement unit mounted on the same motion platform; S2: updating the SAE according to the fixed frequency of the event stream, and using the updated SAE to generate the TS and PTS according to the frequency of the event stream; S3: using the TS as a mask, traversing the event stream, finding the pixel value corresponding to each event in the event stream on the TS, screening out events whose pixel values are greater than a set threshold, and obtaining candidate events; S4: performing candidate feature detection on the candidate event on the SAE to obtain candidate event features; Refining and screening the candidate event features, and screening out event corner points from the candidate event features; S5: obtaining speed information and depth information of the event corner point, and using the position information, speed information, depth information and TS of the event corner point as visual information; S6: constructing pre-integration constraints between adjacent PTSs using the measurement values output by the inertial measurement unit to obtain pre-integration information; S7: Initialize the SLAM system, input the visual information and the pre-integrated information with the same timestamp into the sliding window model of the back end of the SLAM system that has been initialized for processing, and obtain the state variables of the SLAM system at the current moment; S8: Using the state variables of the SLAM system at the current moment obtained in step S7, through the position difference of the PTS at adjacent moments and the parallax of the corresponding event corner points, determine whether the TS generated in step S2 is a key frame; if it is a key frame, execute step S9; if it is not a key frame, return to step S2 to process the event stream at the next moment; S9: Perform loop detection on the key frame. If a historical loop frame of the key frame is detected, execute step S10; if no historical loop frame of the key frame is detected, add the key frame to the key frame database, and then return to step S2 to process the event stream at the next moment; S10: finding a feature matching relationship between the key frame and the historical loop frame, then verifying the feature matching relationship, removing wrong matches, and obtaining a posture change between the key frame and the historical loop frame; S11: Input the posture change between the key frame and the historical loop frame as the loop state quantity into the sliding window model of the SLAM system backend for optimization to obtain the optimized loop state quantity; perform posture graph optimization on the optimized loop state quantity to eliminate the trajectory error between the key frame and the historical loop frame, and use the trajectory after eliminating the error to update the key frame database.
2. The visual SLAM construction method according to claim 1, It is characterized in that The formula for generating TS according to the frequency of the event stream using the updated SAE in step S2 is: The formula for generating PTS according to the frequency of the event stream using the updated SAE in step S2 is: Where (x, y) is the pixel position on the SAE, t is the time when TS is generated, and t last is the timestamp of the previous event at pixel (x, y), δ is the constant decay parameter, exp is the exponential function; p is the polarity of the event at pixel (x, y).
3. The visual SLAM construction method according to claim 2, It is characterized in that In step S4, the Arc* algorithm is used to perform candidate feature detection on the candidate event on the SAE.
4. The visual SLAM construction method according to claim 3, It is characterized in that In step S4, the Harris algorithm is used to refine and screen the candidate event features.
5. The visual SLAM construction method according to claim 4, It is characterized in that The method for obtaining the speed information of the event corner point in step S5 is: for the event corner point, using LK optical flow on PTS to track the event corner point at the previous moment, and obtaining the speed information of the event corner point on the pixel plane.
6. The visual SLAM construction method according to claim 5, It is characterized in that The method for obtaining the depth information of the event corner point in step S5 is: performing dedistortion and triangulation processing on the event corner point to obtain the depth information of the event corner point.
7. The visual SLAM construction method according to claim 6, It is characterized in that In step S7, the specific operation of inputting the visual information and the pre-integration information with consistent timestamps into the sliding window model of the back end of the SLAM system that has been initialized for processing is: using the visual information and IMU pre-integration information in the sliding window model to construct visual residuals and pre-integration residuals, and obtaining the state variables of the SLAM system at the current moment by performing nonlinear optimization on the visual residuals and pre-integration residuals.
8. The visual SLAM construction method according to claim 7, It is characterized in that The nonlinear optimization method in step S7 is the Gauss-Newton method.
9. The visual SLAM construction method according to claim 8, It is characterized in that The specific operation of step S10 is: extract FAST feature points from the key frame, use the Brief descriptor algorithm to calculate the descriptors of the FAST feature points and event corner points of the key frame, find the feature matching relationship between the key frame and the historical loop frame, and then use the PnP method to verify the feature matching relationship, remove incorrect matches, and obtain the posture changes of the key frame and the historical loop frame.
10. The visual SLAM construction method according to claim 9, It is characterized in that The measurement values output by the inertial measurement unit in step S1 are linear acceleration and angular velocity; the pre-integration constraints in step S6 are position constraints, attitude constraints and velocity constraints.
Citation Information
Patent Citations
Panoramic inertial navigation SLAM method based on multiple key frames
CN109307508A
Event and distance fused visual inertial odometer method
CN115479602A