Binocular vision positioning method and system in weak light environment based on deep learning enhancement
The method uses RCF, SuperPoint, and SuperGlue models with LSD to enhance feature extraction and matching in low-light conditions, improving SLAM system precision and stability by filtering noise and optimizing keyframes, suitable for resource-constrained applications.
Patent Information
- Application Number
- CN202510819725.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Under low light conditions, the feature extraction and matching stability of the visual SLAM system leads to a reduced positioning accuracy. The existing geometric and deep learning-based methods have noise interference in low-light environments, affecting the positioning results.
The edge detection model (RCF) based on deep learning is used for edge feature enhancement, combined with linear detection algorithm and ray erase logic to eliminate spurious line segments, point feature extraction and matching is used using SuperPoint and SuperGlue models, and poses are optimized through keyframe selection and reprojection residual model.
It improves the stability of feature extraction and positioning accuracy in low-light environments, effectively suppresses light scattering interference, improves the robustness and real-timeness of the positioning system, and is suitable for application scenarios where computing resources are limited.
Smart Images

Figure CN120318479A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a binocular vision positioning method and system under low-light environment enhanced by deep learning. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] In recent years, with the rapid development of computer vision and robot technologies, Visual SLAM (Simultaneous Localization and Mapping) technology has been widely applied in autonomous driving, robot navigation, augmented reality (AR), virtual reality (VR), 3D reconstruction, industrial fields and even agricultural fields due to its advantages such as low hardware cost and rich information. The Visual SLAM system collects environmental images through a camera, extracts feature points or line features, and estimates the camera motion trajectory by combining multi-frame data, so as to achieve autonomous positioning and environmental modeling.
[0004] Visual SLAM technology is not only affected by the quality of picture data. For example, under low-light conditions, the number of photons received by the camera's photosensitive element decreases, resulting in increased image noise and reduced contrast, which in turn affects the stability of feature extraction. At the same time, how to ensure the stability of extracting picture feature points and the stability of feature matching between picture frames will also have a great impact on the accuracy of backend optimization. At present, the methods for extracting and matching feature points in the front end of Visual SLAM are mainly divided into geometric-based methods and deep learning-based methods. However, in the case of insufficient light, the positioning results are still not ideal. Low light will introduce noise when extracting feature points for traditional geometric-based methods, and will also introduce noise for deep learning-based feature point extraction methods, resulting in a reduction in the positioning accuracy of the SLAM system that calculates the inter-frame pose based on feature points. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a binocular vision localization method and system in low-light environments enhanced by deep learning. An edge detection model based on a convolutional neural network (Richer Convolutional Features, RCF) is used to enhance the edge features of the original camera image. Secondly, an image point feature extraction model, such as the SuperPoint model, and an image matching model, such as the SuperGlue model, are used to extract point features from the original image and perform matching. Then, a line detection algorithm (Line Segment Detector, LSD) is used to extract line features, and a light erasure logic is used during this process to reduce the impact of scattered imaging caused by point light sources. Finally, key frames of the image are selected for inter-frame pose calculation and a graph optimization algorithm is used to obtain the final trajectory.
[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a binocular vision localization method in low-light environments enhanced by deep learning; A binocular vision localization method in low-light environments enhanced by deep learning includes: Obtain binocular vision images; Use an edge detection model to enhance the edge features of the binocular vision images; Use a line detection algorithm to extract line features from the binocular vision images with enhanced edge features, and use a light erasure logic to eliminate false line segments generated by point light source scattering; Use an image point feature extraction model to extract the point features of the binocular vision images; based on the obtained point features, use the SuperGlue model for feature matching; judge the positional relationship between the point features and the line features, and use the relationship of point feature matching to obtain the relationship of line feature matching; Based on a key frame selection strategy, select frames with position, angle changes, and low feature matching numbers for back-end optimization, and output the final localization trajectory; the back-end optimization includes optimizing the inter-frame pose by constructing a reprojection residual model of point and line features.
[0007] As a further technical solution, the edge detection model is based on an improved VGG16 network structure, the fully connected layer and the fifth pooling layer are deleted, and edge information in the image is extracted through multi-scale feature fusion.
[0008] As a further technical solution, using a line detection algorithm to extract line features from the binocular vision images with enhanced edge features, and using a light erasure logic to eliminate false line segments generated by point light source scattering includes: Extract the line features of the binocular vision images with enhanced edge features and the geometric center coordinates of the light source in the image; Extract the pixel values corresponding to the endpoints of all lines, perform contrast stretching on the line endpoints, and set the astigmatism threshold; Perform grayscale value calculation on the endpoint pixel values after contrast stretching, further exclude the lines whose starting point grayscale values are greater than the astigmatism threshold, and calculate the perpendicular distance from the light source center coordinate to each line, and determine whether it is caused by astigmatism through the distance threshold.
[0009] As a further technical solution, use an image point feature extraction model to extract the point features of the binocular vision image, including: The image point feature extraction model compresses the binocular vision image through an encoder, detects the positions of feature points on the compressed image, and outputs descriptors; Pass the result of the feature point position detection through the softmax function and reshape operation to obtain a tensor with the same size as the binocular vision image; Select the feature point information with high confidence, and evenly distribute the feature points on the binocular vision image through grid allocation.
[0010] As a further technical solution, based on the obtained point features, use the SuperGlue model for feature matching; judge the positional relationship between the point features and the line features, and use the relationship of the point feature matching to obtain the relationship of the line feature matching; including: Based on the descriptors obtained by the image point feature extraction model, use the attention mechanism and the graph neural network to optimize the feature point matching, and filter out the wrong matching pairs through the dustbin mechanism; After obtaining the matching relationship of the point features, for the point and line features on the same image, judge which point features are on the line feature or the position of the point is less than the set matching threshold from the position of the line feature, then define that the point feature belongs to the line feature, and based on the matching relationship of the point features, further obtain the matching relationship of the line features.
[0011] As a further technical solution, the process of optimizing the image pose by constructing a reprojection residual model of the point features includes: After extracting and matching the point features, restore the depth information of the image points through binocular stereo vision; Transform the 3D points in the world coordinate system to the camera coordinate system; The pixel error formed by the image position after the 3D point transformation (i.e., reprojection) and the extracted feature points constitutes the reprojection error of the point features.
[0012] As a further technical solution, the process of optimizing the inter-frame pose by constructing a reprojection residual model of the line features includes: Calculate the depth of the line segment endpoints for the extracted line features to obtain the information of the spatial 3D line segments; The Plücker coordinates are used to represent the 3D line segments in space, and the line segment information in the Plücker coordinate system is obtained. The line segment information in the Plücker coordinate system is transformed into the camera coordinate system. The space line in the camera coordinate system is reprojected onto the image plane, and the reprojection error of the line is represented by the perpendicular distance from the line segment endpoint coordinates to the line where the line segment is located.
[0013] The second aspect of the present invention provides a binocular vision positioning system in low-light environments enhanced by deep learning.
[0014] The binocular vision positioning system in low-light environments enhanced by deep learning includes: An image acquisition module, configured to: acquire binocular vision images; An edge feature enhancement module, configured to: enhance the edge features of the binocular vision images using an edge detection model; A line feature extraction module, configured to: extract line features from the binocular vision images with enhanced edge features using a line detection algorithm, and eliminate the false line segments generated by the scattering of point light sources using a ray erasure logic; A point feature extraction and point-line feature matching module, configured to: extract the point features of the binocular vision images using an image point feature extraction model; based on the acquired point features, perform feature matching using the SuperGlue model; judge the positional relationship between the point features and the line features, and use the relationship of point feature matching to obtain the relationship of line feature matching; A backend optimization module, configured to: based on a key frame selection strategy, screen the frames with position, angle changes, and low feature matching numbers for backend optimization, and output the final positioning trajectory; the backend optimization includes optimizing the inter-frame poses by constructing a reprojection residual model of point and line features.
[0015] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in the binocular vision positioning method in low-light environments enhanced by deep learning as described in the first aspect of the present invention are implemented.
[0016] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the binocular vision positioning method in low-light environments enhanced by deep learning as described in the first aspect of the present invention are implemented.
[0017] The above one or more technical solutions have the following beneficial effects: (1) The present invention can effectively improve the stability of feature extraction in low-light environments. By using the RCF (Richer Convolutional Features) model to enhance the edges of images, the extraction effect of line features under low-light conditions is effectively improved, and the problem of feature loss caused by noise interference in traditional edge detection methods is effectively reduced. Based on the combined use of the SuperPoint and SuperGlue models, the extraction and matching of point features can still maintain high accuracy and robustness under low-light conditions, reducing the false matching rate.
[0018] (2) The present invention automatically identifies and eliminates false line features generated by the scattering of point light sources through the light scattering suppression logic, avoiding the influence of such interference on pose calculation, improving the reliability of line features in low-light environments, and effectively suppressing light scattering interference.
[0019] (3) By using Plücker coordinates to represent space lines, the reprojection error calculation of line features is optimized, further improving the accuracy of pose calculation. By dynamically selecting key frames, redundant calculations are reduced, and the real-time performance of the system is improved while ensuring the positioning accuracy, which is applicable to application scenarios with limited computing resources (such as mobile robots, drones, etc.).
[0020] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0022] Figure 1 It is a flowchart of the method for the first embodiment.
[0023] Figure 2 It is a model architecture diagram of the edge detection model in the first embodiment.
[0024] Figure 3 It is an extraction effect diagram of line features after edge enhancement in the first embodiment.
[0025] Figure 4 It is an extraction effect diagram of line features without edge enhancement in the first embodiment.
[0026] Figure 5 It is an effect diagram after removing the light scattering of the point light source in the first embodiment.
[0027] Figure 6 It is an effect diagram before removing the light scattering of the point light source in the first embodiment.
[0028] Figure 7 Schematic diagram of the matching principle of point features and line features in the first embodiment.
[0029] Figure 8 Schematic diagram of the comparison result between the method of the present invention and other methods in the first embodiment.
[0030] Figure 9 System structure diagram of the second embodiment. Detailed implementation manners
[0031] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0032] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention.
[0033] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0034] Embodiment 1 This embodiment discloses a binocular vision positioning method in a low-light environment based on deep learning enhancement; As Figure 1 shown, the binocular vision positioning method in a low-light environment based on deep learning enhancement includes: Step S1, obtaining binocular vision images; the binocular vision images include a left image and a right image, and the left and right view images are synchronously collected by a calibrated binocular camera; and the collected images are subjected to timestamp alignment and geometric correction.
[0035] Step S2, using an edge detection model to enhance the edge features of the binocular vision images.
[0036] The edge detection model (RCF) is a convolutional neural network model useful for image edge inference. The traditional image edge detection idea mainly focuses on the color gradient of the image, and special algorithms are designed to traverse the image to obtain edge features. However, when extracting line features for visual SLAM, traditional edge detection has unstable line features extracted between frames in a darker imaging environment. The edge detection model uses a fully convolutional network to fuse more image details by combining the convolutional features of other intermediate layers.
[0037] Furthermore, the structure and functions of the edge detection model include the following: Combined with Figure 2, the edge detection model is improved based on VGG16, deleting all fully connected layers and the fifth pooling layer; each convolutional layer is connected to a convolutional layer with a kernel of ; followed by an element-wise operation (eltwise) layer; then use a transposed convolutional layer to upsample the feature map of this layer; connect the upsampled result through Sigmoid cross-entropy loss; fuse the upsampled results of all convolutional layers, and then connect them with a convolutional layer with a kernel of and calculate the final loss.
[0038] The kernels of the backbone convolutional layers of the edge detection model are all in size. The dimension of the input image can be arbitrary. After multiple convolutions, padding, and merging, an image with the same dimension as the input image is obtained.
[0039] Step S3: Use a line detection algorithm to extract line features from the binocular vision image with enhanced edge features, and adopt a ray erasure logic to eliminate false line segments generated by the scattering of point light sources; the effect is as Figures 3 - 6 shown. Among them, Figure 3 , Figure 4 show the line feature extraction effect diagrams before and after edge enhancement. The red line segments in the figure represent the extracted line features. Figure 5 , Figure 6 show the comparison diagrams of the effects before and after removing the light scattering of the point light source. The red line segments in the figure represent the extracted line features, and the green serial numbers represent the serial numbers of the extracted line features.
[0040] Use the line detection algorithm (LSD) to extract the line features of the binocular vision image with enhanced edge features and the geometric center coordinates of the light source in the image ; Extract the pixel values corresponding to the endpoints of all lines, perform contrast stretching on the line segment endpoints, and set the astigmatism threshold ,
[0041] Among them, , , are the pixel values of the line segment endpoints after contrast stretching; , , respectively represent the original image channel values of the line segment endpoints; , , are the image stretching coefficients; , , are the pixel value offsets.
[0042] Calculate the grayscale values of the endpoint pixel values after contrast stretching, further exclude the line segments where the grayscale values of the starting points are all greater than the astigmatism threshold, and calculate the perpendicular distance from the center coordinate to each line. Determine whether it is caused by astigmatism based on the distance threshold.
[0043] Calculate the grayscale values of the endpoint pixel values after contrast stretching.
[0044]
[0045] Among them, is the pixel grayscale value corresponding to the endpoint coordinate of the line feature; 、 、 are the weight coefficients of the image channels.
[0046] For the line segments where the grayscale values of the starting points are all greater than the astigmatism threshold perform further exclusion to obtain ( ); and calculate the perpendicular distance from the optical center coordinate to each line. If it is less than a certain threshold, it can be considered that the line is caused by astigmatism.
[0047]
[0048] Among them, represents the straight-line distance from the i-th optical center to the n-th line segment; 、 represent the pixel coordinates of the i-th optical center; 、 、 、 、 represent the starting pixel coordinates of the n-th line segment. When is less than the distance threshold
[0049] it can be considered that the line segment is generated due to light scattering. All the parameters used above can be appropriately adjusted before the program runs to achieve different degrees of suppression effects. After the above processing, the influence of the false line features generated by astigmatism on pose solution can be greatly suppressed. In this embodiment, the image point feature extraction model uses the SuperPoint model; the image matching model uses the SuperGlue model.
[0050] Among them, SuperPoint is a neural network for self-supervised extraction of image point features and descriptors. In this embodiment, the feature points of each frame are directly inferred by the SuperPoint model, and the number of feature points per frame is preset when the program starts.
[0051] Step S41, in the process of using the SuperPoint model to extract the point features of the binocular vision image, first, the binocular vision image is compressed by an encoder. The encoder performs three times of convolution, pooling, and non-linear activation on the binocular vision image through a VGG-style convolutional network to achieve the compression of the binocular vision image. The size of the compressed image becomes 1 / 8 * 1 / 8 of the original image. The encoded tensor will simultaneously detect the positions of the feature points and output the descriptors.
[0052] The detection output dimension of the point features is , where 64 channels correspond to non-overlapping pixel grids in the original image. The extra one-dimensional is called the dustbin dimension, which is used to suppress the large values generated by softmax when there are no feature points in the region. After softmax, it can be reshaped to obtain a tensor with the same size as the binocular vision image.
[0053] For the point feature information output by SuperPoint, select the point feature information with high confidence, and through grid allocation, evenly distribute the feature points on the binocular vision image.
[0054] Step S42, further, for the descriptors obtained by the SuperPoint model, use the attention mechanism and the graph neural network to optimize the feature point matching, and filter the incorrect matching pairs through the dustbin mechanism.
[0055] SuperGlue is a neural network that combines query feature point pairs and excludes mismatched feature points. Cooperating with the output of the SuperPoint model, it can obtain better extraction and matching of feature points in different frames.
[0056] Specifically, based on the descriptors obtained by the image point feature extraction model, use the attention mechanism and the graph neural network to optimize the feature point matching, and filter the incorrect matching pairs through the dustbin mechanism.
[0057] For pictures A and B, that is, the left and right pictures of the binocular vision image, each has a set of feature points p and corresponding descriptors d. Define (p, d) as a set of local features, where the descriptor d is generated by the SuperPoint model. Then the local feature set of the binocular vision image can be expressed as , the local feature set of the binocular vision image can be expressed as .
[0058] Using an attention graph neural network to simulate the behavior of humans observing feature point matching, binding the feature point positions and description information, and initially representing each merged feature point as ,
[0059] where MLP represents a multi-layer perceptron, and are the position and descriptor of the -th feature point, respectively. The low-dimensional feature points are upsampled, and the description of the point and the position of the point are coupled; Using a graph neural network, the nodes on the graph represent each feature point in the image. There are two types of edges in the graph. The intra-image edge is represented as (connecting the feature points of its own graph), and the inter-image edge is represented as (connecting the feature points of this graph and another image). Define as the expression of the -th layer of the -th element in image A. The information represents the result of aggregating all feature points. Define
[0060] where is the residual information of the -th layer of the -th element in image A after feature transfer update; Using the attention mechanism to perform aggregation and calculate the message . The intra-image edge is based on the intra-image edge, and the inter-image edge is based on the inter-image edge. Similar to database retrieval, that is, query and retrieve the value of certain elements according to the attributes of certain elements (key ).
[0061]
[0062] where the attention weight , is the representation of all feature points, and . The key , query and value are calculated as linear projections of the deep features of the graph neural network. Considering that the query feature point i is in image Q, and all source feature points are on image S, where , then the key , query and value are represented as:
[0063] Each layer has its own projection parameters, and the projection parameters can be learned and shared for all feature points of two images.
[0064] After L times of calculation of the inner and outer edges of the image, the output of the attention GNN is obtained. The representation of the final matching descriptor for image A is:
[0065] Secondly, SuperGlue constructs an optimal matching layer, which generates an assignment matrix P. It can be achieved by calculating a score matrix That is, by maximizing the score the P matrix can be obtained. The pairwise scores are represented as the similarity of the matching descriptors:
[0066] where is the inner product operation.
[0067] To enable the network to suppress certain key points, the model expands each group of feature points with a dustbin to explicitly assign unmatched feature points to it. This technique is common in graph matching. SuperPoint also uses a dustbin to account for image units that may not be detected. SuperGlue also sets the last row and last column of the score matrix to the dustbin to obtain which is used to filter out wrongly matched feature point pairs, as shown in the following formula:
[0068] After obtaining the matching relationship of the point features, for the point and line features on the same image, determine which point features are on the line feature or the position of the point is less than the set matching threshold from the position of the line feature, then define that the point feature belongs to the line feature. According to the matching relationship of the point features, the matching relationship of the line features is further obtained.
[0069] Specifically, based on the point features and the matching relationship of the point features, traverse each line feature and calculate whether the point features in the left image are on the line feature. Specifically, the following requirements need to be met: a. The projection of the point feature on the coordinate axis should fall within the projection of the endpoints of the line feature on the coordinate axis; b. The distance from the point to the line where the line segment is located should not be greater than the set distance threshold.
[0070] The logical schematic diagram is as shown in Figure 7As shown, the red dots in the figure are the non - satisfying points: From this, the attribution relationships of point features and line features on the left and right figures can be obtained. Then, according to the matching relationships of point features, the matching relationships of line features can be obtained.
[0071] Step S5: Based on the key - frame selection strategy, frames with position, angle changes, and low feature - matching numbers are selected for back - end optimization, and the final positioning trajectory is output; the back - end optimization includes optimizing the inter - frame pose by constructing a reprojection residual model of point and line features. Among them, the process of optimizing the pose by constructing a reprojection residual model of points includes: Binocular stereo vision restores the depth information of feature points; The 3D points in the world coordinate system are converted into the camera coordinate system and normalized to obtain the pixel coordinates after the reprojection of point features, which is expressed as follows:
[0072] Among them, is the reprojection error of the feature point; is the observed pixel coordinate of this feature point in the k - th key frame; is the position of this feature point in the world coordinate system; and are the transformation matrices from the camera coordinate system to the pixel coordinate system and from the world coordinate system to the camera coordinate system respectively.
[0073] Furthermore, in the process of optimizing the inter - frame pose by constructing a reprojection residual model of line features, the Plücker coordinate system is introduced to construct a reprojection residual model of line features for pose optimization. The process includes: After the extraction of image line features, 2D line - segment information is obtained. Since it is a binocular SLAM system, the depth of the line - segment endpoints is calculated to obtain the information of the 3D line segment in space, that is, the 3D line segment is represented by 3D points in two camera coordinate systems. Using the Cartesian coordinate system to describe a straight line in space has redundant degrees of freedom, so the Plücker coordinate ( ) is used for expression, that is, the spatial straight - line equation is:
[0074] Among them, is the representation of the straight line in the Plücker coordinate system, is the direction vector of this straight line, is the cross - product of the vector formed by the two 3D endpoints of the line segment and the origin.
[0075] The mathematical relationship for transforming a straight line from the world coordinate system to the camera coordinate system is:
[0076] Among them, is the Plücker coordinate representation of the straight line in the camera coordinate system and the world coordinate system, is the rotation matrix and transformation matrix from the world coordinate system to the camera coordinate system, is the skew-symmetric matrix representation of a three-dimensional vector.
[0077] The spatial straight line in the camera coordinate system The transformation formula for reprojection onto the image plane is:
[0078] where is the representation of the spatial straight line in the two-dimensional plane, is its coefficient, is the camera internal parameter, is the projection matrix from the camera coordinate system to the image coordinate system.
[0079] Through the above transformation, the spatial straight line represented by Plücker coordinates can be reprojected from the world coordinate system to the image plane. The reprojection error of the line feature is represented by the perpendicular distance from the endpoint coordinates of the line segment to the straight line where the line segment is located, that is:
[0080] where is the reprojection error of the straight line, is the calculation method of the distance from a point to a straight line, are the coordinates of the two endpoints of the plane line segment.
[0081] Furthermore, by selecting the OIVIO dataset, the OIVIO dataset is an open-source dataset for visual SLAM, and the acquisition environment is low-light underground tunnels, pipelines, etc., and the trajectory ground truth is provided for comparison, as Figure 8 shown. Furthermore, the RMSE results of the absolute pose error between the solution results under the partial interception method and the trajectory ground truth are shown in Table 1: Table 1 RMSE results of absolute pose error
[0082] It can be seen from Table 1 that the RMSE results of the absolute pose error of the method provided by the present invention are significantly better than the calculation results of other methods, and have better performance in low-light and low-texture environments. It further shows that in a low-light environment, simultaneously using point and line features can make the front end of visual SLAM more robust.
[0083] Embodiment 2 This embodiment discloses a binocular vision positioning system under low-light environment enhanced by deep learning; such as Figure 9As shown in the figure, a binocular vision positioning system under low-light environment enhanced by deep learning includes: An image acquisition module, configured to: acquire binocular vision images; An edge feature enhancement module, configured to: perform edge feature enhancement on the binocular vision images using an edge detection model; A line feature extraction module, configured to: use a line detection algorithm to extract line features from the binocular vision images with enhanced edge features, and adopt a light ray erasure logic to eliminate false line segments generated by point light source scattering; A point feature extraction and matching module, configured to: use an image point feature extraction model to extract point features of the binocular vision images; based on the acquired point features, adopt a SuperGlue model for feature matching; judge the positional relationship between the point features and the line features, and use the relationship of point feature matching to obtain the relationship of line feature matching; A backend optimization module, configured to: based on a key frame selection strategy, screen frames with position, angle changes and low feature matching numbers for backend optimization, and output the final positioning trajectory; the backend optimization includes optimizing the inter-frame pose by constructing a reprojection residual model of point and line features. Embodiment III The purpose of this embodiment is to provide a computer-readable storage medium.
[0084] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the binocular vision positioning method under low-light environment enhanced by deep learning as described in Embodiment 1.
[0085] Embodiment IV The purpose of this embodiment is to provide an electronic device.
[0086] An electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the binocular vision positioning method under low-light environment enhanced by deep learning as described in Embodiment 1.
[0087] The steps involved in the devices in Embodiments II, III, and IV above correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0088] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0089] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A binocular vision positioning method in low-light environments enhanced by deep learning, characterized in that Including: Obtain binocular vision images; Adopt an edge detection model to enhance the edge features of the binocular vision images; Use a line detection algorithm to extract line features from the binocular vision images with enhanced edge features, and adopt a light erasure logic to eliminate the false line segments generated by the scattering of point light sources; Use an image point feature extraction model to extract the point features of the binocular vision images; based on the obtained point features, adopt a SuperGlue model for feature matching; judge the positional relationship between the point features and the line features, and use the relationship of point feature matching to obtain the relationship of line feature matching; Based on a key frame selection strategy, screen the frames with position, angle changes and low feature matching numbers for back-end optimization, and output the final positioning trajectory; the back-end optimization includes optimizing the inter-frame pose by constructing a reprojection residual model of point and line features.
2. The binocular vision positioning method in low-light environment enhanced by deep learning as claimed in claim 1, wherein, The adopted edge detection model is based on an improved VGG16 network structure, the fully connected layer and the fifth pooling layer are deleted, and the edge detection ability is enhanced through multi-scale feature fusion.
3. The binocular vision positioning method in low-light environment enhanced by deep learning as claimed in claim 1, wherein Using a line detection algorithm to extract line features from the binocular vision images with enhanced edge features, and adopting a light erasure logic to eliminate the false line segments generated by the scattering of point light sources, including: Extract the line features of the binocular vision images with enhanced edge features and the geometric center coordinates of the light source in the image; Extract the pixel values corresponding to the endpoints of all lines, perform contrast stretching on the line segment endpoints, and set the astigmatism threshold; Calculate the gray values of the endpoint pixel values after contrast stretching, further exclude the line segments whose starting point gray values are greater than the astigmatism threshold, and calculate the perpendicular distance from the light source center coordinates to each line, and judge whether it is caused by astigmatism through the distance threshold.
4. The binocular vision positioning method in low light environment based on deep learning enhancement according to claim 1, characterized in that, Using an image point feature extraction model to extract the point features of the binocular vision images, including: The image point feature extraction model compresses the binocular vision images through an encoder, detects the positions of feature points of the compressed images, and outputs descriptors; The results of feature point position detection are processed through a softmax function and a reshape operation to obtain a tensor with the same size as the binocular vision images; Select high-confidence point feature information, and through grid allocation, evenly distribute the point features on the binocular vision images.
5. The binocular vision positioning method in low-light environment enhanced by deep learning as claimed in claim 1, wherein, Based on the obtained point features, adopt a SuperGlue model for feature matching; judge the positional relationship between the point features and the line features, and use the relationship of point feature matching to obtain the relationship of line feature matching; including: Based on the descriptors obtained by the image point feature extraction model, adopt an attention mechanism and a graph neural network to optimize feature point matching, and filter out incorrect matching pairs through a dustbin mechanism; After obtaining the matching relationship of the point features, for the point and line features on the same image, judge which point features are on the line features or the position of the points is less than the set matching threshold from the position of the line features, then define that the point features belong to the line features, and based on the matching relationship of the point features, further obtain the matching relationship of the line features.
6. The binocular vision positioning method in low-light environment enhanced by deep learning as claimed in claim 1, wherein, The process of optimizing the image pose by constructing a reprojection residual model of point features includes: After extracting and matching the point features, restore the depth information of the image points through binocular stereo vision; Transform the 3D points in the world coordinate system to the camera coordinate system; The pixel error formed by the image position of the 3D points after transformation and the extracted feature points constitutes the reprojection error of the point feature.
7. The binocular vision positioning method in low-light environment enhanced by deep learning as claimed in claim 1, wherein The process of optimizing the inter-frame pose by constructing the reprojection residual model of the line feature includes: Calculate the depth of the endpoints of the line segment for the extracted line feature to obtain the information of the 3D line segment in space; Use Plücker coordinates to represent the 3D line segment in space to obtain the line segment information in the Plücker coordinate system; Transform the line segment information in the Plücker coordinate system to the camera coordinate system; Reproject the space line in the camera coordinate system onto the image plane, and use the perpendicular distance from the endpoint coordinates of the line segment to the line where the line segment is located to represent the reprojection error of the line.
8. A binocular vision positioning system under low-light environments enhanced by deep learning, characterized in that: Including: An image acquisition module, configured to: acquire binocular vision images; An edge feature enhancement module, configured to: enhance the edge features of the binocular vision images using an edge detection model; A line feature extraction module, configured to: use a line detection algorithm to extract line features from the binocular vision images with enhanced edge features, and use a ray erasure logic to eliminate the false line segments generated by the scattering of point light sources; A point feature extraction and point-line feature matching module, configured to: use an image point feature extraction model to extract the point features of the binocular vision images; based on the acquired point features, use the SuperGlue model for feature matching; judge the positional relationship between the point features and the line features, and use the relationship of point feature matching to obtain the relationship of line feature matching; A backend optimization module, configured to: based on a key frame selection strategy, screen the frames with position, angle changes, and low feature matching numbers for backend optimization, and output the final positioning trajectory; the backend optimization includes optimizing the inter-frame pose by constructing the reprojection residual models of point and line features.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the binocular vision positioning method under low-light environment enhanced by deep learning according to any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the binocular vision positioning method under low-light environment enhanced by deep learning according to any one of claims 1-7.
Citation Information
Patent Citations
Binocular vision odometer design method based on optical flow tracking and dot-line feature matching
CN112115980A
Visual SLAM system and method based on panoramic annular lens
CN113705369A
Noise image edge detection method based on deep learning
CN117218147A
Orchard robot binocular vision inertial positioning method, device, equipment and medium
CN117870661A
Monocular vision inertial pose estimation method based on deep learning point and line features
CN118537393A