Localization and mapping method for substation inspection robot based on point-line feature fusion

By integrating point and line features and IMU information into the substation inspection robot, and optimizing map initialization and loop closure detection, the problem of low positioning and mapping accuracy in complex substation scenarios is solved, and higher accuracy and robust navigation map construction are achieved.

CN119444849BActive Publication Date: 2025-10-28HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411446570.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-10-28
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

Existing substation inspection robots suffer from low accuracy and poor robustness in localization and mapping technologies in complex scenarios. In particular, visual SLAM algorithms produce sparse maps with low accuracy in complex outdoor large-scale substation scenarios, have poor system initialization robustness, and are affected by changes in lighting, which impacts the quality of image feature extraction. Inaccurate posture of the inertial measurement unit leads to error accumulation.

Method used

A point-line feature fusion method is adopted, which extracts point-line features through a binocular camera, optimizes the EDLines algorithm using an adaptive criterion to obtain high-quality long-line features, optimizes the initial pose by combining IMU information, trains a point-line feature dictionary and a device target detection model for substation scenarios, performs closed-loop detection and correction, and optimizes map initialization and relocalization.

Benefits of technology

The positioning accuracy and robustness of the substation inspection robot are improved, a map with more spatial semantics is quickly constructed, error accumulation is reduced, and the accuracy of the map and the reliability of navigation are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119444849B_ABST
    Figure CN119444849B_ABST
Patent Text Reader

Abstract

This invention discloses a method for localization and mapping of substation inspection robots based on point-line feature fusion. The invention uses a binocular camera to acquire point-line features in image frames, and uses an adaptively optimized EDLines algorithm to quickly obtain high-quality long-line features. It obtains the spatial depth of point-line features through binocular camera triangulation, and filters redundant line features based on a length response threshold strategy for substation spatial depth. Based on a point-line polar feature polar plane constraint method, it combines IMU information and visual information to independently estimate and optimize gyroscope bias, and updates IMU pre-integration information to obtain a more accurate initial pose for the inspection robot. A point-line feature dictionary based on the substation scene and a target detection model for large equipment in the substation are created. Loop closure detection is performed using substation equipment landmarks and point-line features. Rapid relocalization and correction are performed on the current keyframe and the sliding window with a fixed number of previous keyframes. Finally, global BA is used to optimize the substation inspection robot pose and the substation scene map, solving the problem of large-scale scene error drift accumulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of SLAM in complex scenarios that combine visual and inertial fusion based on point and line features, and to a method for substation inspection robots to perform stable mapping and accurate localization in complex substation environments. Background Technology

[0002] With the continuous and rapid development of the economy and society, my country's reliance on power equipment is increasing. Substations, as a crucial component of modern power systems and future smart grids, are responsible for voltage regulation and power distribution within the power network. Their operational stability and reliability directly determine the stability of the power supply system. Therefore, it is necessary to regularly inspect equipment within substations and eliminate potential safety hazards as early as possible to ensure the safe and stable operation of the power network.

[0003] Traditional inspection methods typically require manual intervention. On one hand, substations are large and have diverse equipment, making manual inspection time-consuming and labor-intensive. On the other hand, manual live-line work poses certain safety risks. Therefore, introducing robotics technology to achieve automated substation inspection has become a trend. A substation inspection robot is a complex system integrating a mobile platform, multiple sensors, a positioning and navigation system, a data processing and decision-making system, and an autonomous control system to achieve precise, efficient, and safe inspection operations, effectively avoiding a series of problems faced by manual inspection.

[0004] Trackless navigation technology is a prerequisite for autonomous positioning and real-time equipment status monitoring. Simultaneous Localization and Mapping (SLAM), as a core technology for autonomous robot navigation, has gradually become a research hotspot for trackless inspection robots in substations. Currently, SLAM research mainly focuses on laser SLAM, visual SLAM, and multi-sensor fusion SLAM. Laser SLAM solutions can be divided into single-line LiDAR or multi-line LiDAR. Single-line LiDAR, due to its 2D dimension limitation, lacks information about most of the scene. Multi-line LiDAR compensates for the dimensionality deficiency of single-line LiDAR, but its large point cloud load and high cost limit its application in substations. Visual SLAM utilizes image information captured by cameras to estimate the robot's position and posture in real time, while simultaneously constructing a 3D map of the environment. It has superior scene recognition capabilities and can obtain rich external texture information. However, existing visual algorithms, such as the ORB-SLAM series and Vins-Fusion, which are based on point features, have certain spatial limitations and are far from adequate for describing the complex structures of substations. Furthermore, perceptual aliasing can affect the mapping and localization of similar scenes within the substation, resulting in low accuracy of sparse maps constructed in complex outdoor large-scale substation scenarios and poor system initialization robustness. Algorithms based on line features, such as PL-SLAM and PL-VIO, suffer from excessively long detection and matching times for line features, weak geometric constraints, and low efficiency in point-line feature fusion, severely impacting the efficiency of substation inspections and map construction accuracy. In addition, outdoor lighting conditions can vary significantly, affecting the quality of camera perception and image feature extraction. Smooth surfaces and metal structures in man-made structures in outdoor substation environments can introduce unreliable texture interference. While the Inertial Measurement Unit (IMU) provides high-frequency updates of equipment motion states, compensating for the difficulty of extracting reliable feature points during rapid camera movement, its inaccurate orientation means that the effects of gravity cannot be removed, and the continuous accumulation of errors affects global map construction. Summary of the Invention

[0005] To address the above challenges and problems, this invention provides a localization and mapping method for substation inspection robots based on point-line feature fusion, aiming to improve the localization accuracy and robustness of substation inspection robots in large-scale substation scenarios. This invention uses a binocular camera to acquire point-line features in image frames, and uses an adaptively optimized EDLines algorithm to quickly obtain high-quality long-line features. It obtains the spatial depth of point-line features through binocular camera triangulation, and filters redundant line features based on a length response threshold strategy for substation spatial depth. Based on a point-line polar feature polar plane constraint method, it combines IMU information and visual information to independently estimate and optimize gyroscope bias, and updates IMU pre-integration information to obtain a more accurate initial pose for the inspection robot. A point-line feature dictionary based on the substation scenario and a target detection model for large equipment in the substation are created. Loop closure detection is performed using substation equipment landmarks and point-line features. Rapid relocalization correction is performed on the current keyframe and the sliding window with a fixed number of previous keyframes. Finally, global BA is used to optimize the substation inspection robot pose and the substation scene map, solving the problem of error drift accumulation in large-scale scenes. The key point of this invention is to improve the line feature extraction method, which can more quickly construct a more spatially semantic map by using substation equipment landmarks and point and line features. It also improves the visual-inertial joint initial optimization method and closed-loop correction method to compensate for and optimize the inaccuracy of mapping and unstable positioning in complex substation scenarios.

[0006] A method for localization and mapping of substation inspection robots based on point-line feature fusion includes the following steps:

[0007] Step 1: Acquire binocular image data, extract point features and line features respectively, and perform matching.

[0008] Step 2: Calculate the initial posture of the substation inspection robot under pure vision based on the point and line features; decouple and optimize the robot's posture by projecting the point and line features onto the polar plane to obtain a more robust and accurate map initialization result.

[0009] Step 3: Train a closed-loop word dictionary and a target detection model for medium and large equipment in the substation scenario, based on point and line features. During the substation inspection robot's inspection process, semantic information and point and line feature descriptors of the equipment in the scene are extracted for closed-loop detection. The weighted similarity score of the point and line features is optimized, and the keyframe with the highest similarity score in the keyframe database is found as the optimal closed-loop keyframe, thus completing the closed-loop detection.

[0010] Step 4: When a corresponding optimal closed-loop keyframe is detected in the current keyframe, the relative pose relationship between the current keyframe and the optimal closed-loop keyframe is calculated through the optimized fast relocation. The closed-loop correction is performed through the nonlinear global optimization in the backend, and the poses of all inspection robots in the current map and the complex outdoor 3D map of the substation are updated and optimized.

[0011] The beneficial effects of this invention are as follows:

[0012] The present invention proposes a substation inspection robot localization and mapping method based on point-line feature fusion for complex substation operation and maintenance scenarios. This method combines the advantages of point and line features to effectively acquire more structural information about the substation scene; it accelerates line feature extraction, removes broken line features and redundant short line features, ensuring map accuracy; and it optimizes the map initialization method to improve system robustness. Facing the cumulative errors caused by large-scale outdoor scenes, the closed-loop detection and correction method using point-line features and substation equipment landmarks significantly reduces errors. This solution can obtain a more structured and stable navigation map, making navigation in substation scenarios faster and more reliable, providing strong support for operation and maintenance personnel.

[0013] Innovation Point 1: Line features are incorporated into visual information, and the EDLines straight line detection algorithm is improved through an adaptive merging criterion. Redundant line features are removed based on a substation spatial depth response threshold strategy. In the complex operation and maintenance environment of substations, line features can effectively convey structural information compared to point features. Based on the EDLines algorithm, its relevant parameters are improved, and broken line features are merged according to an adaptive merging criterion. Based on the substation spatial depth obtained by binocular cameras through left and right eye triangulation, a length response threshold strategy for substation spatial depth is proposed to filter out redundant line features and reduce the computational burden of line feature optimization in the backend.

[0014] Innovation Point 2: An polar plane constraint method based on point and line features optimizes the map initialization method, resulting in a more accurate initial pose for the inspection robot. By combining visual and IMU information, constraints are added on the projection of point and line features onto the polar plane. The gyroscope bias is independently estimated and optimized. The IMU pre-integration information is updated, and the decoupled optimization yields a precise initial pose for the inspection robot. This optimizes map initialization in complex outdoor substation maintenance environments, improving its robustness.

[0015] Innovation Point 3: Spatial constraints are added through a point-line feature dictionary and substation equipment landmarks to optimize loop closure detection and correction methods. Based on a substation image dataset, point-line feature descriptors are used to generate a point-line feature dictionary, which is then used to train a target detection model for large equipment in substations. Loop closure detection is performed on the map based on the point-line feature dictionary and substation equipment landmarks, and a more suitable point-line weighted similarity scoring method is used to evaluate the optimal loop closure keyframe. The residuals of substation equipment landmarks are added to the residual optimization function constructed by the sliding window, and local map optimization is performed through fast relocalization. Attached Figure Description

[0016] Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0018] The present invention provides a method for localization and mapping of substation inspection robots based on point-line feature fusion, such as... Figure 1 As shown, it includes the following steps:

[0019] Step 1: During the inspection, the substation inspection robot uses binocular cameras to acquire image frames from both the left and right eyes in real time as input for the vision thread. The vision thread preprocesses the acquired image frames, extracting point features and line features respectively and performing matching. The image frame is then stored in the image frame database. Simultaneously, the first image frame is set as the initial keyframe, and every 1 second, or when the number of matched point-line feature pairs in a new image frame is less than 1 / 4, a new image frame is set as the keyframe and stored in the keyframe database. Specifically, an adaptive merging principle is used to optimize the EDLines algorithm for line feature extraction, while a length response threshold strategy based on spatial depth is used to filter out redundant short line features. The specific operations of the vision thread are as follows:

[0020] Step 1.1: The left and right eye image frames acquired by the stereo camera are processed for grayscale conversion and histogram equalization. The Shi-Tomasi algorithm is used to obtain the point features of each image frame, and a quadtree is used to achieve feature point homogenization and remove redundant point features. Based on the known parallax between the left and right eyes of the stereo camera, the point features of the left and right eye image frames are merged through triangulation to obtain spatial point features with absolute scale. BRIEF descriptors are used to describe the pixel information surrounding the point features. A coarse similarity matching is performed by comparing the Hamming distance between BRIEF descriptors in adjacent left eye image frames, and the Progressive Sample Consensus (PROSAC) algorithm of RANSAC is used to eliminate incorrectly matched point feature pairs.

[0021] Step 1.2: For the left-eye image frame acquired by the stereo camera, line features are extracted using the optimized EDLines algorithm, and the LBD descriptor is used to describe the local appearance of the line features. KNN is then used to match the line features. Specifically, the steps are as follows:

[0022] Step 1.2.1: Image preprocessing: The left eye image frame acquired by the binocular camera is converted to grayscale. Then, bilateral filtering is performed on the grayscale image to reduce noise and obtain the preprocessed image. Compared with Gaussian filtering, bilateral filtering uses pixel similarity weighting, which removes noise while preserving the edge details of the image, thus preparing for the subsequent line feature extraction.

[0023] Step 1.2.2: Extract line features based on the optimized EDLines algorithm:

[0024] The image gradient function is obtained by taking the gray value of the preprocessed image. Since the image is equivalent to a two-dimensional discrete function, the gray approximation of the image gradient function is calculated using the Sobel discrete difference operator based on the difference method, so as to obtain the gradient value and gradient direction at the pixel (x,y).

[0025] The gradient map is obtained by calculating the image gradient of the preprocessed image using the image gradient function. The entire gradient map is then traversed to select pixels whose gradient values ​​are greater than a gradient threshold compared to their neighboring pixels. This identifies image regions with more edge details. The preferred gradient threshold is TH. gradient =5.22. For the obtained new image region, pixels whose gradient values ​​are greater than the anchor threshold relative to the gradient values ​​of their neighboring pixels are selected as anchor points along the gradient direction. The preferred anchor threshold is TH. anchor =3. For anchor points, based on the surrounding pixels of the current anchor point, find pixels (which may be anchor points) with the same gradient direction as the current anchor point as extension pixels. If the gradient direction of the next pixel in the gradient direction of the current anchor point is inconsistent with the gradient direction of the current anchor point, then detect the two adjacent pixels of the next pixel. Finally, when all extension directions have been extracted, it means that the connection of the current anchor point is over. For the remaining unconnected anchor points, repeat the extension operation to connect the anchor points until all anchor points are connected into line edges. The above pixel extension method reduces the search amount required for anchor point extension and speeds up the extraction of image line edges.

[0026] For the extracted image line edges, based on the minimum straight line length and the maximum mean square line fitting error, the number of false alarms (NFA) is used to fit the image line edges into suitable line segments and retain them in the line segment set.

[0027] After generating a set of line segments containing all line segments of the image, sort them in descending order of line segment length, and then filter the line segments based on two aspects: angle and spatial distance. First, calculate the angle between the current line segment and the direction vector of the remaining line segments in the set, and filter out line segments with excessively large angles:

[0028]

[0029] Where l1 and l2 represent the two line segments being compared, v1 and v2 represent the direction vectors of l1 and l2, and θ min This represents the filtering threshold between the feature angles of the comparison lines; a value of 0.2 rad is preferred here.

[0030] Next, segments with excessively large horizontal and vertical distances are filtered out. If a pair of endpoints between the current segment and another segment satisfies the following condition, then the other segment is considered a pre-merged segment of the current segment. Compared to calculating Euclidean distance, this filtering method saves computational overhead:

[0031] |x 1i -x 2j | <d min

[0032] |y 1i -y 2j | <d min

[0033] i,j∈[1,2]

[0034] Where, x 1i and x 2j The horizontal positions of the two endpoints of line segment 1 and line segment 2 are given by y. 1i and y 2j Let d represent the perpendicular positions of the two endpoints of line segment 1 and line segment 2, respectively. min To compare the filtering threshold of the distance between the endpoints of the line features, a value of 30 is preferred here.

[0035] After traversing and filtering the line segment set using the above method, only line segments that meet the adaptive merging criterion can be merged with their corresponding line segments. Firstly, the Euclidean distance between the endpoint pairs of the pre-merged line segment and its corresponding line segment must be less than d. min Secondly, the angle threshold is adaptively adjusted according to the length of the pre-merged segment, so that longer pre-merged segments have a more stringent angle threshold. A sigmoid function for the angle threshold based on the adaptive length is established, and the specific calculation method is as follows:

[0036]

[0037] Where k and a are the curve steepness and translation position of the sigmoid function, respectively. Here, k is preferably 2.5 and a is 2, so that the adaptive angle threshold has a more stringent restriction on the merging of pre-merged long line segments; n is the normalized length of the pre-merged line segment.

[0038] Once a pre-merged segment meets the above adaptive merging criteria, based on the horizontal direction of the current segment, the farthest endpoints of the current segment and the pre-merged segment are found and used as the endpoints of the merged segment. If both farthest endpoints are endpoints of the current segment, the pre-merged segment is ignored. After selecting new endpoints, the two endpoints are connected to merge the two lines into a new segment, replacing the original corresponding segment, and the pre-merged segment is removed from the segment set. The above operation is repeated sequentially by length for the segments in the segment set until there are no more pre-merged segments to merge in the segment set.

[0039] Step 1.2.3: Based on the spatial depth of the substation, a length response threshold strategy is adopted to filter out redundant short-line features:

[0040] Considering the varying depths of line features in large outdoor substation scenes, and the presence of short line features when the EDLines algorithm extracts image frames from these scenes, a length response threshold strategy based on spatial depth is used to ensure the reliability of line features and reduce system computational costs. A threshold L is set based on the spatial depth of the line features and the image size. min The formula for filtering out redundant short-line features is as follows:

[0041]

[0042] Wherein, min(W) I H I ) represents the smaller of the width and height of the input image; z is the spatial depth of the line feature; b is the left and right eye baselines of the stereo camera; η is the ratio factor, which is preferably set to 0.125 here.

[0043] Step 1.2.4: Based on the LBD descriptor, use KNN to match line features:

[0044] Based on the selected line features, the support regions of each line feature are segmented into pixel strips, and strip descriptors are calculated. The LBD descriptor is constructed using the average vector and standard deviation vector of the strip descriptors. By calculating the Hamming distance between the LBD descriptors of pairwise line features in adjacent frames, a kd-tree is built, and the optimal line feature pair is obtained through nearest neighbor search.

[0045] Step 2: Based on the point and line feature pairs matched by the left and right eyes in the initial keyframe of Step 1, the initial posture of the substation inspection robot in pure vision mode is calculated using the essential matrix. Depending on the required number of point and line features, consecutive keyframes are selected from the initial keyframe to form a continuous keyframe set. Based on the direction vectors of the point and line feature pairs on the left and right eye polar planes of adjacent keyframes and the first-order Taylor expansion of the IMU rotation part pre-integrated, a minimum eigenvalue optimization problem of the point and line features on the polar plane is constructed by combining visual and inertial information. This yields the initial gyroscope bias. Prior and IMU residual functions are established to obtain the optimal estimates of the keyframe velocity, gravity direction, and IMU bias. Finally, the robot's own posture is decoupled and optimized sequentially to obtain a more robust and accurate map initialization result.

[0046] Step 2.1: Obtain the initial pose of the substation inspection robot under purely visual conditions:

[0047] Based on the successfully matched point feature pairs obtained from the initial keyframe in step 1, the image coordinates of the successfully matched point feature pairs are first converted into normalized coordinates. Then, 8 pairs of point features are randomly selected to construct a system of linear equations to estimate the essential matrix. Finally, the rotation matrix and translation vector of the camera are recovered from the singular value matrix obtained by SVD decomposition, which is the initial posture of the substation inspection robot in pure vision.

[0048] Step 2.2: IMU preprocessing and gyroscope zero-bias construction:

[0049] Since the state propagation of the IMU model is performed through pre-integration, but the computational cost of pre-integration is too high, a partial derivative of a first-order Taylor expansion is used to adjust the pre-integration. Specifically, the pre-integration for the rotation portion between frames k and k+1 is... and its first-order partial derivative for:

[0050]

[0051] Among them, b k and b k+1 These are the k-th and (k+1)-th frames of the IMU system, respectively. k and t k+1 Ω(ω) represents the timestamps of the k-th and (k+1)-th frames, respectively, and Ω(ω) is the angular velocity matrix. Let ω be the IMU angular velocity measurement value at time t. x ω y ω z These are the x, y, and z components of the IMU angular velocity vector, respectively. Let be the rotational change from time t to the k-th frame of the IMU system. and These are the gyroscope zero bias and the increment of the gyroscope zero bias in the k-th frame of the IMU, respectively, b ω It is the zero bias of the gyroscope. It is the Jacobian matrix of the pre-integrated rotation term with respect to the zero bias of the gyroscope.

[0052] For a point feature in 3D space, connecting it to the left and right eye centers of the camera in frames k and (k+1) yields the direction vectors. and For a line feature in three-dimensional space, the direction vector f from the camera optical center to the line feature is obtained by cross product of the Plücker coordinate normal vector n and the line direction vector d. l :

[0053]

[0054] Where n is the normal vector of the polar plane formed by the line feature and the camera optical center, and d is the direction vector of the line feature.

[0055] Connecting the line features to the left and right eye centers of the camera in frames k and k+1 yields the direction vectors. and Each pair of vectors forms a different polar plane with the baseline of adjacent frames. This is achieved through a rotation matrix between the IMU coordinate system and the camera coordinate system. After converting the IMU coordinate system where the gyroscope's zero bias is located to the left and right eye coordinate systems of the camera, a coordinate system based on the direction vector and... Minimize the eigenvalue model to optimize the gyroscope zero bias b ω The formula is expressed as follows:

[0056]

[0057] in, It is a set of consecutive keyframes in the initialization phase containing IMU measurement information. The set of consecutive keyframes is composed of consecutive keyframes selected from the initial keyframe according to the required number of point and line features. L M k,k+l and R M k,k+l It is a positive semi-definite matrix, λ min Let N be the smallest eigenvalue of the positive semi-definite matrix, and let N be the product of the matrix formed by the polar plane constraint relationship of the left and right cameras and its transpose. L and N R M represents the total number of point features matched by the left and right eyes, respectively. L and M R These represent the total number of line features matched by the left and right eyes, respectively. and The left and right eye centers and the feature p of the i-th point are respectively located in the k-th frame. i The direction vector between them and The left and right eye centers and the j-th line feature are respectively located in the k-th frame. i The direction vector between them and These are the coordinate systems of the left-eye camera, c. L And the right eye camera coordinate system c R The transformation matrix between the IMU coordinate system b and the IMU coordinate system b.

[0058] Step 2.3: IMU state variable optimization:

[0059] Step 2.2 yields a preliminary gyroscope bias estimate, which is then used to obtain the optimal estimate of the IMU state variables. That is, keyframe velocity, gravity vector, gyroscope zero bias, and accelerometer zero bias b. a The formula is expressed as follows:

[0060]

[0061] in, It refers to the velocity of each of the N+1 keyframes in the world coordinate system. It is the gravity vector in the camera coordinate system.

[0062] Through IMU pre-integration and The residual function is constructed from the state variables to optimize the state variables, as shown in the following formula:

[0063]

[0064] Where, r p It is the IMU bias prior residual. Σ is the residual of the change in IMU measurement information between frame k and frame (k+1). p It is the covariance of the IMU's zero-biased prior. It is the covariance of the change in IMU measurement information between the k-th frame and the (k+1)-th frame.

[0065] Step 2.4: Gradually optimize the posture of the substation inspection robot through decoupling:

[0066] While obtaining the optimal IMU state variables in step 2.3, the residuals of the pre-integrated rotating part are also used. Obtain the rotation matrix between the optimized world coordinate system and the current IMU frame in the k-th frame. Then update the rotation matrix of each keyframe in the left and right eyes:

[0067]

[0068] in, and Let be the rotation matrices from the k-th frame of the left and right cameras to the world coordinate system, respectively. and These are the fixed rotation matrices from the left and right cameras to the IMU coordinate system, respectively.

[0069] For each keyframe of the image, the point and line features extracted are used to construct an error model for point and line reprojection using the least squares method to optimize the translation vector. The formula is expressed as follows:

[0070]

[0071] Among them, X i It is the i-th 3D point feature in the world coordinate system, x i It is X i Two-dimensional point features projected onto the image plane of the k-th frame in the left eye. m j These are the homogeneous coordinates of the midpoint of the j-th two-dimensional line feature. It is the projected two-dimensional line feature obtained by transforming the j-th 3D line feature in the camera coordinate system to the k-th frame image plane of the left eye. and n represents the coefficients of the projected two-dimensional line characteristics. xand n l π represents the number of point features and line features, respectively; π is the projection function; ρ is the Huber kernel function, which only processes outliers; and Σ is the covariance of the residual function.

[0072] Step 3: Train a closed-loop word dictionary and a target detection model for medium and large equipment in the substation scenario, based on point and line features. During the substation inspection robot's inspection process, semantic information and point and line feature descriptors of the equipment in the scene are extracted for closed-loop detection. The weighted similarity score of the point and line features is optimized, and the keyframe with the highest similarity score in the keyframe database is found as the optimal closed-loop keyframe, thus completing the closed-loop detection.

[0073] Step 3.1: Offline training of the point and line feature dictionary for the substation scenario:

[0074] To obtain a sufficient dataset of substation images, it is necessary to cover images taken under different lighting conditions, substation scenes, weather conditions, and seasons. First, all training images are traversed, and Shi-Tomasi point features and EDLines line features are extracted from each image. The feature information is then quantified using BRIEF descriptors and 256-dimensional binary LBD descriptors, respectively. A dictionary tree for point and line features is built using the DBoW2 library, a third-party library for the bag-of-words model. The number of branches K=10 and the depth L=5 are uniformly set. The extracted point and line feature descriptors are clustered into 10 sets using K-means, i.e., all feature samples are clustered into 10 classes. This clustering is repeated four times to generate a point and line feature dictionary offline. Finally, the last layer node of the dictionary is called a word, and each word is assigned a weight based on its frequency in the training set using TF-IDF (frequency-inverse document frequency).

[0075] Step 3.2: Extract semantic information of substation equipment:

[0076] Images of medium and large-sized equipment were selected from a substation image dataset to form the substation equipment dataset, and the data was labeled. Considering the real-time requirements of the SLAM system, a lightweight YOLOv5s model was chosen as the target detection model for medium and large-sized equipment. Target detection was performed only on keyframes to meet the real-time performance requirements of SLAM. In the large-scale substation scene, to make reasonable use of limited visual information processing resources, a coordinate attention module (CA) was inserted into the convolutional layer of the backbone network of the lightweight YOLOv5s model. This module introduces coordinate information to obtain accurate positioning information of substation equipment, enhancing the target location information capture capability of the medium and large-sized equipment target detection model. After training the substation equipment dataset with the medium and large-sized equipment target detection model, the model was built into the industrial control computer of the inspection robot to obtain semantic information of substation equipment in real time. Due to the possibility of false detection, the equipment was only mapped as a landmark onto a map with point and line features when the medium and large-sized equipment target detection model detected the spatial location information of the equipment in three consecutive keyframes.

[0077] Step 3.3: Obtain the loop closure frame through loop closure detection:

[0078] After map initialization is complete, a loop closure detection thread is initiated to perform loop closure detection on the current keyframes obtained in real time during the inspection. If the current keyframe contains substation equipment landmarks, the most recently mapped landmark is identified and compared with landmarks within keyframes in the keyframe database. The spatial proximity between them is calculated to determine whether the keyframes in the keyframe database are candidate keyframes for loop closure.

[0079]

[0080] Among them l cur This represents the 2D projection position of the landmark in the current keyframe. This refers to the set of landmarks within keyframes in the keyframe database. Landmark Collection The three-dimensional position of the i-th landmark in the array, where ||·||2 is the L2 norm.

[0081] Simultaneously, based on the point and line features in the current keyframe obtained from the visual thread, and using the point and line feature dictionary obtained in step 3.1, the descriptors of the point and line features are compared online with the cluster centers of each intermediate node in the tree structure of the dictionary (4 times) to find the word position where the descriptor is located. Furthermore, by comparing the keyframes in the keyframe database, keyframes with more than 80% common words between two frames are added to the closed-loop candidate keyframe group. The point and line features of the current keyframe are converted into bag-of-words vectors for point and line features, respectively. By comparing the bag-of-words vectors between the current keyframe and all closed-loop candidate keyframes in the closed-loop candidate keyframe group, an image similarity score based on point and line features is obtained. In the similarity score, feature size and distribution degree are used as two criteria to weight the similarity scores of the two features, and the normalized feature size I is calculated. p I l and distribution degree D p D l The formula is expressed as follows:

[0082]

[0083] Where, N p N l and N max These are the number of point features, line features, and all features, respectively, D. p D l and D max These are the distribution values ​​for point features, line features, and all features, respectively.

[0084] Based on distance weighting, dispersion is added as a penalty term to the denominator of the similarity score to reduce the impact of high dispersion on the score. The formula is as follows:

[0085]

[0086] Where w p and w l These are the weighting coefficients for point features and line features, with more weighting for point features than line features. Therefore, w is the preferred weighting factor here. p Take 0.6, w l Take 0.4.

[0087] Based on the weighted similarity scores, the candidate keyframe with the highest similarity score to the current keyframe is selected as the loop closure keyframe. However, this loop closure keyframe undergoes a temporal consistency check: it cannot be a keyframe within the sliding window containing the current keyframe. Furthermore, a spatial coherence check is performed: the similarity score between this loop closure keyframe and all other keyframes within the current sliding window must be greater than a set similarity threshold, which is set to 0.75 times the similarity score between the loop closure keyframe and the current keyframe. Only when both conditions are met can the loop closure keyframe be determined as the optimal loop closure keyframe.

[0088] Step 4: When a corresponding optimal closed-loop keyframe is detected in the current keyframe, the relative pose relationship between the current keyframe and the closed-loop keyframe is calculated through the optimized fast relocation. The closed-loop correction is performed through the nonlinear global optimization in the backend, and the poses of all inspection robots in the current map and the complex outdoor 3D map of the substation are updated and optimized.

[0089] Step 4.1: Calculate the relative pose between the current frame and the closed-loop frame using the optimized fast relocalization method:

[0090] Once the optimal closed-loop keyframe is detected, a coarse matching based on descriptors is first performed using the point and line features of the optimal closed-loop keyframe and the current keyframe. The PROSAC algorithm is then used to calculate the homography matrix for high-quality point and line pairs, and the point and line projection residuals are calculated to determine interior and exterior points, eliminating incorrectly matched point and line features. The pose and point and line features of the optimal closed-loop keyframe are used as closed-loop constraints and added to the overall residual function of the back-end nonlinear optimization. A sliding window is used to optimize the pose of the optimal closed-loop keyframe, thereby calculating the relative pose relationship between the current keyframe and the optimal closed-loop keyframe v. The residual function formula is as follows:

[0091]

[0092] Among them, I prior e imu e point and e line These are the marginalized prior residuals, IMU residuals, point feature residuals, and line feature residuals, respectively. loop It is composed of the residual of the closed-loop term, e ldmk It refers to the landmark residual, which is the spatial proximity of landmarks. It is the set of point and line features in the optimal closed-loop keyframe. and These are the measured values ​​of point and line features, respectively. and These are the rotation quaternion and spatial position of the optimal closed-loop keyframe pose, respectively.

[0093] By using the multi-view constraints of the sliding window, the tightly coupled optimized pose is obtained, the accumulated offset error is calculated, and the pose of the keyframes within the sliding window is corrected. All frames within the sliding window are aligned with the past poses to complete fast relocalization.

[0094] Step 4.2: Global closed-loop optimization:

[0095] When a keyframe matching the optimal closed-loop keyframe slides out of the sliding window, closed-loop optimization is performed on all image frames in the image frame database. Considering that the gravity vector has already been optimized in step 2.3, and the pitch and roll angles in the rotation quaternion can be observed from the gravity vector, the global closed-loop optimization only needs to optimize the position q. w And yaw angle ψ. The global residual objective function of the position and yaw angle of the sequence frame and the optimal closed-loop key frame is constructed by the global BA function, and the Ceres library is called to solve the optimization problem to obtain the optimal attitude of each frame globally, and the spatial position of its 3D points and 3D lines is optimized accordingly, finally obtaining a complex outdoor 3D map of the substation that can be used for stable navigation.

[0096] The above description, in conjunction with specific / preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. Those skilled in the art can make various substitutions or modifications to these described embodiments without departing from the inventive concept, and all such substitutions or modifications should be considered within the scope of protection of the present invention.

[0097] The parts of this invention not described in detail are well-known to those skilled in the art.

Claims

1. A method for localization and mapping of substation inspection robots based on point-line feature fusion, characterized in that, Includes the following steps: Step 1: Acquire binocular image data, extract point features and line features respectively, and perform matching; Step 2: Calculate the initial pose of the substation inspection robot under pure vision conditions based on point and line features; decouple and optimize the robot's pose by using the constraints of point and line feature projection onto the polar plane to obtain more robust and high-precision map initialization results; the specific method is as follows: Based on the point feature pairs and line feature pairs matched by the left and right eyes in the initial keyframe in step 1, the initial posture of the substation inspection robot in pure vision mode is calculated using the essential matrix. Depending on the required number of point and line features, consecutive keyframes are selected from the initial keyframe to form a continuous keyframe set. Based on the direction vectors of the point and line feature pairs on the left and right eye polar planes of adjacent keyframes and the first-order Taylor expansion of the IMU rotation part pre-integrated, a minimum eigenvalue optimization problem for the point and line features on the polar plane is constructed using visual and inertial information. This yields the initial gyroscope zero bias. Prior and IMU residual functions are established to obtain the optimal estimates of the keyframe velocity, gravity direction, and IMU zero bias. Finally, the robot's own posture is decoupled and optimized sequentially to obtain a more robust and accurate map initialization result. Step 3: Train a closed-loop word dictionary and a target detection model for medium and large equipment based on point and line features in the substation scenario; during the inspection process of the substation inspection robot, extract the semantic information of the equipment in the scene and point and line feature descriptors to perform closed-loop detection, optimize the weighted similarity score of point and line features, find the key frame with the highest similarity score in the key frame database as the optimal closed-loop key frame, and complete the closed-loop detection. Step 4: When a corresponding optimal closed-loop keyframe is detected in the current keyframe, the relative pose relationship between the current keyframe and the optimal closed-loop keyframe is calculated through the optimized fast relocation. The closed-loop correction is performed through the nonlinear global optimization in the backend, and the poses of all inspection robots in the current map and the complex outdoor 3D map of the substation are updated and optimized.

2. The substation inspection robot localization and mapping method based on point-line feature fusion according to claim 1, characterized in that, The specific method for step 1 is as follows: During the inspection, the substation inspection robot uses a binocular camera to acquire image frames from the left and right eyes in real time as input to the vision thread. The vision thread preprocesses the acquired image frames, acquires point features and line features respectively, and performs matching, and stores the image frame in the image frame database. Simultaneously, the first image frame is set as the initial keyframe, and every 1 second or when the number of point-line feature pairs matched in a new image frame is less than 1 / 4, a new image frame is set as the keyframe and stored in the keyframe database. Among these, the EDLines algorithm is optimized using an adaptive merging principle for line feature extraction, while redundant short line features are filtered out using a length response threshold strategy based on spatial depth.

3. The substation inspection robot localization and mapping method based on point-line feature fusion according to claim 2, characterized in that, The specific operations of the visual thread are as follows: Step 1.1: For the left and right eye image frames acquired by the stereo camera, perform grayscale conversion and histogram equalization. Use the Shi-Tomasi algorithm to obtain the point features of each image frame, and use a quadtree to achieve feature point homogenization and remove redundant point features. Based on the known parallax of the left and right eyes of the stereo camera, merge the point features of the left and right eye image frames through triangulation to obtain spatial point features with absolute scale. Use the BRIEF descriptor to describe the pixel information around the point features. Perform coarse similarity matching by comparing the Hamming distance between the BRIEF descriptors in the adjacent left eye image frames, and use the RANSAC progressive sample consensus algorithm to remove incorrectly matched point feature pairs. Step 1.2: For the left-eye image frame acquired by the stereo camera, line features are extracted using the optimized EDLines algorithm, and LBD descriptors are used to describe the local appearance of the line features. KNN is used to match the line features; specifically, the following steps are included: Step 1.2.1: Image preprocessing: The left eye image frame acquired by the binocular camera is converted to grayscale, and then bilateral filtering noise reduction is performed on the grayscale image to obtain the preprocessed image; Step 1.2.2: Extract line features based on the optimized EDLines algorithm: The image gradient function is obtained by analyzing the grayscale values ​​of the preprocessed image. Since the image is equivalent to a two-dimensional discrete function, the grayscale approximation of the image gradient function is calculated using the Sobel discrete difference operator based on the finite difference method, thus obtaining the pixel values. The gradient value and gradient direction are specified. After obtaining the image gradient function, a gradient map is obtained by calculating the image gradient of the preprocessed image. The entire gradient map is traversed, and pixels whose gradient values ​​are greater than a gradient threshold compared to their neighboring pixels are selected to obtain image regions with more edge details. For the obtained new image regions, pixels whose gradient values ​​are greater than an anchor threshold compared to their neighboring pixels are selected as anchors in the gradient direction. For anchors, based on the surrounding pixels of the current anchor, pixels with the same gradient direction as the current anchor are found as extension pixels. If the gradient direction of the next pixel in the gradient direction of the current anchor is not consistent with the gradient direction of the current anchor, the two adjacent pixels of the next pixel are checked. Finally, when all extension directions have been extracted, it means that the connection of the current anchor is over. For the remaining unconnected anchors, the extension operation is repeated to connect the anchors until all anchors are connected into a line edge. For the extracted image line edges, based on the minimum straight line length and the maximum mean square line fitting error, the false detection line segment number NFA is used to fit the image line edges into line segments and retain them in the line segment set. After generating a set of line segments containing all line segments in the image, sort them in descending order of line segment length, and then filter the line segments based on both angle and spatial distance. First, calculate the angle between the current line segment and the direction vector of the remaining line segments in the set, and filter out line segments with excessively large angles. in, , The two line segments being compared are... , express , directional vector, This represents the filtering threshold between the feature angles of the comparison lines; Next, filter out line segments with excessively large horizontal and vertical distances. If there exists a pair of endpoints between the current line segment and another line segment that satisfy the following conditions, then the other line segment is a pre-merged line segment of the current line segment: in, and , and Let be the perpendicular positions of the two endpoints of line segment 1 and line segment 2, respectively. Filtering threshold for the distance between the endpoints of the comparison line features After traversing and filtering the line segment set using the above method, the pre-merging line segments can only be merged with the corresponding line segments if they meet the adaptive merging criteria; firstly, the Euclidean distance between the endpoint pairs of the pre-merging line segment and the corresponding line segment is less than 1. Secondly, the angle threshold is adaptively changed according to the length of the pre-merged line segment, and an angle threshold sigmoid function based on the adaptive length is established. The specific calculation method is as follows: Where k and a are the curve steepness and translation position of the sigmoid function, respectively; n is the normalized length of the pre-merged line segment; Once the pre-merged line segments meet the above adaptive merging criteria, based on the horizontal direction of the current line segment, find the farthest two endpoints of the current line segment and the pre-merged line segment, and use them as the endpoints of the merged line segment; if both farthest two endpoints are endpoints of the current line segment, then ignore the pre-merged line segment; after selecting the new endpoints, connect the two endpoints to merge the two lines into a new line segment, which replaces the original corresponding line segment, and delete the pre-merged line segment from the line segment set; repeat the above operation on the line segments in the line segment set according to their length, until there are no pre-merged line segments that can be merged in the line segment set; Step 1.2.3: Based on the spatial depth of the substation, a length response threshold strategy is adopted to filter out redundant short-line features: A length response thresholding strategy based on spatial depth is used to set thresholds according to the spatial depth of the line features and the image size. The formula for filtering out redundant short-line features is as follows: in, represents the smaller of the width and height of the input image; z is the spatial depth of the line feature, and b is the left and right eye baselines of the stereo camera. It is a ratio factor; Step 1.2.4: Based on the LBD descriptor, use KNN to match line features: Based on the selected line features, the support regions of each line feature are segmented into pixel strips, and strip descriptors are calculated. The LBD descriptor is constructed by the average vector and standard deviation vector of the strip descriptors. By calculating the Hamming distance between the LBD descriptors of each pair of line features in adjacent frames, a kd-tree is constructed, and the optimal line feature pair is obtained by nearest neighbor search.

4. The substation inspection robot localization and mapping method based on point-line feature fusion according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: Obtain the initial pose of the substation inspection robot under purely visual conditions: Based on the point feature pairs successfully matched in the initial keyframe obtained in step 1, the image coordinates of the successfully matched point feature pairs are first converted into normalized coordinates. Then, 8 pairs of point features are randomly selected to construct a system of linear equations to estimate the essential matrix. Finally, the rotation matrix and translation vector of the camera are recovered from the singular value matrix obtained by SVD decomposition, which is the initial posture of the substation inspection robot in pure vision. Step 2.2: IMU preprocessing and gyroscope zero-bias construction: Since the state propagation of the IMU model takes the form of pre-integration, the pre-integration is adjusted using the partial derivative of a first-order Taylor expansion; where the pre-integration is the rotation part between frame k and frame k+1. and its first-order partial derivative for: in, and These are the k-th and (k+1)-th frames of the IMU system, and These are the timestamps of the k-th frame and the (k+1)-th frame, respectively. The angular velocity matrix is... Let be the IMU angular velocity measurement value at time t. , , These are the x, y, and z components of the IMU angular velocity vector, respectively. Let be the rotational change from time t to the k-th frame of the IMU system. and These are the gyroscope zero bias and the increment of the gyroscope zero bias in the k-th frame of the IMU, respectively. It is the zero bias of the gyroscope. It is the Jacobian matrix of the pre-integrated rotation term with respect to the zero bias of the gyroscope; For a point feature in 3D space, connecting it to the left and right eye centers of the camera in frames k and (k+1) yields the direction vectors. and For line features in three-dimensional space, the normal vector in Plücker coordinates is used. and the direction vector of the line Cross product yields the direction vector from the camera's optical center to the line feature. : in, Let be the normal vector of the polar plane formed by the line feature and the camera optical center. The characteristic direction vector of the line; Connecting the line features to the left and right eye centers of the camera in frames k and k+1 yields the direction vectors. and Each pair of vectors forms a different polar plane with the baseline of adjacent frames; the rotation matrix between the IMU coordinate system and the camera coordinate system is used to achieve this. After converting the IMU coordinate system where the gyroscope's zero bias is located to the left and right eye coordinate systems of the camera, a coordinate system based on the direction vector and... Minimize the eigenvalue model to optimize the gyroscope zero bias The formula is expressed as follows: in, The initialization phase of MU measurement information consists of a set of consecutive keyframes, selected from the initial keyframe based on the required number of point and line features. and It is a positive semi-definite matrix. Let be the smallest eigenvalue of the positive semi-definite matrix, and be the product of the matrix formed by the polar plane constraint relationship of the left and right cameras and its transpose. and These represent the total number of point features matched by the left and right eyes, respectively. and These represent the total number of line features matched by the left and right eyes, respectively. and These represent the center of gravity of the left and right eyes in the k-th frame and the feature of the i-th point, respectively. The direction vector between them and These represent the center of gravity of the left and right eyes in the k-th frame and the feature of the j-th line. The direction vector between them and These are the coordinate systems of the left-eye camera. And the right eye camera coordinate system With IMU coordinate system The transformation matrix between them; Step 2.3: IMU state variable optimization: Step 2.2 yields a preliminary gyroscope bias estimate, which is then used to obtain the optimal estimate of the IMU state variables. That is, keyframe velocity, gravity vector, gyroscope zero bias, and accelerometer zero bias. The formula is expressed as follows: in, It refers to the velocity of each of the N+1 keyframes in the world coordinate system. It is the gravity vector in the camera coordinate system; Through IMU pre-integration and The residual function is constructed from the state variables to optimize the state variables, as shown in the following formula: in, It is the IMU bias prior residual. It is the residual of the change in IMU measurement information between the k-th frame and the (k+1)-th frame. It is the covariance of the IMU's zero-biased prior. It is the covariance of the change in IMU measurement information between the k-th frame and the (k+1)-th frame; Step 2.4: Gradually optimize the posture of the substation inspection robot through decoupling: While obtaining the optimal IMU state variables in step 2.3, the residuals of the pre-integrated rotating part are also used. This yields the rotation matrix between the optimized world coordinate system and the current IMU frame in the k-th frame. This updates the rotation matrix of each keyframe in the left and right eyes: in, and Let be the rotation matrices from the k-th frame of the left and right cameras to the world coordinate system, respectively. and These are the fixed rotation matrices from the left and right cameras to the IMU coordinate system, respectively; For each keyframe of the image, the point and line features extracted are used to construct an error model for point and line reprojection using the least squares method to optimize the translation vector. The formula is expressed as follows: in, It is the i-th three-dimensional point feature in the world coordinate system. Two-dimensional point features projected onto the image plane of the k-th frame of the left eye. These are the homogeneous coordinates of the midpoint of the j-th two-dimensional line feature. It is the projected two-dimensional line feature obtained by transforming the j-th 3D line feature in the camera coordinate system to the k-th frame image plane of the left eye. , and The coefficients represent the characteristics of the projected two-dimensional line. and These are the number of point features and line features, respectively. It is a projection function. It's the Huber kernel function, which only processes out-of-view points. It is the covariance of the residual function.

5. The substation inspection robot localization and mapping method based on point-line feature fusion according to claim 1, characterized in that, The specific method for step 3 is as follows: Step 3.1: Offline training of the point and line feature dictionary for the substation scenario: To obtain a sufficient dataset of substation images, it is necessary to cover images taken under different lighting conditions, substation scenes, weather conditions, and seasons. First, all training images are traversed, and Shi-Tomasi point features and EDLines line features are extracted from each image. The feature information is then quantified using BRIEF descriptors and 256-dimensional binary LBD descriptors, respectively. A dictionary tree for point and line features is built using the DBoW2 library, a third-party library for bag-of-words models. The number of branches K=10 and the depth L=5 are uniformly set in the dictionary tree. The extracted point and line feature descriptors are clustered into 10 sets using K-means, that is, all feature samples are clustered into 10 classes. This clustering is repeated 4 times to generate a point and line feature dictionary offline. Finally, the last layer node of the dictionary is called a word, and TF-IDF is used to assign a certain weight to each word based on its frequency of occurrence in the training set. Step 3.2: Extract semantic information of substation equipment: Images of medium and large-sized equipment were selected from the substation image dataset as the substation equipment dataset and labeled. Considering the real-time requirements of the SLAM system, a lightweight YOLOv5s model was selected as the target detection model for medium and large-sized equipment. Target detection was performed only on keyframes to meet the real-time performance requirements of SLAM. In the large-scale substation scene, to make reasonable use of limited visual information processing resources, the attention module CA was inserted into the convolutional layer of the backbone network of the lightweight YOLOv5s model. This module introduces coordinate information to obtain accurate positioning information of substation equipment, enhancing the target position information capture capability of the medium and large-sized equipment target detection model. After training the substation equipment dataset with the medium and large-sized equipment target detection model, the model was built into the industrial control computer of the inspection robot to obtain the semantic information of the substation equipment in real time. Due to the possibility of false detection, the equipment was only mapped as a landmark onto a map with point and line features when the medium and large-sized equipment target detection model detected the spatial position information of the equipment in three consecutive keyframes. Step 3.3: Obtain the loop closure frame through loop closure detection: After map initialization is complete, a loop closure detection thread is started to perform loop closure detection on the current keyframes obtained in real time during the inspection. If the current keyframe contains substation equipment landmarks, the most recently mapped landmark is identified and compared with the landmarks in the keyframe database. The spatial proximity between them is calculated to determine whether the keyframe in the keyframe database is a loop closure candidate keyframe. in This represents the 2D projection position of the landmark in the current keyframe. The set of landmarks within keyframes in the keyframe database Landmark Collection The three-dimensional position of the i-th landmark in the map. It is a norm 2; Simultaneously, based on the point and line features in the current keyframe obtained from the visual thread, and using the point and line feature dictionary obtained in step 3.1, the descriptors of the point and line features are compared online with the cluster centers of each intermediate node in the tree structure of the dictionary. This allows the word positions of the descriptors to be found. Furthermore, by comparing the keyframes in the keyframe database, keyframes with more than 80% of the words in the current keyframe are found and stored in the closed-loop candidate keyframe group. The point and line features of the current keyframe are converted into bag-of-word vectors for point and line features, respectively. By comparing the bag-of-word vectors between the current keyframe and all closed-loop candidate keyframes in the closed-loop candidate keyframe group, an image similarity score based on point and line features is obtained. In the similarity score, the similarity scores of the two features are weighted by two criteria: feature size and distribution degree. The normalized feature size is calculated. , and distribution degree , The formula is expressed as follows: in, , and These are the number of point features, line features, and the total number of features. , and These are the distribution values ​​for point features, line features, and all features, respectively. Based on distance weighting, dispersion is added as a penalty term to the denominator of the similarity score to reduce the impact of high dispersion on the score. The formula is as follows: in and These are the weighting coefficients for point features and line features; Based on the weighted similarity score, the candidate keyframe with the highest similarity score to the current keyframe is selected as the closed-loop keyframe. However, this closed-loop keyframe must undergo a temporal consistency check: it cannot be a keyframe within the sliding window of the current keyframe. Furthermore, a spatial coherence check is required: the similarity score between this closed-loop keyframe and all other keyframes within the current sliding window must be greater than a set similarity threshold, which is set to 0.75 times the similarity score between the closed-loop keyframe and the current keyframe. Only when both conditions are met can this closed-loop keyframe be determined as the optimal closed-loop keyframe.

6. The substation inspection robot localization and mapping method based on point-line feature fusion according to claim 5, characterized in that, Step 4 is explained in detail below: Step 4.1: Calculate the relative pose between the current frame and the closed-loop frame using the optimized fast relocalization method: Once the optimal closed-loop keyframe is detected, a coarse matching based on descriptors is first performed using the point and line features of the optimal closed-loop keyframe and the current keyframe. The PROSAC algorithm is then used to calculate the homography matrix for high-quality point and line pairs, and the point and line projection residuals are calculated to determine interior and exterior points, eliminating incorrectly matched point and line features. The pose and point and line features of the optimal closed-loop keyframe are then used as closed-loop constraints and added to the overall residual function of the back-end nonlinear optimization. A sliding window is used to optimize the pose of the optimal closed-loop keyframe, thereby calculating the relationship between the current keyframe and the optimal closed-loop keyframe. The relative pose relationship between them is expressed by the residual function formula as follows: in, , , and These are the marginalized prior residuals, IMU residuals, point feature residuals, and line feature residuals, respectively. It is the residual of the closed-loop term. It refers to the landmark residual, which is the spatial proximity of landmarks. It is the set of point and line features in the optimal closed-loop keyframe. and These are the measured values ​​of point and line features, respectively. and These are the rotation quaternion and spatial position of the optimal closed-loop keyframe pose, respectively; By using the multi-view constraints of the sliding window, the tightly coupled optimized pose is obtained, the accumulated offset error is calculated, and the pose of the key frames in the sliding window is corrected. All frames in the sliding window are aligned with the past poses to complete fast relocalization. Step 4.2: Global closed-loop optimization: When a keyframe that matches the optimal closed-loop keyframe slides out of the sliding window, closed-loop optimization is performed on all image frames in the image frame database; in global closed-loop optimization, only the position needs to be optimized. and yaw angle The global residual objective function of position and yaw angle of sequence frames and optimal closed-loop key frames is constructed by global BA function, and the Ceres library is called to solve the optimization problem to obtain the optimal attitude of each frame. The spatial position of its three-dimensional points and three-dimensional lines is optimized accordingly, and finally a complex outdoor three-dimensional map of substation that can be used for stable navigation is obtained.

Citation Information

Patent Citations

  • Visual sense simultaneous localization and mapping method based on dot and line integrated features

    CN106909877A

  • Pose optimization method and device

    CN112444242A