A Method for Accurately Extracting Visual SLAM Static Features in a Parking Scenario
By combining the deep learning SuperPoint and YOLOv5 networks, dynamic feature points in parking scenes are eliminated, and the dynamic feature points influence of visual SLAM in parking scenes is solved, and static feature extraction with higher accuracy and robustness is achieved, improving the reliability of the system.
Patent Information
- Application Number
- CN202211028947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-23
AI Technical Summary
The existing visual SLAM algorithm is difficult to effectively remove dynamic feature points in parking scenarios, resulting in failed positioning, and the traditional feature extraction algorithm is not robust in lighting and scene changes.
The SuperPoint network adopting deep learning combines the YOLOv5 object detection network, and eliminates dynamic object feature points through mask mask and adjacent frame comparison, improves the use of deep separable convolution of SuperPoint network, achieves lightweight and robustness enhancement, and uses multi-threaded parallel technology to improve efficiency.
It improves the accuracy of static feature extraction of visual SLAM in parking scenarios, weakens the impact of dynamic objects on map building and positioning, and enhances the reliability and robustness of the system.
Smart Images

Figure CN115439743B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual SLAM and deep learning, and in particular relates to a method for removing dynamic feature points by using deep learning target detection in visual SLAM in a parking scene, which can accurately extract visual SLAM static features in a parking scene to facilitate map construction. Background Art
[0002] Simultaneous localization and mapping (SLAM) technology uses the robot's own sensors to acquire and process information about the surrounding environment without prior knowledge of the environment, thereby completing map construction and the robot's own positioning. Using camera sensors to perceive the surrounding environment is called visual SLAM. Visual sensors have become a common sensor in modern SLAM research due to their low cost and rich information collection.
[0003] With the continuous improvement and development of visual SLAM technology, a number of excellent open source SLAM frameworks such as ORB-SLAM2 and OpenVSLAM have emerged. Classic visual SLAM mainly consists of several modules: sensor data input, front-end visual odometer, back-end optimization, loop detection, and mapping. ORB-SLAM2 runs in parallel with three threads: tracking, local mapping, and loop detection. It uses the traditional ORB algorithm for feature extraction. It has been verified that its robustness to lighting is poor. In recent years, some feature extraction algorithms based on deep learning have also emerged. Feature extraction plays a vital role in the entire visual SLAM system, ensuring that the extracted feature points have a good representative effect on the scene. Most of the current visual SLAM can only extract all the feature points in the current scene, but the feature points on dynamic objects such as vehicles and pedestrians may not be matched during the next positioning, resulting in positioning failure. Summary of the invention
[0004] In view of the shortcomings of the above-mentioned traditional visual SLAM algorithm in terms of dynamic feature points and feature point accuracy and robustness, the purpose of the present invention is to provide a method for accurately extracting visual SLAM static features in parking scenarios. This scheme can improve the accuracy of feature point extraction, and can effectively reduce the influence of dynamic object feature points on visual SLAM mapping and positioning, thereby improving the reliability of visual SLAM.
[0005] In order to solve the above technical problems, the present invention provides a method for accurately extracting visual SLAM static features in a parking scene, comprising the following steps:
[0006] Step 1: Extract an image of the parking lot scene in front of the vehicle. After preprocessing the image, input it into the object detection network for object detection to obtain the detection bounding boxes of the objects.
[0007] Step 2: Filter the detection bounding boxes of dynamic objects output in Step 1, and form a mask. Combine it with the feature points extracted by SuperPoint, eliminate the feature points of the dynamic object detection bounding boxes, and obtain the key points and descriptors. The SuperPoint network includes a shared encoder for key points and descriptors, a key point decoder, and a descriptor decoder. The shared encoder is used to encode the image to obtain a feature map. The key point decoder is used to obtain the coordinates of the key points in the obtained image. The descriptor decoder is used to obtain the descriptor vectors of the key points. Among them, the improvement of the SuperPoint network includes: changing all the convolutions in the encoder to depthwise separable convolutions. Among them, multi-thread parallel technology is used for object detection and feature extraction, and object detection is performed while feature extraction is in progress.
[0008] Step 3: If the mask represents a pedestrian, the SuperPoint network eliminates the feature points within the mask; if it is a car, compare the car object detection regions in adjacent frames. Retain the feature points in the non-common part of two adjacent object detection regions, and eliminate the feature points in the common part to obtain the filtered static feature points.
[0009] Step 4: Use the SuperPoint network to extract and eliminate the key points and descriptors within the mask, use the remaining feature points for feature matching, continue to execute the tracking module of visual SLAM, calculate the camera pose and build a map to complete the entire SLAM work.
[0010] In Step 1, the YOLOv5 network is used as the object detection network. The object detection algorithm based on YOLOv5 is as follows: Input an RGB image of 608*608*3, scale the input image to the input size of the network, and use Mosaic for data augmentation. Mosaic randomly selects 4 pictures for scaling, rotation, and arrangement to form a new picture, which not only greatly increases the number of pictures but also speeds up the training speed to achieve the effect of data augmentation; The Backbone module uses the CSPDarknet53 structure and the Focus structure to extract some general features; The extracted general features are sent to the Neck network to extract more diverse and robust features, input into the CSP2_X and CBL structures, and after upsampling, concatenate with the features output by the backbone network to enhance the ability of feature fusion; Finally, at the output end, CIoU_LOSS is used instead of the previous GIoU_LOSS as the loss function for the Bounding Box. The CIoU formula is as follows:
[0011]
[0012]
[0013]
[0014] CIou considers the size ratio of the ground truth box and the predicted box. In the formula, v ∈ [0, 1] represents the normalized representation of the ratio difference between the length and width of the predicted box and the corresponding ground truth box. α represents the loss balance factor.
[0015] In step 2, before training the SuperPoint network adopted in step 2,
[0016] The SuperPoint network is extracted in a self-supervised manner. First, a regular geometric shape is used as the dataset to train a fully convolutional network - Base Detector; the detection results of the Base Detector network are used as the pseudo Ground Truth Keypoint (pseudo true value keypoint) for the unlabeled real images. In order to make the pseudo Ground Truth Keypoint more robust and accurate, the Homographic Adaptation technology (homography technology) is used to extract features from the unlabeled real images at different sizes to generate pseudo labels; after generating the pseudo labels, the real unlabeled images can be put into the SuperPoint network for training. In the image input stage, data augmentation means such as flipping are adopted.
[0017] The SuperPoint network includes a shared encoder for keypoints and descriptors, a keypoint decoder, and a descriptor decoder. Further, in step 2, the process of the SuperPoint network detecting keypoints and descriptors is as follows:
[0018] Input an image frame of H*W*3, grayscale it and convert it to H*W*1, and then input the image into the improved and more lightweight shared encoder. After passing through the encoder, the input image size is converted to H c = H / 8, W c = W / 8, so as to reduce the image size;
[0019] The keypoint decoder performs sub-pixel convolution operations, and through the depth to space process, the input vector is converted from H / 8*W / 8*65 to H*W, and the final output is the probability that each pixel point is a Keypoint;
[0020] The descriptor decoder uses a convolutional network to obtain semi-dense descriptors, then uses bicubic differences to obtain the remaining descriptors, and finally obtains descriptors of a unified length (H*W*D) through L2 normalization.
[0021] In step 2, an improved SuperPoint shared encoder is used. The original SuperPoint encoder uses convolutional network layers similar to VGG6, but the computational complexity and the number of training parameters are huge. In the present invention, all convolutions in the encoder part are changed to depthwise separable convolutions. The normal convolution process is as Figure 3 shown. Assuming the input image size is H*W*3 and m layers of feature maps are output, the number of parameters of the ordinary convolution kernel is 3*f*f*m;
[0022] The depthwise separable convolution (as Figure 4 shown) is divided into two consecutive processes: depthwise convolution and pointwise convolution. Depthwise convolution gives each channel a separate convolution kernel for convolution, transforms the convolution process into a two-dimensional plane, and finally generates a mid feature map. The number of parameters of the convolution kernel in this link is f*f*3. The generated mid feature map is subjected to pointwise convolution using a 1*1*3 convolution kernel, which has the effect of data fusion and finally outputs m layers of feature maps. The number of parameters in this part is 1*3*m. Then the number of parameters of the depthwise kernel and the pointwise convolution kernel is 3*(f*f + m), which is one order of magnitude lower than that of the direct convolution 3*f*f*m, and the time efficiency will be greatly improved. Although the number of learned parameters is reduced compared with ordinary convolution, the accuracy does not drop too much.
[0023] In step 2, the loss function of the improved SuperPoint consists of two parts: the key point extraction loss and the descriptor detection loss:
[0024]
[0025] In the formula, the loss function consists of two parts: is the key point loss, and the cross-entropy loss function is used. is the descriptor loss, and λ is the balance factor.
[0026] In step 2, the operation mechanism of object detection and feature extraction is as follows: Using multi-thread parallel technology, object detection is carried out while feature extraction is in progress. The tracking thread does not need to wait for the detection results of YOLOv5, which improves the utilization rate of the CPU and enhances the operation efficiency.
[0027] In step 3, the feature points in the detection frames of dynamic objects (predetermined vehicles and pedestrians) in steps 1 and 2 are removed. If the feature points in each dynamic object frame are directly removed, it will be difficult to match due to too few feature points. AsFigure 7 As shown, the feature points of the detected dynamic target boxes for two adjacent frames are respectively ( Figure 7 area A in ( Figure 7 area B in indicating the i-th feature point in the n-th frame. The intersection of the detected same dynamic object target boxes in two adjacent frames is used as the final dynamic target feature points, that is, D = D n D n+1 ( Figure 7 area C of
[0028] In step 4, the process of calculating the camera pose is as follows: The feature points and descriptors screened out in the above process are used for image matching, and the RANSAC random sampling consensus algorithm is used to eliminate the mismatched feature points, and the epipolar geometry problem from 2D points to 2D points is transformed according to the matching relationship. Assume that x1 and x2 are the normalized coordinates of the corresponding matching points on two images, R is the camera rotation matrix, and t is the translation matrix, then it satisfies
[0029] x2 = Rx1 + t
[0030] Multiply x T 2t on the left:
[0031] If the left side of the equation is 0, then: That is the epipolar constraint expression, and the camera pose can be obtained according to the minimum reprojection error. Let the fundamental matrix E = t^R, and the essential matrix F = K -T EK -1 . Solving the camera pose can be transformed into the following two steps: obtaining the fundamental matrix E or the essential matrix F; solving R and t according to E or F.
[0032] In step 4, the screening of key frames has a greater impact on information redundancy and the release of computing resources. If the system is in the positioning mode, the local map is occupied, or the relocalization has just ended, no key frames are inserted.
[0033] Compared with the prior art, the present invention can at least achieve the following beneficial effects:
[0034] The present invention uses the SuperPoint network based on deep learning in combination with an object detection network (such as YOLOv5) to extract key points, descriptors, and dynamic object detection respectively. Most traditional solutions use ORB and SURF for feature extraction, but the feature extraction effect is not good when the parking lot scene changes and the light intensity changes significantly. The improved SuperPoint used in the present invention first makes the network model more lightweight, and finally enables the feature extraction to be more robust to different scene changes, and the extracted feature points are more uniform and reasonable. Description of the Drawings
[0035] Figure 1 It is a schematic flow chart of a method for accurately extracting static features of visual SLAM in a parking scenario provided by an embodiment of the present invention.
[0036] Figure 2 It is a core flow chart of the improved lightweight SuperPoint provided by an embodiment of the present invention.
[0037] Figure 3 It is a schematic diagram of ordinary convolution.
[0038] Figure 4 It is a schematic diagram of depthwise separable convolution, where (a) is a schematic diagram of per-channel convolution and (b) is a schematic diagram of pointwise convolution.
[0039] Figure 5 It is a detection effect diagram of YOLOv5 in a parking lot scene provided by an embodiment of the present invention.
[0040] Figure 6 It is a feature extraction effect diagram of SuperPoint provided by an embodiment of the present invention.
[0041] Figure 7 It is a schematic diagram of dynamic feature screening provided by an embodiment of the present invention (the area marked C is the removed area). Detailed Embodiments
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in conjunction with the present application.
[0043] Traditional visual SLAM feature extraction is based on the assumption of a static environment. However, for each parking space in a parking lot environment, there is not always a car parked. Therefore, it is necessary to remove the feature points of some dynamic objects on the parking spaces to ensure that these dynamic features are not retained in the final map. The dynamic feature points proposed in the prior art generally use traditional feature extraction algorithms such as ORB and SURF. However, the effects of these feature points vary greatly for different scene changes and are not very robust. In the present invention, the SuperPoint network based on deep learning is combined with the YOLOv5 algorithm to extract key points, descriptors, and detect dynamic targets respectively, and the SuperPoint is improved to make the network model more lightweight. Finally, the feature extraction is more robust to different scene changes, and the extracted feature points are more uniform and reasonable. The method provided by the present invention will be specifically introduced below.
[0044] As Figures 1 to 6 shown, the present invention discloses a method for accurately extracting static features of visual SLAM in a parking scenario, including the following steps:
[0045] Step 1: Extract an image of the parking lot scene in front of the vehicle. After preprocessing the image, input the image into a target detection network for target detection to obtain the detection box of the target object.
[0046] In some embodiments of the present invention, the target detection network used is the YOLOV5 network model. It can be understood that in other embodiments, other target detection networks can also be used.
[0047] Use a monocular camera to collect images of the parking lot scene in front of the vehicle in real time. After performing image preprocessing operations (such as filtering, image enhancement, etc.), input an RGB image of H*W*3 (H and W respectively represent the number of pixels of the length and width of an image) into the YOLOV5 network model. The input image is scaled to the input size of the network, and Mosaic is used for data augmentation. Mosaic randomly selects 4 images for scaling, rotation, and arrangement to form a new image, which not only greatly increases the number of images but also can accelerate the training speed to achieve the effect of data augmentation; the backbone network extracts image features to generate a feature map; the Backbone module uses the CSPDarknet53 structure and the Focus structure to extract some general features; the extracted general features are sent to the Neck network to extract more diverse and robust features, input into the CSP2_X and CBL structures, and after upsampling, they are spliced with the features output by the backbone network to enhance the ability of feature fusion; finally, the CIoU_LOSS is used instead of the previous GIoU_LOSS as the loss function of the Bounding Box at the output end. The CIoU formula is as follows:
[0048]
[0049]
[0050]
[0051] Wherein, CIou takes into account the size ratio of the ground truth box and the predicted box, and CIoU(B pre , B GT ) represents the CIOU intersection over union ratio between the predicted box and the ground truth target box; Iou(B pre , B Gr ) represents the intersection over union ratio between the predicted box and the ground truth target box; B pre represents the predicted box; B GT represents the ground truth detection box; ρ(B pre , B GT ) represents the distance between the center points of the predicted box and the ground truth box; v represents the similarity of the width-to-height ratio of the predicted box and the ground truth box, and v ∈ [0, 1]; w GT represents the width of the ground truth box; h GT represents the height of the ground truth box; GT represents the ground truth box information; w and h respectively represent the width and height of the predicted box; α represents the loss balance factor.
[0052] Step 2: Screen the dynamic object detection boxes output in Step 1, form a mask, and use it in combination with the feature points extracted by SuperPoint to eliminate the feature points of the dynamic object detection boxes.
[0053] In some embodiments of the present invention, the dynamic objects in the dynamic object detection boxes include: vehicles, pedestrians, and a small number of animals.
[0054] Meanwhile, another thread performs the feature extraction process. The collected RGB image with a size of H*W*3 is grayscaled and then input into the lightweight SuperPoint network. The SuperPoint network is extracted in a self-supervised manner. First, a regular geometric shape is used as a dataset to train a fully convolutional network Base Detector; the detection results of the Base Detector network are used as pseudo ground truth keypoints (pseudo Ground Truth Keypoint) for the unlabeled real images. In order to make the pseudo ground truth keypoints more robust and accurate, the homographic adaptation technique is used to extract features from the unlabeled real images at different sizes to generate pseudo labels; after generating the pseudo labels, the real unlabeled images can be put into the SuperPoint network for training. In the image input stage, data augmentation means such as flipping are adopted.
[0055] The SuperPoint network includes a shared encoder for keypoints and descriptors, a keypoint decoder, and a descriptor decoder. The shared encoder is used to encode an image to obtain a feature map. The keypoint decoder is used to obtain the coordinates of keypoints in the image, and the descriptor decoder is used to obtain the descriptor vectors of the keypoints.
[0056] Specifically, an image frame of H*W*3 is input. After grayscaling, it is converted into H*W*1. Then the image is input into an improved and more lightweight shared encoder. In the encoder, all ordinary convolution operations are converted into depthwise separable convolutions. For ordinary convolution, if the output feature map is set to have m layers, the number of convolution kernel parameters is 3*f*f*m (f represents the convolution kernel size, and m represents the final output number of channels). After depthwise separable convolution, the number of parameters becomes 3*(f*f + m), which is an order of magnitude lower than 3*f*f*m of direct convolution, and the time efficiency will be greatly improved. Although the number of learned parameters is decreased compared with ordinary convolution, the accuracy does not decrease much.
[0057] In some embodiments of the present invention, the depthwise separable convolution includes two consecutive processes: depthwise convolution and pointwise convolution. Depthwise convolution is to perform convolution on each channel with a separate convolution kernel, and the convolution process is transformed into a two-dimensional plane for convolution, finally generating an intermediate feature map. The number of convolution kernel parameters in this part is f*f*3. The generated intermediate feature map is subjected to pointwise convolution using a 1*1*3 convolution kernel, which has the effect of data fusion, and finally m layers of feature maps are also output. The number of parameters in this part is 1*3*m. Then the number of convolution kernel parameters of depthwise convolution and pointwise convolution is 3*(f*f + m), which is an order of magnitude lower than 3*f*f*m of direct convolution, and the time efficiency will be greatly improved. Although the number of learned parameters is decreased compared with ordinary convolution, the accuracy does not decrease much.
[0058] After passing through the encoder, the input image size is transformed into H c = H / 8, W c = W / 8, thereby reducing the image size. The keypoint decoder performs sub-pixel convolution operations. Through the depth to space process (moving the data in the depth dimension to the space dimension), the input vector is transformed from H / 8*W / 8*65 into an H*W vector. The H*W vector undergoes softmax operation, and then through Reshape for dimension conversion, finally the probability that each pixel point is a keypoint (KeyPoint) is output, represented in vector form. The places that are selected as keypoints through the probability threshold are the keypoint coordinates. The descriptor detector uses a convolutional network to obtain semi-dense descriptors, then uses bicubic differences to obtain the remaining descriptors, and finally obtains descriptors of a unified length (H*W*D) through L2 normalization.
[0059] The loss function for improving SuperPoint consists of two parts: the key-point extraction loss and the descriptor detection loss:
[0060]
[0061] In the formula, the loss function consists of two parts: is the key-point loss, is the key-point loss of the flipped image, and the cross-entropy loss function is used. is the descriptor loss, represents the response to the key points after the image passes through the encoding network model, and the size is represents the response to the key points after the original image passes through the encoding network model after flipping; D represents the response to the descriptors after the image passes through the encoding network model; D′ represents the response to the descriptors after the flipped image passes through the encoding network model; Y represents the key-point coordinate label; Y′ represents the key-point coordinate label of the flipped image; S represents the image pair composed of the original image and the flipped image; λ is the balance factor.
[0062] Step3: If the mask represents a pedestrian, the SuperPoint network removes the feature points within the mask; if it is a car, then compare the car target detection regions in adjacent frames. The feature points in the non-common part of two adjacent target detection regions are retained, and the feature points in the common part are removed to obtain the filtered static feature points.
[0063] The object detection box and depth features use multi-thread parallel technology to perform object detection while extracting features. The tracking thread does not need to wait for the detection results of the object detection network, which improves the utilization rate of the CPU and enhances the operation efficiency.
[0064] Remove the feature points in the detection boxes of dynamic objects (such as vehicles and pedestrians) in Steps 1 and 2. If the feature points in each dynamic object box are directly removed, it will be difficult to match due to too few feature points. Therefore, in some embodiments of the present invention, for the feature points in the detected dynamic target boxes in two adjacent frames, they are respectively (such as Figure 7 in Region A), (such as Figure 7 in Region B), represents the i-th feature point in the n-th frame, p represents the total number of feature points in the n-th frame, and q represents the total number of feature points in the (n + 1)-th frame. The intersection of the detected dynamic object target boxes in two adjacent frames is used as the final dynamic target feature points, that is, D = D n D n+1 (such as Figure 7In the C area), the feature points in set D are used as the final set of dynamic feature points, and the feature points in set D are removed, which reduces the probability of mis-removing dynamic feature points and also retains some suspected static feature points. The remaining A and B areas are reserved areas, which increases the reliability of the tracking thread. The filtered static feature points are saved for subsequent feature matching and pose calculation.
[0065] Step4: Use the SuperPoint network to extract and remove the key points and descriptors within the mask, use the remaining feature points for feature matching, continue to execute the Tracking module of visual SLAM, calculate the camera pose using the minimum reprojection error and build a map to complete the entire SLAM work.
[0066] Perform image matching on the feature points and descriptors filtered in the above process. In some embodiments of the present invention, the RANSAC (Random Sample Consensus) algorithm is used to remove the mismatched feature points, and the epipolar geometry problem from 2D points to 2D points is transformed according to the matching relationship. Assume that x1 and x2 are the normalized coordinates of the corresponding matching points on two images, R is the camera rotation matrix, and t is the translation matrix, then it satisfies
[0067] x2 = Rx1 + t
[0068] Left multiply by x T 2t:
[0069] If the left side of the equation is 0, then: That is the epipolar constraint expression, and the camera pose can be obtained according to the minimum reprojection error. Let the fundamental matrix E = t^R, and the essential matrix F = K -T EK -1 , where K is the camera intrinsic matrix, E is the fundamental matrix, T represents the transpose operation of the matrix, t is the translation matrix, and R is the rotation matrix. Solving the camera pose can be transformed into the following two steps: find the fundamental matrix E or the essential matrix F; according to E or F, solve for R and t. If the system is in the localization mode, the local map is occupied, or the relocalization has just ended, no key frame is inserted. On the premise of not satisfying the above three situations, if the number of inlier points matched in the current frame exceeds the set threshold, the current frame can be set as a key frame, the tracking thread is completed, and then the mapping and loop closure detection continue, and finally the entire map is established.
[0070] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for accurately extracting visual SLAM static features in a parking scenario, characterized in that It includes the following steps: Step 1: Extract an image of the parking lot scene in front of the vehicle. After preprocessing the image, input it into a target detection network for target detection to obtain the detection bounding boxes of the targets; Step 2: Screen the detection bounding boxes of dynamic objects output in Step 1, and form a mask. Combine it with the feature points extracted by the SuperPoint network, eliminate the feature points of the dynamic object detection bounding boxes, and obtain key points and descriptors. The SuperPoint network includes a shared encoder for key points and descriptors, a key point decoder, and a descriptor decoder. The shared encoder is used to encode the image to obtain a feature map. The key point decoder is used to obtain the coordinates of the key points in the image. The descriptor decoder is used to obtain the descriptor vectors of the key points. Among them, the improvement of the SuperPoint network includes: changing all the convolutions in the encoder to depthwise separable convolutions. Among them, multi-thread parallel technology is used for target detection and feature point extraction, and target detection is performed while feature points are being extracted; Step 3: If the mask represents a pedestrian, the SuperPoint network eliminates the feature points within the mask; if it is a car, then compare the car target detection regions of adjacent frames. Retain the feature points in the non-common part of two adjacent target detection regions, and eliminate the feature points in the common part to obtain the filtered static feature points; Step 4: Use the SuperPoint network to extract and eliminate the key points and descriptors within the mask, perform feature matching using the remaining feature points, continue to execute the tracking module of visual SLAM, calculate the camera pose and build a map to complete the entire SLAM work; Among them, the depthwise separable convolution in Step 2 includes two consecutive processes: depthwise convolution and pointwise convolution. Depthwise convolution is to perform convolution on each channel with a separate convolution kernel, convert the convolution process into a two-dimensional plane, and finally generate an intermediate feature map. The number of convolution kernel parameters in this link is f*f*3. The generated intermediate feature map is subjected to pointwise convolution using a 1*1*3 convolution kernel, and finally m layers of feature maps are also output. The number of parameters in this part is 1*3*m; the number of convolution kernel parameters of the depthwise kernel and pointwise convolution is 3*(f*f + m), where f represents the convolution kernel size.
2. The method for accurately extracting visual SLAM static features in a parking scenario according to claim 1, characterized in that, In Step 1, the YOLOv5 network is used as the target detection network. The process of target detection includes: Input an RGB image, scale the input image to the input size of the network, and perform data augmentation; The backbone network extracts image features to generate a feature map. The Backbone module uses the CSPDarknet53 structure and the Focus structure to extract general features; convey the extracted general features to the Neck network to extract more diverse and robust features, input them into the CSP2_X and CBL structures, and after upsampling, splice them with the features output by the backbone network; finally, the CIoU_LOSS is used as the loss function of the Bounding Box at the output end.
3. A method for accurately extracting visual SLAM static features in a parking scenario according to claim 1, characterized in that, Before training, the SuperPoint network adopted in Step 2 The SuperPoint network is extracted in a self-supervised manner. First, a fully convolutional network is trained using regular geometric shapes as the dataset. The detection results of the fully convolutional network are used as pseudo-ground truth key points for unlabeled real images, and the homography technique is used to extract features from unlabeled real images at different scales to generate pseudo-labels. After generating the pseudo-labels, the real unlabeled images can be put into the SuperPoint network for training.
4. A method for accurately extracting visual SLAM static features in a parking scenario according to claim 1, characterized in that, In step 2, the process of the SuperPoint network detecting key points and descriptors is as follows: Input an image frame of H*W*3, grayscale it and convert it to H*W*1, then input the image into an improved and more lightweight shared encoder. After passing through the shared encoder, the input image size is converted to H c = H / 8, W c = W / 8; The key point decoder performs sub-pixel convolution operations and transforms the input vector from H / 8*W / 8*65 to H*W, and finally outputs the probability that each pixel point is a key point. The descriptor decoder uses a convolutional network to obtain semi-dense descriptors, then uses bicubic differences to obtain the remaining descriptors, and finally obtains descriptors of a unified length through L2 normalization.
5. A method for accurately extracting visual SLAM static features in a parking scenario according to claim 1, characterized in that In step 2, the loss function of the improved SuperPoint network consists of two parts: key point extraction loss and descriptor detection loss. Composition: Wherein, is the key point loss, is the key point loss of the flipped image, is the descriptor loss, χ represents the response of the image to the key points after passing through the encoding network model; χ′ represents the response of the original image to the key points after flipping and passing through the encoding network model; D represents the response of the image to the descriptor after passing through the encoding network model; D′ represents the response of the flipped image to the descriptor after passing through the encoding network model; Y represents the key point coordinate label; Y′ represents the key point coordinate label of the flipped image; S represents the image pair composed of the original image and the flipped image; λ is the balance factor.
6. The method for accurately extracting visual SLAM static features in a parking scenario according to claim 1, wherein The method of removing feature points in step 3 is: For two adjacent frames, the feature points of the detected dynamic object bounding boxes are respectively denotes the i-th feature point in the n-th frame, p denotes the total number of feature points in the n-th frame, q denotes the total number of feature points in the (n + 1)-th frame. The intersection of the detected dynamic object bounding boxes in two adjacent frames is taken as the final dynamic object feature points, that is, D = D n D n+1 , and the feature points in the D set are used as the final dynamic feature point set.
7. A method for accurately extracting visual SLAM static features in a parking scenario, characterized in that, according to any one of claims 1-6 In step 4, the process of calculating the camera pose is as follows: The selected feature points and descriptors are used for image matching, and the mis-matched feature points are removed. According to the matching relationship, it is transformed into an epipolar geometry problem from 2D points to 2D points. Assuming that x1 and x2 are the normalized coordinates of the corresponding matching points on two images, R is the camera rotation matrix, t is the translation matrix, and T represents the transpose operation of the matrix, then it satisfies x2 = Rx1 + t If the left side of the equation is 0, then: This is the epipolar constraint expression, and the camera pose can be obtained according to the minimum reprojection error.
8. A method for accurately extracting visual SLAM static features in a parking scenario, characterized in that, The RANSAC (Random Sample Consensus) algorithm is used to remove the mis-matched feature points.
9. A method for accurately extracting visual SLAM static features in a parking scenario, characterized in that, In step 4, when building the map, if the system is in the localization mode, the local map is occupied, or the relocalization has just ended, no key frames are inserted.
Citation Information
Patent Citations
Monocular vision inertia SLAM method for dynamic scene
CN111156984A
Semantic vision SLAM positioning method based on target detection in indoor dynamic scene
CN114677323A
Cited By
Dynamic isolation slam system and method based on semantic-enhanced mixed gaussian splatting
CN122650938A