A vehicle positioning method, device and computer-readable storage medium
By combining visual inertial odometers and semantic segmentation networks, stable semantic features are extracted for vehicle positioning and map updates, the problems of low positioning accuracy and poor repositioning effect in automatic parking technology are solved, and higher accuracy and stable vehicle positioning are achieved.
Patent Information
- Application Number
- CN202110016641.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-06
AI Technical Summary
In the existing automatic parking technology, the vehicle positioning accuracy is low and the repositioning effect is poor. Especially in the parking lot environment, it is greatly affected by changes in visual feature points and light, resulting in unstable positioning results.
Combining the visual inertial odometer and semantic segmentation network, relatively stable semantic features of driving scene images are extracted, global pose map optimization is performed through loop detection, and the map is updated to improve positioning accuracy and robustness.
By introducing a deep learning semantic segmentation network to obtain static area features in the environment, the accuracy and stability of vehicle positioning are improved, the accuracy of repositioning is enhanced, and the impact of light and dynamic objects is reduced.
Smart Images

Figure CN114723779B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of positioning and navigation, and particularly to a vehicle positioning method, device, and computer-readable storage medium. Background Art
[0002] In recent years, the development of automotive autonomous driving technology has been rapid, and the automatic parking technology has become one of the key research and development focuses in the field of autonomous driving. Not limited to the operation of parking into the garage, the automatic parking technology has been extended to a comprehensive parking system including autonomous low-speed cruising, parking space searching, parking, and summon response.
[0003] The existing automatic parking technology is mainly implemented based on the vehicle positioning algorithm of visual SLAM: by using an on-vehicle camera to acquire images, corresponding visual feature points are extracted and tracked, and the pose of the vehicle itself and the spatial positions of the feature points are estimated; then loop detection is used for global pose optimization to achieve vehicle positioning, and at the same time, the feature point information and pose information are saved to construct a map for subsequent vehicles to load and reuse when repositioning.
[0004] However, in actual applications, the areas where visual feature points in the parking lot are concentrated change at different times, such as moving vehicles or pedestrians, etc. This results in that after a period of time, the vehicle cannot achieve repositioning at the current moment through the map constructed by the existing technology. On the other hand, the feature points and vehicle pose information obtained through the camera sensor carried by the vehicle are greatly affected by the measurement distance of the camera, the parking lot environmental light, and rapid movement, and the accuracy of the obtained vehicle positioning result is often low. Summary of the Invention
[0005] The present invention provides a vehicle positioning and map construction method and system, which can achieve vehicle repositioning and map construction through a visual inertial odometer and a semantic segmentation network model, and solve the problems of low vehicle positioning accuracy and poor repositioning effect existing in the prior art.
[0006] In a first aspect, an embodiment of the present invention provides a vehicle positioning method, including:
[0007] Obtain a driving scene image and pose information of the vehicle in the current driving state;
[0008] Extract semantic features of the current driving scene image by using a semantic segmentation network, where the semantic features include scene features that are relatively stable both spatially and temporally;
[0009] Perform loop detection on the semantic features in the recently updated historical map loaded. If a loop is detected, perform global pose graph optimization on the driving trajectory of the vehicle according to the pose information and update the map.
[0010] In one embodiment, the pose information is output by a visual inertial odometer and is associated with key frames of the driving scene image; wherein, the driving scene image is acquired by a front-view camera on the vehicle; the visual inertial odometer includes an omnidirectional fisheye camera and an IMU on the vehicle.
[0011] In one embodiment, the method further includes: extracting the pose information corresponding to the current frame and the key frame of the scene image acquired from the omnidirectional fisheye camera on the vehicle, and performing image encoding on the differences between the scene images according to the pose information;
[0012] When the return result of the visual inertial odometer is normal, if the current frame is a key frame, store the image and pose information of the current frame; if the current frame is a non-key frame, discard the current frame and continue to read the next frame of image;
[0013] When the return result of the visual inertial odometer is abnormal, match the key frame most similar to the current frame in the key frame set according to the image coding threshold, and re-perform visual inertial odometer calculation. If there is no matching result, discard the current frame and continue to read the next frame of image.
[0014] In one embodiment, after the return result of the visual inertial odometer is normal and it is determined that the current frame is a key frame, it further includes:
[0015] Align the visual information of the key frame and the pre-integration result of the IMU in a loose coupling manner, and calculate the visual feature point information and the IMU state variable information through a non-linear method;
[0016] Perform non-linear optimization on the feature point information and the IMU state variable information to obtain the pose information at the current moment corresponding to the scene image.
[0017] In one embodiment, the extraction of the semantic features of the current driving scene image by using a semantic segmentation network includes:
[0018] Input the current driving scene image into the semantic segmentation network, identify the static features, semi-static features and dynamic features in the scene features of the current driving scene image, and generate a mask for removing the semi-static features and the dynamic features;
[0019] Under the action of the mask, output the static features as semantic features.
[0020] In one embodiment, the driving scene image in the current driving state includes the key frame image acquired by the current front-view undistorted camera;
[0021] The loop closure detection of the semantic features in the recently updated historical map loaded includes:
[0022] Extract the static features of the key frame image obtained by the current front view undistorted camera, where the static features include the spatial coordinates of the feature points of the key frame image;
[0023] Extract the BRIEF descriptor corresponding to the feature points, convert the descriptor into a word vector of the bag-of-words model, and then perform similarity matching between the word vector and the feature point word vectors of the corresponding historical frames in the historical map;
[0024] Determine whether the key frame obtained by the current front view undistorted camera and the corresponding historical frame in the historical map form a loop according to a preset similarity threshold.
[0025] In one embodiment, the global pose graph optimization of the driving trajectory of the vehicle according to the pose information and updating the map include:
[0026] Solve the change amount of the vehicle pose at the current moment relative to the historical moment according to the spatial coordinates of the feature points that have been matched between the key frame and the historical frame obtained by the current front view undistorted camera, and the external parameters between the corresponding camera and the IMU, and perform global pose graph non-linear optimization and update the map by minimizing the pose residuals between the two loop frames and the pose residuals between all key frames under the global trajectory.
[0027] In a second aspect, an embodiment of the present invention further provides a vehicle positioning device, including:
[0028] A driving data acquisition unit, configured to acquire the driving scene image and pose information of the vehicle in the current driving state;
[0029] A semantic feature extraction unit, configured to extract the semantic features of the current driving scene image by using a semantic segmentation network, where the semantic features include scene features that are relatively stable both spatially and temporally;
[0030] A positioning unit, configured to perform loop detection on the semantic features in the recently updated historical map loaded. If a loop is detected, global pose graph optimization of the driving trajectory of the vehicle is performed according to the pose information and the map is updated.
[0031] In one embodiment, the pose information is output by a visual inertial odometer and is associated with the key frames of the driving scene image; where the driving scene image is obtained by a front view camera on the vehicle; the visual inertial odometer includes an omnidirectional fish-eye camera and an IMU on the vehicle.
[0032] In one embodiment, the driving data acquisition unit is specifically configured to: extract the pose information corresponding to the current frame and the key frame of the scene image obtained from the omnidirectional fish-eye camera on the vehicle, and perform image encoding on the difference between the scene images according to the pose information;
[0033] When the return result of the visual inertial odometer is normal, if the current frame is a key frame, store the image and pose information of the current frame; if the current frame is a non-key frame, discard the current frame and continue to read the next frame of image;
[0034] When the return result of the visual inertial odometer is abnormal, match the key frame most similar to the current frame in the key frame set according to the image coding threshold, and re-perform visual inertial odometer calculation. If there is no matching result, discard the current frame and continue to read the next frame of image.
[0035] In one embodiment, the driving data acquisition unit is further configured to: after the return result of the visual inertial odometer is normal and it is determined that the current frame is a key frame:
[0036] Align the visual information of the key frame and the IMU pre-integration result in a loose coupling manner, and calculate the visual feature point information and the IMU state variable information through a non-linear method;
[0037] Perform non-linear optimization on the feature point information and the IMU state variable information to obtain the pose information at the current moment corresponding to the scene image.
[0038] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the method described in any one of the above embodiments.
[0039] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows:
[0040] The vehicle positioning method, device and computer-readable storage medium proposed by the present invention combine four fisheye cameras and IMU information to obtain vehicle pose information. On this basis, a deep learning semantic segmentation network is introduced to obtain relatively stable static regions in the environment, so that the positioning result retains static feature points, improves the robustness and accuracy of the positioning method, and obtains a more stable vehicle repositioning effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for implementation will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 is a schematic flowchart of the vehicle positioning method provided by an embodiment of the present invention;
[0043] Figure 2It is a schematic diagram of the distribution positions of vehicle-mounted sensors provided by an embodiment of the present invention;
[0044] Figure 3 It is a schematic structural diagram of a vehicle positioning device provided by an embodiment of the present invention. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0047] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0048] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0049] The term " / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0050] As Figure 1 shown, an embodiment of the present invention provides a vehicle positioning method, which specifically includes the following steps:
[0051] S11: Obtain the driving scene image and pose information of the vehicle in the current driving state.
[0052] In this embodiment, the driving scene image is obtained by a front-view undistorted camera on the vehicle.
[0053] Generally, the vehicle-mounted sensors installed in an automobile usually include four-way surround-view fisheye cameras, an inertial measurement unit (IMU), and a front-view undistorted camera; As Figure 2As shown in the figure, the installation position of the four-way surround-view fisheye camera is consistent with the installation position of the 360° surround-view imaging function of ordinary cars. They are installed near the front and rear logos of the vehicle and below the left and right rearview mirrors, and are directed obliquely outward; the IMU sensor is installed at the center of the rear axle of the vehicle; and the one-way forward-view distortion-free camera is installed behind the rearview mirror below the windshield of the vehicle, and is directed toward the front of the vehicle.
[0054] Specifically, the calibration tool can be used to obtain the internal parameters of the vehicle-mounted sensors and the external parameters of the system, where the internal parameters include the pixel center, focal length and distortion coefficient of the vehicle-mounted camera, the white noise and random walk noise errors of the accelerometer and gyroscope of the IMU, and the external parameters of the system include the rotation and translation matrices between the coordinate systems of each sensor and the vehicle coordinate system. The present invention sets the position of the IMU as the origin of the vehicle coordinate system, and calibrates the rotation and translation matrices between the four fisheye cameras and the IMU, and between the front-view camera and the IMU, respectively.
[0055] In a specific embodiment, the posture information is output by a visual inertial odometer and is associated with a key frame of the driving scene image; wherein the driving scene image is acquired by a vehicle-mounted forward-looking camera; and the visual inertial odometer includes a vehicle-mounted surround-view fisheye camera and an IMU.
[0056] In this embodiment, the pose information corresponding to the current frame and the key frame of the scene image acquired from the vehicle-mounted surround fisheye camera is extracted, and image encoding is performed on the difference between the scene images according to the pose information.
[0057] When the visual inertial odometer returns a normal result, if the current frame is a key frame, the image and posture information of the current frame are stored; if the current frame is a non-key frame, the current frame is discarded and the next frame image is read.
[0058] When the visual inertial odometry returns an abnormal result, the key frame most similar to the current frame is matched in the key frame set according to the image encoding threshold, and the visual inertial odometry calculation is performed again. If there is no matching result, the current frame is discarded and the next frame image is read.
[0059] After the visual inertial odometry returns a normal result and the current frame is determined to be a key frame, the visual information of the key frame and the IMU pre-integration result are aligned in trajectory using a loosely coupled method, and the visual feature point information and the IMU state variable information are calculated using a nonlinear method; the feature point information and the IMU state variable information are nonlinearly optimized to obtain the current pose information corresponding to the scene image.
[0060] In this embodiment, the visual inertial odometer can preferably be implemented based on the extended open-source monocular visual inertial fusion positioning algorithm VINS-Mono to obtain vehicle pose information. The specific implementation steps of the preferred method are as follows:
[0061] First, sparse edge features are extracted from the images obtained by the four fisheye cameras. The edge features mainly focus on static identifiers such as parked vehicles and parking lot environments.
[0062] When the acquired image is the first frame image, N Harris feature points of the image are respectively extracted. The pixel distance between the N feature points is not less than d, where N is a preset hyperparameter of the number of feature points, and d is a hyperparameter of the minimum pixel spacing of the feature points.
[0063] When the acquired image is not the first frame image, the LK optical flow method is used to track the pixel coordinates of the feature points in the previous frame in the current frame image, and a hyperparameter of the number of feature points is preset. If the number of tracked feature points m < N, then (N - m) Harris feature points are extracted from the current image, and the number of feature points in the current frame is maintained as N.
[0064] Obtain the pixel coordinates of the feature points tracked between two frames, and the external parameters R i and T i (R i is the rotation parameter, T i is the translation parameter, i = 1, 2, 3, 4). According to the epipolar geometry relationship, triangulate the pixel coordinates of the feature points and the external parameters, and calculate the spatial coordinates of the feature points tracked in each image, and the pose information of the vehicle in the world coordinate system at the current frame.
[0065] Insert the visual feature point information corresponding to each frame of the image captured by the four fisheye cameras into the sliding window until the number of key frames in the sliding window is equal to the preset sliding window size hyperparameter W. Then use PnP to solve the spatial coordinates of the feature points observed by all key frames in the sliding window and the vehicle pose at the corresponding time of the key frames to complete visual initialization.
[0066] Align the visual feature point information with the state vector obtained by IMU pre-integration, and iteratively solve the estimated values of the state variables of the visual feature point information and the IMU tightly coupled system through a non-linear method based on the principle of minimizing the visual reprojection error and the IMU observation residual. Among them, the state variables include the poses of the key frames in the sliding window, the spatial positions of the visual feature points observed by the key frames, the IMU biases, and the state estimation covariance matrix.
[0067] When a new frame of image is inserted into the sliding window, by determining whether the average parallax of the tracked feature points between two frames of images meets a preset threshold, and further determining whether the penultimate frame in the sliding window is a key frame. If so, marginalize the first frame in the sliding window; if not, marginalize the penultimate frame and insert the new frame into the sliding window.
[0068] When a new key frame is inserted into the sliding window, it is necessary to perform non-linear optimization on the visual feature point information and IMU state variables of all key frames in the sliding window to obtain the optimal value of the system's state variable estimation. The optimal value is the vehicle pose information at the current moment, specifically including the vehicle position P and the vehicle attitude R.
[0069] Using the visual-inertial odometer proposed in the above embodiment can effectively make up for the instability caused by vision being affected by light, dynamic objects, and fast movement, and further improve the vehicle positioning accuracy.
[0070] S12: Use a semantic segmentation network to extract the semantic features of the current driving scene image, where the semantic features include scene features that are relatively stable both spatially and temporally.
[0071] In this embodiment, for areas such as cars in motion or statically parked and pedestrians in the scene image that will change after a period of time, they do not belong to the semantic features.
[0072] In a specific embodiment, the current driving scene image can be input into a semantic segmentation network to identify the static features, semi-static features, and dynamic features in the scene features of the current driving scene image, and generate a mask for removing the semi-static features and the dynamic features; under the action of the mask, output the static features as semantic features.
[0073] Among them, the semantic segmentation network used can preferably be the U-Net network, and the construction process of the U-Net network is as follows:
[0074] Extract the scene data collected by the front-view undistorted camera, label the labels of the static regions in the scene data, and obtain the training data set of the semantic segmentation network model according to the scene data and the labels. Use the obtained training data set to train the semantic segmentation network model and output an image mask.
[0075] Specifically, use the front-view undistorted camera to collect the scene image of the underground parking lot, and define the static environment region of the parking lot in the scene image as the effective region, including walls, columns, and signs, etc. At the same time, define the regions where pedestrians and vehicles in motion in the scene image change their positions over time as dynamic regions, and define the regions where statically parked cars will change after a period of time as semi-static regions.
[0076] Next, label the valid region of the image obtained by the front-view undistorted camera to generate an image mask label. Specifically, the label value of the valid region of the image can be set to 0, and the label value of the dynamic region can be set to 255 to obtain a training dataset, and the labeled training dataset is used to train the U-Net network.
[0077] In this embodiment, the resolution of the original image captured by the front-view undistorted camera is 640*360. The input image can be scaled to 0.5 times the size of the original image. After dynamic scene segmentation through training the U-Net network, a mask image of size 320*180 is output.
[0078] After dynamic scene segmentation using the U-Net semantic segmentation network model, a mask that only retains the static region can be generated. Image feature points are extracted within the mask region, and the semantic features of the scene image can be obtained. These semantic features are fixed static features and do not change easily with the change of parked vehicles, so the vehicle repositioning accuracy can be effectively improved.
[0079] In a specific embodiment, after obtaining the vehicle pose information through the visual inertial odometer, the spatial position information of the feature points of the current scene image can be obtained by triangulating the vehicle key frame pose information, the feature point coordinates of the current scene image obtained through the semantic segmentation network model, and the system external parameters, realizing the association between the feature points and the vehicle pose information.
[0080] Specifically, obtain the scene image captured by the front-view undistorted camera, scale the size of the image to 320*180, input the scaled image into the trained semantic segmentation network, and output an image mask. At this time, the obtained image mask only retains the valid region in the image, and the value of the valid region in the mask is 0, and the value of the dynamic region is 255.
[0081] Further extract image feature points using the image mask: when the image is the first frame, extract N s Harris feature points in the image, and the distance between the feature points is not less than d s , where N s is a hyperparameter of the number of feature points within the valid region of the preset front-view camera, and d s is the corresponding minimum pixel spacing hyperparameter of the feature points; when the image is not the first frame, use the LK optical flow method to track the feature point pixel coordinates of the previous frame in the current frame image, and preset the feature point number hyperparameter. If the number of tracked feature points m s < N s , then extract (N s - m s ) Harris feature points for the current image, while keeping the number of feature points in the current frame as Ns 。
[0082] For each frame after the initialization of the visual inertial odometer is completed, specifically including the vehicle pose information P and R of the key frame and the external parameters of the relative position between the front-view camera and the IMU, triangulation is performed to associate the vehicle pose information of the key frame with the coordinates of the image feature points observed by the front-view camera, and the spatial coordinates of the current image feature points are obtained.
[0083] S13: Perform loop detection on the semantic features in the most recently updated historical map that is loaded. If a loop is detected, perform global pose graph optimization on the driving trajectory of the vehicle according to the pose information and update the map.
[0084] In a specific embodiment, when the vehicle is first positioned, extract the static features of the key frame image obtained by the current front-view undistorted camera. The static features include the spatial coordinates of the key frame image feature points; extract the BRIEF descriptors corresponding to the feature points, convert the descriptors into word vectors of the bag-of-words model, and then perform similarity matching between the word vectors and the feature point word vectors of the historical frames in the historical map; determine whether the key frame obtained by the current front-view undistorted camera and the historical frames in the historical map form a loop according to a preset similarity threshold. If a loop is formed, perform global pose graph optimization; if not, there is no need to correct the global pose graph, thus completing the vehicle positioning.
[0085] Among them, the global pose graph optimization includes: According to the spatial coordinates of the feature points that have been matched in the two frames of images of the key frame obtained by the current front-view undistorted camera and the historical frames in the historical map, and the external parameters (rotation parameter R i , and translation parameter T i , i = 1, 2, 3, 4) between the corresponding cameras and the IMU, solve the change amount of the vehicle pose at the current moment relative to the historical moment, and perform global pose graph non-linear optimization by minimizing the pose residuals between the two frames of the loop and the pose residuals between all key frames in the global trajectory. That is, use the relative pose relationship of the loop to correct the global trajectory and eliminate the cumulative error of the visual inertial odometer.
[0086] In this embodiment, save the vehicle pose information and the observed image feature point word vectors corresponding to each key frame after the global pose graph optimization as a binary file, and construct a map according to the binary file. The map can be loaded for vehicle repositioning.
[0087] In the case of vehicle repositioning, load the historical map, match the associated image and the map data through loop detection. When the matching result forms a loop, correct the vehicle pose information according to the relative pose relationship between the associated image and the map data, and update the map according to the corrected vehicle pose information and the associated image.
[0088] Specifically, before the vehicle enters the parking lot again, load the most recently updated historical map, extract the BRIEF descriptors corresponding to the static feature points of the key frame images obtained by the current front-view undistorted camera, convert the BRIEF descriptors into word vectors of the bag-of-words model, and perform similarity matching between the word vectors and the feature point word vectors of the loaded map. Determine whether the current key frame and the map key frame form a loop according to a preset similarity threshold, and optimize the global pose graph according to the loop result, thereby completing vehicle relocalization.
[0089] The above embodiments can optimize the global pose graph through loop detection, eliminate the cumulative deviation of the visual inertial odometer. At the same time, compared with the prior art, the map constructed by the key frame poses and corresponding feature point information in the global driving trajectory of the vehicle in this embodiment has a higher matching success rate in vehicle relocalization.
[0090] As Figure 3 shown, another embodiment of the present invention further provides a vehicle positioning device, including a driving data acquisition unit 101, a semantic feature extraction unit 102, and a positioning unit 103.
[0091] The driving data acquisition unit 101 is used to acquire the driving scene images and pose information of the vehicle in the current driving state.
[0092] The semantic feature extraction unit 102 is used to extract the semantic features of the current driving scene image by using a semantic segmentation network, and the semantic features include scene features that are relatively stable both spatially and temporally.
[0093] The positioning unit 103 is used to perform loop detection on the semantic features in the most recently updated historical map loaded. If a loop is detected, optimize the global pose graph of the vehicle's driving trajectory according to the pose information and update the map.
[0094] Regarding the information interaction and execution process among the above units, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.
[0095] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the method described in any one of the above embodiments is implemented.
[0096] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0097] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A vehicle positioning method, characterized in that, Including: Obtain the driving scene image and pose information of the vehicle in the current driving state; Use a semantic segmentation network to extract the semantic features of the driving scene image, including inputting the driving scene image into the semantic segmentation network, identifying the static features, semi-static features, and dynamic features in the scene features of the driving scene image, and generating a mask for removing the semi-static features and the dynamic features; under the action of the mask, output the static features as semantic features, and the semantic features include scene features that are relatively stable both spatially and temporally; Perform loop detection on the semantic features in the most recently updated historical map loaded. If a loop is detected, globally optimize the driving trajectory of the vehicle according to the pose information and update the map; Among them, the driving scene image in the current driving state includes the key frame image obtained by the current front-view undistorted camera; The performing loop detection on the semantic features in the most recently updated historical map loaded includes: Extract the static features of the key frame image obtained by the current front-view undistorted camera, and the static features include the spatial coordinates of the feature points of the key frame image; Extract the BRIEF descriptors corresponding to the feature points, convert the descriptors into word vectors of the bag-of-words model, and then perform similarity matching between the word vectors and the feature point word vectors of the corresponding historical frames in the historical map; Determine whether the key frame obtained by the current front-view undistorted camera and the corresponding historical frame in the historical map form a loop according to a preset similarity threshold; The method for obtaining the pose information includes: Perform sparse edge feature extraction on the images obtained by four fisheye cameras, and the edge features include static identifiers of parked vehicles and the parking lot environment; When the obtained image is the first frame image, respectively extract N Harris feature points from the image, and the pixel distance between the N feature points is not less than d, where N is a preset hyperparameter of the number of feature points, and d is a hyperparameter of the minimum pixel spacing of the feature points; When the obtained image is not the first frame image, use the LK optical flow method to track the pixel coordinates of the feature points of the previous frame image in the current frame image. If the number of tracked feature points m < N, then extract (N - m) Harris feature points from the current frame image; Obtain the pixel coordinates of the feature points tracked between two frames of images, and the external parameters of the camera corresponding to the feature points. According to the epipolar geometry relationship, triangulate the pixel coordinates of the feature points and the external parameters, calculate the spatial coordinates of the feature points tracked in each image, and the pose information of the vehicle in the world coordinate system in the current frame image.
2. The vehicle positioning method according to claim 1, wherein, The pose information is output by the visual inertial odometer and is associated with the key frame of the driving scene image; among them, The driving scene image is obtained by the on-vehicle front-view camera; the visual inertial odometer includes the on-vehicle surround-view fisheye camera and the IMU.
3. The vehicle positioning method according to claim 2, wherein, It also includes: Extract the pose information corresponding to the current frame and the key frame of the scene image obtained from the on-vehicle surround-view fisheye camera, and perform image coding on the differences between the scene images according to the pose information; When the return result of the visual inertial odometer is normal, if the current frame is a key frame, store the image and pose information of the current frame; if the current frame is a non-key frame, discard the current frame and continue to read the next frame of image; When the return result of the visual inertial odometer is abnormal, match the key frame most similar to the current frame in the key frame set according to the image coding threshold, and re-calculate the visual inertial odometer. If there is no matching result, discard the current frame and continue to read the next frame of image.
4. The vehicle positioning method according to claim 3, wherein After the return result of the visual inertial odometer is normal and it is determined that the current frame is a key frame, it further includes: Align the visual information of the key frame and the IMU pre-integration result in a loose coupling manner, and calculate the visual feature point information and IMU state variable information through a non-linear method; Non-linearly optimize the feature point information and the IMU state variable information to obtain the pose information at the current moment corresponding to the scene image.
5. The vehicle positioning method according to claim 1, wherein, The global pose graph optimization and map update of the vehicle's driving trajectory according to the pose information includes: Solve the change amount of the vehicle pose relative to the historical moment at the current moment according to the space coordinates of the feature points matched by the key frame and the historical frame obtained by the current front view undistorted camera, and the external parameters between the corresponding camera and IMU, and perform global pose graph non-linear optimization and map update by minimizing the pose residuals between two loop closure frames and the pose residuals between all key frames in the global trajectory.
6. A vehicle positioning device, characterized in that, It includes: A driving data acquisition unit for acquiring the driving scene image and pose information of the vehicle in the current driving state; A semantic feature extraction unit for extracting the semantic features of the driving scene image by using a semantic segmentation network, including inputting the driving scene image into the semantic segmentation network, identifying the static features, semi-static features and dynamic features in the scene features of the driving scene image, and generating a mask for removing the semi-static features and the dynamic features; under the action of the mask, output the static features as semantic features, and the semantic features include scene features that are relatively stable both spatially and temporally; A positioning unit for performing loop closure detection on the semantic features in the recently updated historical map loaded. If a loop closure is detected, perform global pose graph optimization and map update on the vehicle's driving trajectory according to the pose information; Among them, the driving scene image in the current driving state includes the key frame image obtained by the current front view undistorted camera; The loop closure detection of the semantic features in the recently updated historical map loaded includes: Extract the static features of the key frame image obtained by the current front view undistorted camera, and the static features include the space coordinates of the key frame image feature points; Extract the BRIEF descriptor corresponding to the feature points, convert the descriptor into a word vector of the bag-of-words model, and then perform similarity matching between the word vector and the feature point word vector of the corresponding historical frame in the historical map; Determine whether the key frame obtained by the current front view undistorted camera and the corresponding historical frame in the historical map form a loop closure according to a preset similarity threshold; The method for obtaining the pose information includes: Extract sparse edge features from the images obtained by four fisheye cameras, where the edge features include parked vehicles and static signs of the parking lot environment; When the obtained image is the first frame image, extract N Harris feature points of the image respectively, and the pixel distance between the N feature points is not less than d, where N is a preset hyperparameter of the number of feature points, and d is a hyperparameter of the minimum pixel spacing of feature points; When the obtained image is not the first frame image, use the LK optical flow method to track the pixel coordinates of the feature points of the previous frame image in the current frame image. If the number of tracked feature points m < N, then extract (N - m) Harris feature points from the current frame image; Obtain the pixel coordinates of the feature points tracked between two frames of images, and the external parameters of the camera system corresponding to the feature points. According to the epipolar geometry relationship, triangulate the pixel coordinates of the feature points and the external parameters of the system, and calculate the spatial coordinates of the feature points tracked in each image, as well as the pose information of the vehicle in the world coordinate system in the current frame image.
7. The vehicle positioning device according to claim 6, characterized in that, The pose information is output by the visual inertial odometer and is associated with the key frames of the driving scene image; where The driving scene image is obtained by a front-view camera on the vehicle; the visual inertial odometer includes a surround-view fisheye camera and an IMU on the vehicle; The driving data acquisition unit is specifically used for: Extract the pose information corresponding to the current frame and the key frames of the scene image obtained from the surround-view fisheye camera on the vehicle, and encode the differences between the scene images according to the pose information; When the return result of the visual inertial odometer is normal, if the current frame is a key frame, store the image and pose information of the current frame. If the current frame is a non-key frame, discard the current frame and continue to read the next frame of image; When the return result of the visual inertial odometer is abnormal, match the key frame most similar to the current frame in the key frame set according to the image coding threshold, and recalculate the visual inertial odometer. If there is no matching result, discard the current frame and continue to read the next frame of image.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, where when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the vehicle positioning method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Positioning method and system based on visual inertial navigation information fusion
CN107869989A
Vehicle positioning and mapping method and device based on camera
CN110136199A
Monocular vision inertia SLAM method for dynamic scene
CN111156984A
Microminiature unmanned aerial vehicle visual navigation method in high dynamic scene
CN111693047A