A multi-view fusion SLAM parking positioning and mapping method for underground parking lot
By employing a multi-view fusion SLAM method, combining a forward-looking camera, four fisheye surround-view cameras, and an IMU, the problem of insufficient positioning and mapping accuracy in underground parking lots was solved. Stable pose estimation and semantic map construction of parking spaces were achieved, enhancing the system engineering value of automated parking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-21
AI Technical Summary
Problems such as missing GNSS signals, repetitive textures, uneven lighting, and frequent occlusion in underground parking environments lead to insufficient accuracy in vehicle positioning and mapping.
A multi-view fusion SLAM method is adopted, which combines information from a forward-looking camera, four fisheye surround-view cameras, IMU inertial measurement, and parking space semantic information. Under a unified optimization framework, high-precision vehicle localization and parking space semantic map construction are carried out. Through semantic validity scoring, cross-view semantic association, semantic weighted fusion, and semantic loop triggering mechanism, more robust parking localization and mapping are achieved.
Stable pose estimation and parking space semantic map construction were achieved in underground parking lots, reducing the accumulation of positioning drift. The output map contains geometric trajectory and parking space semantic attributes, making it suitable for automatic parking planning and parking space management.
Smart Images

Figure CN122435565A_ABST
Abstract
Description
Technical Field
[0001] This invention pertains to the fields of autonomous driving technology, computer vision, and SLAM. It is primarily applied to autonomous parking and localization of vehicles in underground parking lots, specifically involving a multi-view fusion SLAM parking localization and mapping method for underground parking lots. Background Technology
[0002] In autonomous driving and autonomous parking (AVP) applications, vehicles need to stably complete environmental perception, pose estimation, parking space identification, and path planning within a limited space. Underground parking lots, due to characteristics such as missing GNSS signals, uneven lighting conditions, frequent wall and pillar obstructions, strong ground reflections, high repetition of parking space lines, and weak local textures, are prone to problems such as initialization difficulties, increased cumulative drift, loop closure instability, and missing map semantics in their localization and mapping methods, which are often based on monocular vision, binocular vision, or a single laser sensor.
[0003] Existing visual SLAM methods typically focus on the extraction and matching of geometric feature points. However, in underground parking garages, targets with more structured semantic meaning, such as parking space corners, directional arrows, and column edges, are often more suitable as long-term stable localization constraints than random texture points. Meanwhile, the vehicle's forward-looking camera has strong long-range observation capabilities, the surround-view fisheye camera has a large field of view coverage of the vehicle's near-field area, and the IMU can provide high-frequency continuous motion priors. If these three elements can be spatiotemporally and uniformly modeled and dedicated constraints oriented towards parking space semantics can be introduced, it is expected that localization and mapping performance in underground parking garages can be significantly improved while controlling costs.
[0004] Therefore, it is necessary to propose a multi-view fusion SLAM parking localization and mapping method for underground parking lots, enabling forward-looking geometric observation, surrounding-looking semantic observation, and inertial motion observation to work collaboratively within a unified optimization framework. Through original semantic validity scoring, cross-view semantic association, semantic weighted fusion, semantic confidence updating, and semantic loop triggering mechanisms, this method achieves more robust parking localization and more practically valuable parking space semantic map construction. Summary of the Invention
[0005] This invention aims to address the insufficient accuracy of vehicle localization and mapping in underground parking environments caused by issues such as missing GNSS signals, repetitive textures, uneven lighting, and frequent occlusion. This invention proposes a multi-view fusion SLAM parking localization and mapping method. By fusing visual observations from a forward-looking camera, visual observations from four fisheye surround-view cameras, IMU inertial measurement information, and parking space semantic information, high-precision vehicle localization and parking space semantic map construction are achieved within a unified optimization framework. The technical means employed in this invention are as follows: A multi-view fusion SLAM parking localization and mapping method for underground parking lots includes the following steps: S1: Acquire data from the vehicle's forward-looking camera, four-eye fisheye surround-view camera, and IMU inertial data, and perform timestamp alignment and joint calibration of spatial extrinsic parameters; S2: Perform distortion correction on the front view image and the four-eye fisheye image, extract image point features and parking space semantic features, and perform feature matching and tracking between multi-view frames; S3: Based on IMU pre-integration results, multi-view visual reprojection error and parking space semantic constraint error, a joint cost function is constructed and local sliding window optimization is performed to output high-frequency local vehicle pose. S4: Based on multi-view local pose and semantic features, keyframe selection, loop closure detection and global pose graph optimization are performed to generate a global parking space semantic map and global optimized pose for underground parking lots.
[0006] Furthermore, in S1, the first... Original timestamps of each sensor Perform time compensation to obtain a synchronization timestamp. ; and based on the extrinsic parameter matrix camera coordinate system The three-dimensional points below Transform to IMU coordinate system I, satisfying ;in, For the first The time offset of each sensor relative to a reference clock and These are the rotation matrix and the translation vector, respectively.
[0007] Furthermore, in S2, when performing distortion correction on the front view image and the four-eye fisheye image: based on the fisheye projection function... For the Distortion modeling is performed using observations from a fisheye camera, and the image points satisfy... ; and through homography matrix Projecting effective semantic points onto the bird's-eye view plane satisfies ,in As a scale factor, These are the pixel coordinates of the original image. The coordinates are for the bird's-eye view.
[0008] Furthermore, semantic features of parking spaces are extracted, and an effectiveness score is established for the semantic detection results of parking spaces. ,satisfy ,in For the first Classification confidence of individual parking space targets For geometric consistency error, For timing stability error, , , For weighting coefficients; when Not less than the association threshold The semantic result for the parking space is retained if it is valid; otherwise, it is discarded.
[0009] Furthermore, the joint association between visual point features and parking space semantic features established across multiple viewpoint frames incurs a significant cost. ,satisfy ,in , The center position of the candidate semantic target. , Oriented towards the target The overlap rate of the candidate boxes. , Score semantic validity; when Less than the association threshold At that time, candidate targets and They are identified as the same semantic landmark.
[0010] Furthermore, for the IMU in adjacent keyframes and Pre-integration is performed between these parameters, and the rotation, velocity, and position increments satisfy the following conditions:
[0011]
[0012]
[0013] in and These are the measured values of angular velocity and acceleration, respectively. and These are the gyroscope bias and the accelerometer bias, respectively. This is the gravity vector.
[0014] Furthermore, the local joint cost function constructed by S3 is: ;in To marginalize prior residuals, For IMU pre-integration residuals, For multi-view visual reprojection residuals, For parking space semantic residuals, For robust kernel functions, These are the semantic constraint weights.
[0015] Furthermore, the parking space semantic residual It consists of the difference between the observed and predicted corner points of the parking space in the bird's-eye view plane, satisfying ; and construct fusion weights based on the semantic validity scores from each perspective. This allows us to obtain the global semantic landmark location. ,in To prevent positive numbers with a denominator of zero.
[0016] Furthermore, in S4, a loop closure score is established. ,in For keyframe appearance similarity, The positional difference between candidate keyframes. For parking space semantic similarity, , , For weighting coefficients; when Not less than the threshold When this occurs, a loop closure detection is triggered and a loop closure constraint is generated.
[0017] Furthermore, in S4, the objective function of the global pose graph is minimized. This achieves joint optimization of odometer edges and loopback edges; and employs a semantic confidence recursive formula. Update # Map confidence of semantic landmarks, when Not less than the threshold It is then written into the global parking space semantic map.
[0018] This invention also provides a multi-view fusion SLAM parking localization and mapping method for underground parking lots. This method fully integrates forward-looking long-range vision, surrounding-looking near-range vision, and IMU continuous motion information, maintaining stable pose estimation even in turning, occlusion, and sparse texture areas of underground parking lots. By directly incorporating structured semantic information such as parking space corners, parking space edges, and available parking space status into the local and global graph optimization processes, it can significantly reduce the drift accumulation caused by relying solely on random texture points. The invention proposes semantic validity scoring, cross-view semantic association cost, semantic weighted fusion formula, semantic loop closure scoring formula, and semantic confidence update formula, enabling semantic information to have a quantifiable, optimizable, filterable, and long-term maintainable mathematical expression. The output map not only contains geometric trajectory information but also parking space semantic attributes that can directly serve automatic parking planning and parking space management, improving the practical value of the system engineering. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the overall technical route of the method of the present invention.
[0021] Figure 2 This is a schematic diagram of the vision-inertial tightly coupled sliding window optimization model in the method of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] like Figure 1 As shown, this invention provides a multi-view fusion SLAM parking localization and mapping method for underground parking lots. In this embodiment, the vehicle is equipped with one forward-looking camera, four fisheye surround-view cameras, and one IMU. Let the world coordinate system be... The IMU coordinate system is , No. Each camera coordinate system is .in, Indicates a forward-facing camera. to These represent the front, rear, left, and right fisheye surround-view cameras, respectively. The system state vector can be represented as: (1) in, , , Representing time respectively attitude, position and speed , These represent the gyroscope bias and accelerometer bias, respectively. This represents the set of semantic landmarks that are participating in the optimization within the current sliding window.
[0025] Specifically, it includes the following steps: Step S1: Multi-source sensor data acquisition and spatiotemporal synchronization In step S1, images from each camera and IMU measurements are first acquired. Because different sensors have different hardware trigger frequencies and bus transmission delays, a unified reference clock needs to be established, and time offset compensation is applied to each sensor channel. The first sensor The synchronization timestamp of frame data is defined as follows: (2) in, This is the original timestamp. For sensors The time offset relative to a reference clock. This is achieved by... Offline calibration or online fine-tuning can unify multi-source data onto the same timeline.
[0026] After time synchronization is completed, the spatial extrinsic parameters between the camera and the IMU are jointly calibrated. The transformation relationship of 3D points from the camera coordinate system to the IMU coordinate system is as follows: (3) The corresponding homogeneous transformation matrix can be expressed as: .in, For rotation matrix, This is the translation vector. The above relationship ensures that geometric and semantic points observed by different cameras can be uniformly mapped to the IMU reference frame, providing a foundation for subsequent multi-view fusion.
[0027] Step S2: Multi-view visual feature and semantic extraction For a four-eye fisheye camera, the first step is to establish a fisheye imaging model. Let the spatial point be in the camera coordinate system as... Its normalization perspective is After considering the radial distortion of the fisheye lens, the distortion angle satisfy: (4) in , , , This represents the fisheye distortion coefficient. From this, the pixel coordinate projection function can be obtained. This enables accurate distortion modeling and distortion correction of fisheye images.
[0028] To unify ground semantic information from different fisheye perspectives, this invention projects effective parking space corner points, parking space edge midpoints, or other ground semantic features onto the bird's-eye view plane, satisfying the following: (5) in, For the first Homography matrix of a fisheye camera to the local ground plane of a vehicle. As a scale factor, These are the pixel coordinates of the original image. This corresponds to the coordinates of the bird's-eye view. Through this inverse perspective mapping, parking space information from multiple perspectives can be uniformly expressed in the same planar coordinate system.
[0029] Regarding the candidate parking space results output by the parking space detection network, this invention does not directly incorporate all of them into subsequent optimization. Instead, it proposes a semantic validity scoring formula to jointly evaluate classification confidence, geometric rationality, and temporal stability. (6) in, For the first The semantic validity score of each candidate parking space. To test the classification confidence of the network output, To account for geometric consistency errors such as parking space length, width, angle, and symmetry. This represents the time-series discrepancy between current observations and historical tracking results. , , These are non-negative weighting coefficients. When... Less than the threshold If the candidate parking space is not reliable enough, it should be eliminated before optimization.
[0030] To achieve joint tracking of visual point features and semantic features between adjacent frames and different viewpoints, a cross-viewpoint semantic association cost is constructed: (7) in, , The center coordinates of the two candidate semantic targets, , For the target orientation angle, The overlap rate of candidate regions. , Score the corresponding semantic validity. to For weight parameters. When Less than the threshold At that time, it was determined that the two belonged to the same parking space semantic landmark, thereby establishing a stable semantic observation sequence.
[0031] Step S3: Vision-Inertial-Semantic Tightly Coupled Odometry Estimation Since the IMU frequency is typically higher than the camera frequency, this invention employs a pre-integration method to compress inertial measurement information between adjacent keyframes. Let adjacent keyframes be... and Then the pre-integral quantities for rotation, velocity, and position are respectively: (8) (9) (10) in, and They are time points Angular velocity and acceleration measurements, and These are the corresponding biases. The gravity vector is used to construct the IMU residuals, which, together with the visual reprojection residuals and semantic observation residuals, are then optimized via a sliding window.
[0032] The local joint cost function constructed in this invention is as follows: (11) in, To marginalize prior residuals, For IMU pre-integration residuals, For multi-view visual reprojection error, For parking space semantic residuals, , , , These are the corresponding information matrices. For robust kernel functions, Equation (11) represents the semantic constraint weights. Compared with the traditional method that only optimizes geometric points, Equation (11) directly incorporates the semantic observation of parking spaces into the state estimation process, which is beneficial to improving the constraint stability of weak texture regions.
[0033] Furthermore, the semantic residual of parking spaces is defined using the bird's-eye view difference between the observed corner point and the predicted corner point: (12) in, Indicates the first The first parking space The actual observed values of each corner point This represents the predicted corner position based on the current state. This definition facilitates a direct measurement of the consistency of parking space geometry within a local map.
[0034] For the same semantic landmark observed from multiple perspectives simultaneously, a semantic validity score is used to construct the fusion weight: (13) (14) in, semantic landmarks In the Effectiveness rating from multiple perspectives To prevent positive numbers with a denominator of zero, This represents the inverse projection function, which projects pixel coordinates back onto a 3D point. This represents the position of the fused semantic landmark in the world coordinate system. This weighted fusion method allows high-confidence perspectives to contribute more to the estimation of semantic landmark positions.
[0035] V. Step S4: Semantic Global Mapping and Optimization of Parking Spaces After completing local pose estimation, this invention performs global maintenance on keyframes and semantic landmarks. To improve the accuracy of loop closure detection, a loop closure scoring method combining appearance, spatial, and semantic information is proposed. (15) in, For keyframes and The similarity in appearance between them The difference in displacement between the two. For parking space semantic similarity, For distance scale factor, , , For weighting coefficients. When Greater than the loopback threshold When a valid loop is detected, a loop edge is added to the global pose graph.
[0036] The objective function for global pose graph optimization is defined as: (16) in, Denotes the odometer edge set, Denotes the set of cyclic edges. This indicates a partial odometer measurement. Indicates closure constraint measurement. and Represents the information matrix. Let be the mapping operator from Lie group to Lie algebra. By optimizing equation (16), the cumulative drift caused by long-distance vehicle travel can be eliminated, making the global trajectory consistent with the map.
[0037] To mitigate the long-term impact of occasional false detections on map quality, this invention further proposes a recursive update formula for semantic landmark confidence: (17) in, For the first The semantic landmark in the first Map confidence level after the latest update. For update rate, Score the semantic validity corresponding to the current observation. Greater than the threshold When a landmark is added to the global parking space semantic map, it is removed from the map maintenance set if its value remains below the removal threshold for an extended period. This mechanism enables the system to maintain a high-quality parking space semantic map over the long term.
[0038] Workflow of the Implementation Example In one specific embodiment, after the vehicle enters the underground parking lot, the forward-looking camera continuously provides geometric observations of the forward passage, intersections, and distant obstacle areas, while the four-eye fisheye surround-view camera continuously provides observations of parking space lines, parking space corners, and surrounding ground markings in the near-field area of the vehicle. The IMU provides high-frequency angular velocity and acceleration measurements. The system first performs spatiotemporal alignment and extrinsic parameter transformation corresponding to equations (2) and (3), then performs distortion correction, inverse perspective mapping, and parking space semantic extraction on the fisheye images, and uses equations (6) and (7) to filter reliable semantic observations and establish cross-view correlations.
[0039] Subsequently, the system jointly optimizes the vehicle pose, velocity, offset, and semantic landmark position within a sliding window according to equations (8) to (14). When the vehicle passes through a mapped area, equation (15) is used to determine whether a loop closure is triggered, and global pose map optimization is performed according to equation (16) after the loop closure is triggered. Finally, the global parking space semantic map is maintained according to equation (17). The output results include the vehicle's global pose, local / global trajectory, parking space corner position, parking space edge parameters, parking space availability status, and its corresponding confidence level.
[0040] Compared to existing methods, this embodiment exhibits higher pose continuity and lower loop closure mismatch rate in underground parking lot turning areas, ramp entrances and exits, column-occluded areas, and weakly textured areas. Furthermore, since the semantic landmarks of parking spaces directly participate in the optimization, the output map is more suitable for automatic parking trajectory planning, vacant parking space retrieval, and parking guidance.
[0041] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-view fusion SLAM parking localization and mapping method for underground parking lots, characterized in that... include: S1: Acquire data from the vehicle's forward-looking camera, four-eye fisheye surround-view camera, and IMU inertial data, and perform timestamp alignment and joint calibration of spatial extrinsic parameters; S2: Perform distortion correction on the front view image and the four-eye fisheye image, extract image point features and parking space semantic features, and perform feature matching and tracking between multi-view frames; S3: Based on IMU pre-integration results, multi-view visual reprojection error and parking space semantic constraint error, a joint cost function is constructed and local sliding window optimization is performed to output high-frequency local vehicle pose. S4: Based on multi-view local pose and semantic features, keyframe selection, loop closure detection and global pose graph optimization are performed to generate a global parking space semantic map and global optimized pose for underground parking lots.
2. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 1, characterized in that: S1 for the first Original timestamps of each sensor Perform time compensation to obtain a synchronization timestamp. ; and based on the extrinsic parameter matrix camera coordinate system The three-dimensional points below Transform to IMU coordinate system I, satisfying ;in, For the first The time offset of each sensor relative to a reference clock and These are the rotation matrix and the translation vector, respectively.
3. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 1, characterized in that: When performing distortion correction on the front view image and the four-eye fisheye image in S2: based on the fisheye projection function. For the Distortion modeling is performed using observations from a fisheye camera, and the image points satisfy... ; and through homography matrix Projecting effective semantic points onto the bird's-eye view plane satisfies ,in As a scale factor, These are the pixel coordinates of the original image. The coordinates are for the bird's-eye view.
4. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 1, characterized in that: Extract semantic features of parking spaces and establish an effectiveness score for the semantic detection results of parking spaces. ,satisfy ,in For the first Classification confidence of individual parking space targets For geometric consistency error, For timing stability error, , , For weighting coefficients; when Not less than the associated threshold The semantic result for the parking space is retained if it is valid; otherwise, it is discarded.
5. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 4, characterized in that: The cost of establishing a joint association between visual point features and parking space semantic features across multiple viewpoint frames is significant. ,satisfy ,in , The center position of the candidate semantic target. , Oriented towards the target The overlap rate of the candidate boxes. , Score semantic validity; when Less than the association threshold At that time, candidate targets and They are identified as the same semantic landmark.
6. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 4, characterized in that: For IMU in adjacent keyframes and Pre-integration is performed between these parameters, and the rotation, velocity, and position increments satisfy the following conditions: in and These are the measured values of angular velocity and acceleration, respectively. and These are the gyroscope bias and the accelerometer bias, respectively. This is the gravity vector.
7. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 4, characterized in that: The local joint cost function constructed by S3 is as follows: ;in To marginalize prior residuals, For IMU pre-integration residuals, For multi-view visual reprojection residuals, For parking space semantic residuals, For robust kernel functions, These are the semantic constraint weights.
8. The multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 4, characterized in that: The semantic residual of the parking space It consists of the difference between the observed and predicted corner points of the parking space in the bird's-eye view plane, satisfying ; and construct fusion weights based on the semantic validity scores from each perspective. This allows us to obtain the global semantic landmark location. ,in To prevent positive numbers with a denominator of zero.
9. A multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 4, characterized in that: In S4, a loop closure score is established. ,in For keyframe appearance similarity, The positional difference between candidate keyframes For parking space semantic similarity, , , For weighting coefficients; when Not less than the threshold When this occurs, a loop closure detection is triggered and a loop closure constraint is generated.
10. A multi-view fusion SLAM parking localization and mapping method for underground parking lots according to claim 4, characterized in that: In S4, the objective function of the global pose graph is minimized. This enables joint optimization of the odometer edge and the loopback edge. And a semantic confidence recursive formula is used. Update # Map confidence of semantic landmarks, when Not less than the threshold It is then written into the global parking space semantic map.