Garage positioning method and device based on UWB and visual tight coupling

CN122544762APending Publication Date: 2026-08-11WUHAN JIANGXIA CHUNENG AUTOMOBILE TECHNOLOGY R&D CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]有鉴于此,有必要提供一种基于UWB与视觉紧耦合的车库定位方法及装置,用以解决现有的车库定位方案存在的动态遮挡和极端环境下鲁棒性不足,无法实现复杂车库环境下的稳定厘米级定位的技术问题

Benefits of technology

[0014] The beneficial effects of the above implementation are as follows: The garage positioning method and apparatus based on tight coupling of UWB and vision provided by the present invention determine UWB coarse positioning based on UWB base station network configuration files and metal multipath error feature library, as well as the measurement timing of UWB and the real-time motion data of IMU; based on visual map database files, the current frame image captured by the vehicle camera, the pre-integration result of IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning; the UWB coarse positioning and the visual positioning are input to the tight coupling unit for fusion positioning to obtain the fused pose of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122544762A_ABST
    Figure CN122544762A_ABST
Patent Text Reader

Abstract

The application provides a garage positioning method and device based on UWB and visual close coupling, and belongs to the field of intelligent driving.The method comprises the following steps: determining UWB rough positioning based on a UWB base station network configuration file and a metal multipath error feature library, and real-time motion data of UWB measurement timing and IMU; performing visual positioning and map matching based on a visual map database file, a current frame image captured by a vehicle-mounted camera, pre-integration results of IMU from the last frame to the current frame, and the UWB rough positioning, so as to obtain visual positioning; and inputting the UWB rough positioning and the visual positioning into a close coupling unit for fusion positioning, so as to obtain the fusion pose of the vehicle.The application can solve the problem that the existing garage positioning scheme has insufficient robustness under dynamic occlusion and extreme environment, and cannot realize stable centimeter-level positioning in a complex garage environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, specifically to a garage positioning method and device based on UWB and vision tight coupling. Background Technology

[0002] With the development of intelligent driving and automatic parking technologies, accurate positioning in indoor environments such as underground parking garages has become a key technological bottleneck. Traditional GPS (Global Positioning System) signals are completely ineffective indoors, and single-sensor positioning solutions have significant limitations: UWB (Ultra Wide Band) positioning is significantly affected by multipath effects and non-line-of-sight errors in metal-dense parking garage environments, with positioning accuracy deteriorating from centimeter-level to meter-level; visual positioning is prone to feature matching failures and scale drift under conditions such as uneven lighting and repetitive textures in parking garages. Existing fusion solutions mostly adopt loose coupling methods, failing to fully utilize the complementary characteristics between sensors, and lack robustness in dynamic occlusion and extreme environments, making it impossible to achieve stable centimeter-level positioning in complex parking garage environments. Summary of the Invention

[0003] In view of this, it is necessary to provide a garage positioning method and device based on UWB and vision tight coupling to solve the technical problems of existing garage positioning schemes, such as insufficient robustness in dynamic occlusion and extreme environments, and inability to achieve stable centimeter-level positioning in complex garage environments.

[0004] To address the aforementioned problems, in a first aspect, the present invention provides a garage positioning method based on UWB and vision tightly coupled, comprising: Based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, the coarse positioning of the UWB is determined. Based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning. The UWB coarse positioning and the visual positioning are input into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.

[0005] In one possible implementation, the garage localization method based on UWB and vision tightly coupled also includes: Images of the garage area and pose information of the vehicle-mounted camera are collected, and a dense 3D point cloud model of the garage is generated based on the images of the garage area and pose information of the vehicle-mounted camera. Structural and textural features are extracted from the dense 3D point cloud model; the structural features include planar features and linear features. The feature point indexes of the structural and texture features are stored to construct a visual map database file.

[0006] In one possible implementation, the garage localization method based on UWB and vision tightly coupled also includes: The garage environment information is imported into the ray tracing simulation software for coverage simulation and blind spot analysis. The constraint is to ensure that any target area maintains a line-of-sight connection with at least 3 UWB base stations. The objective function is to minimize the geometric dilution of the average positioning accuracy of all areas. The pose of the UWB base stations is optimized using the particle swarm optimization algorithm. Then, the coordinates of all UWB base stations are time-synchronized and calibrated to generate a UWB base station network configuration file.

[0007] In one possible implementation, the garage localization method based on UWB and vision tightly coupled also includes: Acquire channel impulse response data collected by a UWB channel sounder on the metal structure inside the garage; Based on the channel impulse response data, the statistical distribution of the potential error in ranging caused by metal multipath propagation is calculated. Based on the statistical distribution of the potential errors, a feature library of metal multipath errors is generated.

[0008] In one possible implementation, based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, a rough UWB localization is determined, including: A chi-square test is performed on the UWB measurement time series to remove outliers and obtain the verified measurement sequence. The verified measurement sequences are subjected to NLOS identification to determine base station links with suspected NLOS and base station links with suspected LOS. When the number of base station links in LOS is greater than or equal to a preset link number threshold, multiple hyperbolic equations are constructed using one base station in the LOS base station link as a reference. The initial position estimate of the LOS base station link is obtained by solving the hyperbolic equations. When the number of base station links in LOS is less than a preset link number threshold, the measurement sequences of suspected base station links in NLOS are assigned weights and the weighted least squares method is used to calculate the initial location estimate of the suspected base station links in NLOS. Based on the initial location estimation of base station links using LOS and the initial location estimation of suspected base station links using NLOS, as well as the metal multipath error feature library, error prediction results for base station link locations are obtained. Based on the error prediction results, the initial location estimation of the base station link in LOS and the initial location estimation of the suspected base station link in NLOS are offset to obtain the coarse UWB positioning, and the covariance matrix of the coarse UWB positioning and the health status of the UWB are determined.

[0009] In one possible implementation, visual localization is obtained by performing visual localization and map matching based on a visual map database file, the current frame image captured by the vehicle-mounted camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse localization, including: Feature extraction is performed on the current frame image captured by the vehicle-mounted camera to obtain image feature points; Based on the UWB coarse positioning, the visual map database file is queried to retrieve the map feature points of the area corresponding to the UWB coarse positioning. The image feature points are coarsely matched with the map feature points to form a preliminary set of matching pairs. Based on the preliminary matching pair set, the reprojection error between the current frame image of the vehicle camera and the visual map database file is calculated, and based on the reprojection error, mismatches that do not conform to geometric constraints are eliminated to obtain the interior point set. Based on the aforementioned set of interior points, and using the iterative nearest point algorithm to perform fine registration on the 3D point cloud, a more accurate vehicle camera pose can be obtained. Scale recovery is performed on the more accurate pose of the vehicle-mounted camera to obtain visual positioning, and the covariance matrix of the visual positioning and the health status of the vehicle-mounted camera are determined.

[0010] In one possible implementation, feature extraction is performed on the current frame image captured by the vehicle-mounted camera to obtain image feature points, including: A Gaussian pyramid is constructed for the current frame image captured by the vehicle-mounted camera, and the average brightness of the current frame image at each pyramid level is determined. When the average brightness of the current frame image is greater than a preset brightness threshold, the FAST corner detector combined with the BRIEF descriptor is used to extract and describe the features of the current frame image captured by the vehicle camera, and the running line segment detector is used to extract the long line segment features in the current frame image captured by the vehicle camera to obtain image feature points. When the average brightness of the current frame image is less than or equal to a preset brightness threshold, the SIFT feature detector and descriptor are used to extract and describe the features of the current frame image captured by the vehicle camera, and the running line segment detector is used to extract the long line segment features in the current frame image captured by the vehicle camera to obtain image feature points.

[0011] In one possible implementation, the UWB coarse positioning and the visual positioning are input into a tightly coupled unit for fusion positioning to obtain the vehicle's fused pose, including: The UWB coarse positioning and its covariance matrix, along with the UWB health status, the visual positioning and its covariance matrix, the vehicle camera health status, the IMU raw data, and the pre-integration results are spatiotemporally aligned and then input into a tightly coupled unit. The tightly coupled unit processes the data using an adaptive tightly coupled extended Kalman filter architecture to obtain the vehicle's fused pose.

[0012] In one possible implementation, the input parameters are spatiotemporally aligned using tightly coupled sensors, and then processed using an adaptive tightly coupled extended Kalman filter architecture to obtain the vehicle's fused pose, including: After the input parameters are spatiotemporally aligned by tightly coupled sensors, an adaptive tightly coupled extended Kalman filter architecture is used for processing to obtain the vehicle's preliminary fused pose. The preliminary fused pose is then subjected to low-pass filtering or fixed-hysteresis smoothing to obtain the final fused pose of the vehicle.

[0013] Secondly, the present invention also provides a garage positioning device based on UWB and vision tightly coupled, comprising: The UWB positioning module is used to determine a rough UWB positioning based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU. The visual positioning module is used to perform visual positioning and map matching based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, so as to obtain visual positioning. The fusion positioning module is used to input the UWB coarse positioning and the visual positioning into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.

[0014] The beneficial effects of the above implementation are as follows: The garage positioning method and apparatus based on tight coupling of UWB and vision provided by the present invention determine UWB coarse positioning based on UWB base station network configuration files and metal multipath error feature library, as well as the measurement timing of UWB and the real-time motion data of IMU; based on visual map database files, the current frame image captured by the vehicle camera, the pre-integration result of IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning; the UWB coarse positioning and the visual positioning are input to the tight coupling unit for fusion positioning to obtain the fused pose of the vehicle.

[0015] In typical garage environments containing complex metal structures, this invention achieves centimeter-level positioning accuracy by tightly coupling and fusing coarse UWB positioning and related data with visual positioning and related data. Compared to solutions relying on a single sensor, this significantly improves accuracy, providing a solid foundation for applications such as automatic parking and parking space-level navigation. Secondly, this invention incorporates a metal multipath error feature library during UWB positioning, comprehensively enhancing the system's environmental adaptability and robustness. The established metal multipath error feature library enables dynamic compensation for metal multipath interference, specifically suppressing signal interference caused by steel columns and pipes within the garage, reducing positioning errors in densely metal areas. Simultaneously, the visual layered matching mechanism, i.e., visual positioning and map matching, effectively addresses the challenges of varying lighting and texture repetition, ensuring stable operation in various harsh scenarios. Therefore, this invention solves the technical problems of existing garage positioning solutions, such as insufficient robustness in dynamic occlusion and extreme environments, which prevent stable centimeter-level positioning in complex garage environments. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of an embodiment of the garage positioning method tightly coupled with UWB and vision provided by the present invention; Figure 2 A schematic diagram of a system architecture for a garage positioning method tightly coupled with UWB and vision provided by the present invention; Figure 3 A flowchart of another embodiment of the garage positioning method tightly coupled with UWB and vision provided by the present invention; Figure 4 A schematic diagram of one embodiment of the garage positioning device tightly coupled with UWB and vision provided by the present invention; Figure 5 A schematic diagram of an embodiment of the electronic device provided by the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] In the description of the embodiments of this application, unless otherwise stated, "a plurality of" means two or more.

[0020] In this embodiment of the invention, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, apparatus, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product or device.

[0021] The naming or numbering of steps in the embodiments of the present invention does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved.

[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0023] This invention provides a garage positioning method and device based on UWB and vision tight coupling, which will be described below.

[0024] This invention provides a garage positioning method based on tight coupling of UWB and vision. The method can be implemented by executing an application program on an onboard controller, which can be a vehicle airspace controller, an intelligent driving domain controller, or a vehicle controller, etc. Figure 1 As shown, the method includes: S101. Based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, a rough UWB positioning is determined.

[0025] Understandably, UWB base stations are key infrastructure in Ultra-Wide Band (UWB) positioning systems. Serving as reference nodes for known locations, they communicate with UWB tags worn by people or objects, achieving centimeter-level positioning through high-precision ranging. UWB base stations receive nanosecond-level pulse signals from the tags and calculate the distance between the tag and the base station using methods such as Time-of-Flight (ToF), Time-Differentiation of Arrival (TDoA), or Two-Way Ranging (TWR).

[0026] The timing of UWB measurements includes: the raw time of flight (ToF) or time difference of arrival (TDOA) measurements received by the UWB tag from each base station.

[0027] The IMU (Inertial Measurement Unit) provides real-time motion data, including acceleration and angular velocity. The IMU can be installed on a vehicle.

[0028] The UWB base station network configuration file can be obtained by measuring parameters such as base station clock deviation and transmission power at known coordinate points using a vehicle-mounted mobile tag.

[0029] Metal multipath error refers to the phenomenon where wireless signals (such as GNSS, UWB, Wi-Fi, etc.) encounter strong reflections from metal surfaces during propagation, forming multiple propagation paths. This causes the receiver to simultaneously receive both direct and reflected signals, resulting in ranging or positioning errors.

[0030] S102. Based on the visual map database file, the current frame image captured by the vehicle-mounted camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning.

[0031] Understandably, matching images from an onboard camera with a pre-built visual map achieves high-precision pose estimation and addresses the scale blur problem inherent in monocular vision. Inputs include the current frame image, pre-integrated data from the inertial measurement unit, a visual map database, and a coarse position prior provided by UWB. Processing begins with adaptive image feature extraction. The system dynamically selects a feature extraction strategy based on the average image brightness: using fast ORB features to ensure real-time performance in well-lit conditions, and switching to more robust SIFT features in dim lighting. Simultaneously, for the structured environment of a garage, additional line segment features such as column edges are extracted to form a multimodal feature description.

[0032] S103. Input the UWB coarse positioning and the visual positioning into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.

[0033] Understandably, by tightly coupling UWB, vision, and inertial measurement unit (IMU) data, a stable, continuous, and high-precision final pose—the fused pose of the vehicle—is output. The inputs include the processed UWB and vision positioning results and their uncertainties, as well as the raw inertial data. Processing first involves temporal synchronization and spatial alignment of the multi-sensor data, ensuring all observations are unified to the same timestamp and vehicle coordinate system. Fusion is implemented within an adaptive extended Kalman filter framework.

[0034] In some embodiments, the garage positioning method based on UWB and vision tight coupling further includes: Images of the garage area and pose information of the vehicle-mounted camera are collected, and a dense 3D point cloud model of the garage is generated based on the images of the garage area and pose information of the vehicle-mounted camera. Structural and textural features are extracted from the dense 3D point cloud model; the structural features include planar features and linear features. The feature point indexes of the structural and texture features are stored to construct a visual map database file.

[0035] Understandably, a framework based on ORB-SLAM3 is used for visual odometry calculation and loop closure detection to generate a sparse point cloud map. Furthermore, a semi-global matching (SGM) algorithm is used to calculate the disparity map, and combined with camera pose, a dense 3D point cloud model of the parking garage is generated.

[0036] Structural features can be features such as the ground, ceiling, and walls, while linear features can be features such as the edge of a column.

[0037] Dense point clouds can more completely preserve the surface details and spatial structure of objects, and the extracted features can significantly enhance the ability to recognize complex terrain, obstacle shapes, and minute structures. Under extreme conditions such as strong light, shadow, and occlusion, dense point clouds combined with feature extraction can maintain stable environmental perception and reduce false positives and false negatives.

[0038] In some embodiments, the garage positioning method based on UWB and vision tight coupling further includes: The garage environment information is imported into the ray tracing simulation software for coverage simulation and blind spot analysis. The constraint is to ensure that any target area maintains a line-of-sight connection with at least 3 UWB base stations. The objective function is to minimize the geometric dilution of the average positioning accuracy of all areas. The pose of the UWB base stations is optimized using the particle swarm optimization algorithm. Then, the coordinates of all UWB base stations are time-synchronized and calibrated to generate a UWB base station network configuration file.

[0039] Understandably, garage environment information can be a garage BIM model or CAD drawings of the garage; coverage simulation and blind spot analysis involve importing the garage BIM model into ray tracing simulation software (such as Wireless InSite). A UWB signal model (such as an IEEE 802.15.4a channel model) is set up to simulate signal propagation in a simulated environment. By analyzing the received signal strength (RSS) and the power of the first-arrival path, areas where positioning accuracy may deteriorate (e.g., severe multipath propagation or direct path congestion) are identified.

[0040] Base station location and attitude optimization: With the constraint of maintaining line-of-sight (LOS) connectivity between any target area and at least three base stations, and minimizing the geometric dilution of average positioning accuracy (GDOP) across all areas as the objective function, the optimal 3D coordinates of the base station are solved using the Particle Swarm Optimization (PSO) algorithm. During base station installation, the orientation should avoid the main lobe direction directly facing a large metal surface; instead, it should be tilted downwards at 15-30 degrees to cover the vehicle height area.

[0041] System-level synchronization and coordinate calibration: All base stations achieve nanosecond-level time synchronization via wired Ethernet or a precise wireless clock synchronization protocol (such as 802.1AS). After deployment, a high-precision total station is used to measure the actual coordinates of the phase center of each base station antenna, which serve as the known coordinates for the system. Measurements are then taken at these known coordinate points using vehicle-mounted mobile tags to calibrate parameters such as base station clock deviation and transmit power, resulting in a base station configuration parameter file.

[0042] The objective function is to minimize the geometric dilution of the average positioning accuracy across all regions. After optimizing the pose of UWB base stations using the particle swarm optimization algorithm, the coordinates of all UWB base stations are then time-synchronized and calibrated. By simulating the cooperative behavior of the group, it can effectively avoid getting trapped in local optima. It is suitable for high-dimensional nonlinear pose optimization problems, with few parameters, high iteration efficiency, and can quickly approximate the optimal pose.

[0043] In some embodiments, the garage positioning method based on UWB and vision tight coupling further includes: Acquire channel impulse response data collected by a UWB channel sounder on the metal structure inside the garage; Based on the channel impulse response data, the statistical distribution of the potential error in ranging caused by metal multipath propagation is calculated. Based on the statistical distribution of the potential errors, a feature library of metal multipath errors is generated.

[0044] Understandably, channel impulse response (CIR) data is collected by selecting typical metal structures (such as square steel columns, circular pipes, and roller shutters) in a garage and using a UWB channel detector at different distances (0.5m, 1m, 2m) and angles from the metal surface.

[0045] Multipath feature parameterization: For each CIR data point, extract the amplitude A_f and delay τ_f of the first arrival path, as well as the relative delay Δτ and relative amplitude A_mp of the strongest multipath component. Calculate the statistical distribution (mean and variance) of the potential error ε_mp of multipath on ranging.

[0046] Establish an error prediction lookup table: Parameterize the above measurement results to generate a lookup table indexed by "metal type", "label-metal surface distance", and "incident angle", which is the metal multipath error feature library.

[0047] Traditional channel models are typically based on simplified environmental assumptions and struggle to fully reflect the complex multipath effects in the presence of metallic objects. The Metal Multipath Error Feature Library, by integrating channel impulse response data from actual measurements or high-fidelity simulations, can accurately characterize the unique propagation characteristics introduced by metallic reflectors.

[0048] In some embodiments, based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, a rough UWB localization is determined, including: A chi-square test is performed on the UWB measurement time series to remove outliers and obtain the verified measurement sequence. The verified measurement sequences are subjected to NLOS identification to determine base station links with suspected NLOS and base station links with suspected LOS. When the number of base station links in LOS is greater than or equal to a preset link number threshold, multiple hyperbolic equations are constructed using one base station in the LOS base station link as a reference. The initial position estimate of the LOS base station link is obtained by solving the hyperbolic equations. When the number of base station links in LOS is less than a preset link number threshold, the measurement sequences of suspected base station links in NLOS are assigned weights and the weighted least squares method is used to calculate the initial location estimate of the suspected base station links in NLOS. Based on the initial location estimation of base station links using LOS and the initial location estimation of suspected base station links using NLOS, as well as the metal multipath error feature library, error prediction results for base station link locations are obtained. Based on the error prediction results, the initial location estimation of the base station link in LOS and the initial location estimation of the suspected base station link in NLOS are offset to obtain the coarse UWB positioning, and the covariance matrix of the coarse UWB positioning and the health status of the UWB are determined.

[0049] Understandably, NLOS (Non-Line-of-Sight Identification) identification employs a joint identification method based on channel characteristics, analyzing the CIR (Channel Impulse Response) corresponding to each measurement value: calculating the ratio of first-path power to total received power (P_R), and the ratio of first-path arrival time to root mean square delay spread (T_R). Thresholds η_P and η_T are set; if P_R < η_P and T_R > η_T, the path is considered highly likely to be NLOS. Simultaneously, short-term motion consistency predicted by the IMU is used for auxiliary judgment: if the difference between UWB velocity measurement and IMU solution velocity consistently exceeds the threshold, the base station link is marked as suspicious for NLOS.

[0050] Line-of-sight (LOS) location calculation: When the number of base stations marked as LOS, N_LOS ≥ 4, the TDOA algorithm is used. Taking one of the base stations as a reference, N_LOS - 1 hyperbolic equations are constructed, and solved by iterative least squares method through Taylor series expansion to obtain the initial location estimate of the tag (x_uwb, y_uwb, z_uwb).

[0051] Non-line-of-sight (NLOS) suppression: When N_LOS < 4, measurements from NLOS base stations are introduced, but with a dynamic weight w_i assigned. The weight w_i is inversely proportional to the confidence level that the link is classified as NLOS and is correlated with historical measurement residuals. Weighted least squares (WLS) is used for calculation to suppress the contamination of positioning results by NLOS measurements.

[0052] This embodiment identifies suspicious NLOS base station links by comparing the number of LOS base station links with a preset link number threshold. Weights are assigned to the measurement sequences of the suspicious NLOS base station links, and weighted least squares method is used for calculation to suppress the contamination of positioning results by NLOS measurement values, which can further improve the reliability of UWB coarse positioning results.

[0053] In some embodiments, visual localization and map matching are performed based on a visual map database file, the current frame image captured by the vehicle-mounted camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse localization to obtain visual localization, including: Feature extraction is performed on the current frame image captured by the vehicle-mounted camera to obtain image feature points; Based on the UWB coarse positioning, the visual map database file is queried to retrieve the map feature points of the area corresponding to the UWB coarse positioning. The image feature points are coarsely matched with the map feature points to form a preliminary set of matching pairs. Based on the preliminary matching pair set, the reprojection error between the current frame image of the vehicle camera and the visual map database file is calculated, and based on the reprojection error, mismatches that do not conform to geometric constraints are eliminated to obtain the interior point set. Based on the aforementioned set of interior points, and using the iterative nearest point algorithm to perform fine registration on the 3D point cloud, a more accurate vehicle camera pose can be obtained. Scale recovery is performed on the more accurate pose of the vehicle-mounted camera to obtain visual positioning, and the covariance matrix of the visual positioning and the health status of the vehicle-mounted camera are determined.

[0054] It is understood that this embodiment employs a two-level feature matching process, including: Coarse matching: For each feature in the current image feature point set F_t, its descriptor is compared with the descriptors of map points in M_local using Hamming distance (ORB) or Euclidean distance (SIFT) to find the nearest and second nearest neighbors. If the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than a preset ratio (e.g., 0.8), the matching pair is accepted, forming a preliminary matching pair set C_rough.

[0055] Geometric verification and fine matching: Using the RANSAC algorithm, the fundamental matrix or homography matrix between the current frame and the map (depending on the motion model) is estimated based on C_rough, and the reprojection error is calculated. Mismatches that do not meet geometric constraints are eliminated, resulting in a set of inliers, C_inliers. If the number of inliers is sufficient, the Iterative Closest Point (ICP) algorithm is used to perform fine registration on the 3D point cloud, solving for a more accurate camera pose T_cam_map (rotation matrix R and translation vector t).

[0056] This embodiment employs a dual-layer matching mechanism of coarse and fine matching, which can further improve the accuracy of the vehicle-mounted camera pose. Then, the vehicle-mounted camera pose is scaled back, thereby solving the problem of monocular vision scale blur.

[0057] In some embodiments, feature extraction is performed on the current frame image captured by the vehicle-mounted camera to obtain image feature points, including: A Gaussian pyramid is constructed for the current frame image captured by the vehicle-mounted camera, and the average brightness of the current frame image at each pyramid level is determined. When the average brightness of the current frame image is greater than a preset brightness threshold, the FAST corner detector combined with the BRIEF descriptor is used to extract and describe the features of the current frame image captured by the vehicle camera, and the running line segment detector is used to extract the long line segment features in the current frame image captured by the vehicle camera to obtain image feature points. When the average brightness of the current frame image is less than or equal to a preset brightness threshold, the SIFT feature detector and descriptor are used to extract and describe the features of the current frame image captured by the vehicle camera, and the running line segment detector is used to extract the long line segment features in the current frame image captured by the vehicle camera to obtain image feature points.

[0058] Understandably, a Gaussian pyramid is constructed for the image I_t. At each pyramid level, a feature detector is dynamically selected based on the average brightness L_avg of the current image. If L_avg > threshold_th (good lighting), use the FAST corner detector combined with the BRIEF descriptor (i.e., ORB features) to pursue real-time performance.

[0059] If L_avg <= threshold_th (dim lighting), use SIFT feature detectors and descriptors to pursue robustness.

[0060] Simultaneously, the Line Segment Detector (LSD) extracts long line segment features to match structured column edges.

[0061] In this embodiment, a Gaussian pyramid is constructed on the image to extract feature points. By analyzing the image at different scales (resolutions), the Gaussian pyramid makes the extracted feature points (such as SIFT keypoints) robust to image scaling, rotation, and lighting changes, ensuring stable matching under different conditions. The pyramid structure can simultaneously preserve the global structure (low-level) and local details (high-level) of the image, avoiding missed or false detections caused by feature size mismatch at a single scale.

[0062] In some embodiments, the UWB coarse positioning and the visual positioning are input to a tightly coupled unit for fusion positioning to obtain the fused pose of the vehicle, including: The UWB coarse positioning and its covariance matrix, along with the UWB health status, the visual positioning and its covariance matrix, the vehicle camera health status, the IMU raw data, and the pre-integration results are spatiotemporally aligned and then input into a tightly coupled unit. The tightly coupled unit processes the data using an adaptive tightly coupled extended Kalman filter architecture to obtain the vehicle's fused pose.

[0063] Understandably, this refers to tightly coupled units, or a tightly coupled fusion method. In this method, raw data or low-level observations from multiple sensors are jointly optimized, rather than being solved independently first and then fused.

[0064] By integrating UWB, vision, and inertial measurement unit (IMU) data through tight coupling, a stable, continuous, and high-precision final pose is output. Inputs include the processed UWB and vision positioning results and their uncertainties, as well as the raw inertial data. Processing first performs time synchronization and spatial alignment of the multi-sensor data, ensuring all observations are unified to the same timestamp and vehicle coordinate system.

[0065] The process is implemented within an adaptive extended Kalman filter framework. The filter state vector incorporates vehicle position, velocity, attitude, and zero bias from the inertial sensors. State prediction is driven by the inertial measurement unit (IMU), which recursively derives the state and covariance for the next time step using its pre-integration results. During the observation update phase, UWB and visual data are tightly coupled: UWB provides raw distance observations to each base station, and its observation model directly depends on the filter's position state; visual data provides reprojection error observations of feature points. This tight coupling approach allows for more fundamental utilization of raw sensor information and flexible handling of situations where some sensors fail.

[0066] The system dynamically adjusts the noise covariance of each observation based on UWB link quality indicators and visual health status to achieve adaptive fusion. When a UWB link is determined to be non-line-of-sight or has few visually matching feature points, its observation noise is amplified, thus automatically reducing its weight in the fusion process. When visual tracking is completely lost, the filter can temporarily rely solely on UWB and inertial data to maintain positioning. After filtering and updating, the optimal state estimate is obtained. Finally, the pose output from the filter is smoothed before being output.

[0067] This embodiment uses an adaptive tightly coupled extended Kalman filter architecture for fusion positioning. Through tight coupling, multi-sensor information is directly fused at the original observation layer, avoiding information loss caused by independent calculation in loose coupling. The adaptive mechanism can automatically reduce the weight of abnormal observations and maintain positioning continuity.

[0068] In some embodiments, after spatiotemporally aligning the input parameters using tightly coupled sensors, an adaptive tightly coupled extended Kalman filter architecture is employed for processing to obtain the fused pose of the vehicle, including: After the input parameters are spatiotemporally aligned by tightly coupled sensors, an adaptive tightly coupled extended Kalman filter architecture is used for processing to obtain the vehicle's preliminary fused pose. The preliminary fused pose is then subjected to low-pass filtering or fixed-hysteresis smoothing to obtain the final fused pose of the vehicle.

[0069] Understandably, the preliminary fused pose is subjected to low-pass filtering or fixed-hysteresis smoothing to obtain the final fused pose of the vehicle, in order to suppress high-frequency noise and output the final stable and smooth vehicle position, velocity and attitude estimates.

[0070] In some embodiments, the present invention provides a garage centimeter-level positioning method and system based on UWB and vision tight coupling. By constructing a garage-specific visual map database, establishing a metal environment UWB error compensation model, and designing an adaptive tight coupling fusion framework, stable centimeter-level positioning in complex garage environments can be achieved.

[0071] refer to Figure 2 As shown, the system architecture is as follows: This system adopts a layered architecture design. The bottom layer consists of vehicle-mounted perception units (UWB tags, cameras, IMU), the middle layer is a data processing and fusion engine, and the upper layer is a positioning application service. The core innovation lies in the construction of a garage prior information database, which includes a visual map database, a UWB base station layout model, and a metal reflection characteristic database, providing environment-specific prior knowledge for each processing module.

[0072] refer to Figure 3 As shown, the technical implementation steps include: Step 1: Building and Deploying Prior Information for the Garage Environment enter: A building information model (BIM) of the garage or precise CAD drawings, including the geometric dimensions and spatial coordinates of all structural columns, walls, ventilation ducts, and fire protection facilities.

[0073] A list of building materials, especially the surface material, roughness, and electromagnetic properties of metal components (such as steel columns, fireproof roller shutters, and metal pipes).

[0074] The UWB base station pre-deployed coordinate point set contains the initial three-dimensional coordinates (x, y, z) of at least 4 non-coplanar locations.

[0075] The high-precision data acquisition vehicle captures a sequence of surround view images (covering at least all driving lanes and parking spaces) while driving in the garage, along with the corresponding high-precision trajectory data (with an accuracy better than 5 centimeters) recorded by the vehicle's GNSS / INS integrated navigation system, for visual map ground truth calibration.

[0076] Processing procedure: Offline construction of visual map database: Data Acquisition and Preprocessing: A data acquisition vehicle equipped with four fisheye cameras (190° field of view) traverses all areas of the garage at speeds below 10 km / h. Images from all cameras are simultaneously recorded (30 frames per second) and the high-precision pose of the integrated navigation system. The images undergo distortion correction, illumination normalization, and feature enhancement preprocessing.

[0077] Dense 3D Reconstruction: Visual odometry and loop closure detection are performed using an ORB-SLAM3-based framework to generate a sparse point cloud map. Furthermore, a disparity map is calculated using a semi-global matching (SGM) algorithm, and combined with camera pose, to generate a dense 3D point cloud model of the parking garage.

[0078] Hierarchical feature extraction and storage: Structural Feature Layer: From the dense point cloud, planes and cylinders are fitted using the RANSAC (Random Sample Consensus) algorithm to extract key planar features (such as ground, ceiling, and walls) and linear features (such as cylinder edges). For each planar feature, its plane equation coefficients, boundary polygons, and normal vectors are stored; for linear features, their endpoint coordinates and direction vectors are stored.

[0079] Texture Feature Layer: On sparse feature points, ORB features (256-bit descriptors for fast matching) and SIFT features (128-dimensional descriptors for robust matching) are computed in parallel. Simultaneously, for the garage environment, additional semantic features such as arrows, parking lines, and numbers on the ground are extracted. Each feature point stores its 3D coordinates, two feature descriptors, the normal vector of its plane (for viewpoint invariance compensation), and a list of keyframes that observed that point.

[0080] Map indexing and compression: The entire garage space is divided into a 1m x 1m x 1m voxel grid. Each voxel stores the index of the feature points belonging to that space. Product quantization (PQ) is performed on the feature descriptors to compress them, and a visual bag-of-words model is constructed to achieve fast scene retrieval in subsequent real-time matching.

[0081] Optimization and calibration of UWB base station networks: Coverage Simulation and Blind Spot Analysis: Import the garage BIM model into ray tracing simulation software (such as WirelessInSite). Set up a UWB signal model (such as an IEEE 802.15.4a channel model) to simulate signal propagation in a simulation environment. By analyzing the received signal strength (RSS) and the power of the first-arrival path, identify areas where positioning accuracy may deteriorate (e.g., severe multipath propagation or direct path congestion).

[0082] Base station location and attitude optimization: With the constraint of maintaining line-of-sight (LOS) connectivity between any target area and at least three base stations, and minimizing the geometric dilution of average positioning accuracy (GDOP) across all areas as the objective function, the optimal 3D coordinates of the base station are solved using the Particle Swarm Optimization (PSO) algorithm. During base station installation, the orientation should avoid the main lobe direction directly facing a large metal surface; instead, it should be tilted downwards at 15-30 degrees to cover the vehicle height area.

[0083] System-level synchronization and coordinate calibration: All base stations achieve nanosecond-level time synchronization via wired Ethernet or a precise wireless clock synchronization protocol (such as 802.1AS). After deployment, a high-precision total station is used to measure the actual coordinates of the phase center of each base station antenna, which serve as the known coordinates for the system. Measurements are then taken at these known coordinate points using vehicle-mounted mobile tags to calibrate parameters such as base station clock deviation and transmit power, resulting in a base station configuration parameter file.

[0084] Construction of Metal Environmental Feature Library and Error Model: Channel impulse response (CIR) acquisition: Select typical metal structures (such as square steel columns, circular pipes, and roller shutters) in the garage and use a UWB channel detector to collect a large amount of channel impulse response (CIR) data at different distances (0.5m, 1m, 2m) and angles from the metal surface.

[0085] Multipath feature parameterization: For each CIR data point, extract the amplitude A_f and delay τ_f of the first arrival path, as well as the relative delay Δτ and relative amplitude A_mp of the strongest multipath component. Calculate the statistical distribution (mean and variance) of the potential error ε_mp of multipath on ranging.

[0086] Establish an error prediction lookup table: Parameterize the above measurement results to generate a lookup table indexed by "metal type", "label-metal surface distance", and "incident angle". The entries in the table contain the expected value and uncertainty (variance) of the ranging error ε_mp in this scenario. This table will serve as prior knowledge for real-time error compensation in step two.

[0087] Output: Visual map database file: contains hierarchical feature data, voxel index, and visual vocabulary; UWB base station network configuration file: contains the precise three-dimensional coordinates, transmit power, synchronization parameters, and calibrated system deviation parameters of each base station; Metal multipath error feature library: a structured lookup table or a fitted error prediction function f(type, distance, angle).

[0088] Key points: Visual map construction must ensure absolute scale accuracy, which can be achieved using ground markers of known dimensions for scale correction. UWB base station deployment must consider the feasibility of power supply and network cabling. Metal error models must cover the effects of different surface treatments (e.g., smooth, corroded).

[0089] This step processes the raw UWB measurements and performs robust processing and dynamic error compensation for the complex environment of the garage. Inputs include arrival time information from each base station received by the UWB tag, inertial measurement unit (IMU) data, and the aforementioned metal error feature library. The processing begins with preprocessing of the measurements and non-line-of-sight (NLS) propagation identification. First, the measurement sequence undergoes consistency verification, eliminating abnormal jump values. Then, the channel impulse response characteristics are jointly analyzed with the motion consistency provided by the IMU to determine the line-of-sight or NLS state for each base station link. Specifically, by calculating channel characteristic parameters such as first-path power ratio and delay spread, and comparing them with thresholds, obstructed NLS propagation paths can be effectively identified.

[0090] In the robust location calculation stage, the system adaptively selects the algorithm based on the number of base stations identified as line-of-sight. When there are sufficient line-of-sight base stations, a hyperbolic positioning model based on time difference of arrival is used, and an iterative algorithm is employed to solve for the initial estimate of the tag location. When there are insufficient line-of-sight base stations, a weighted least squares method is introduced to assign lower confidence weights to the measurements of base stations identified as non-line-of-sight, thereby suppressing their contamination of the positioning results and ensuring the stability of the calculation in partially obstructed environments.

[0091] The core component is the dynamic compensation of metal multipath errors. Based on the current approximate location, the system queries prior environmental information to determine the type and spatial relationship of nearby metal objects. Using this relationship, it queries a metal multipath error feature library to obtain the expected ranging error values ​​and their uncertainties for each base station link. Subsequently, within an extended Kalman filter framework, the ranging deviation of each link is estimated as an augmented state. The predicted error values ​​are used to apply soft constraints to the deviation state, guiding the filter to quickly and accurately estimate the true multipath error, which is then subtracted in real time during state updates. This step ultimately outputs the compensated and robustly processed UWB location coordinates and their covariance matrix, along with the communication quality status of each link, providing highly reliable distance observation information with corrected system errors for subsequent fusion.

[0092] Step 2: Input for UWB coarse positioning and dynamic error compensation: enter: The raw time of flight (ToF) or time difference of arrival (TDOA) measurement τ_i received by the UWB tag from each base station; Real-time motion data (acceleration a, angular velocity ω) from the IMU.

[0093] The metal multipath error feature library and base station coordinates generated in step one.

[0094] Processing procedure: Preprocessing of measured values ​​and identification of NLOS: Data verification: Perform a chi-square test on the measurement sequence of each base station to remove outliers with obvious jumps.

[0095] NLOS Identification: A joint identification method based on channel characteristics is employed. The CIR corresponding to each measurement is analyzed: the ratio of first-path power to total received power (P_R) and the ratio of first-path arrival time to root mean square delay spread (T_R) are calculated. Thresholds η_P and η_T are set; if P_R < η_P and T_R > η_T, the path is considered highly likely to be NLOS. Simultaneously, short-term motion consistency predicted by the IMU is used for auxiliary judgment: if the difference between UWB velocity measurement and IMU resolution velocity consistently exceeds the threshold, the base station link is marked as potentially NLOS.

[0096] Robust positional solution: Line-of-sight (LOS) location calculation: When the number of base stations marked as LOS, N_LOS ≥ 4, the TDOA algorithm is used. Taking one of the base stations as a reference, N_LOS - 1 hyperbolic equations are constructed, and solved by iterative least squares method through Taylor series expansion to obtain the initial location estimate of the tag (x_uwb, y_uwb, z_uwb).

[0097] Non-line-of-sight (NLOS) suppression: When N_LOS < 4, measurements from NLOS base stations are introduced, but with a dynamic weight w_i assigned. The weight w_i is inversely proportional to the confidence level that the link is classified as NLOS and is correlated with historical measurement residuals. Weighted least squares (WLS) is used for calculation to suppress the contamination of positioning results by NLOS measurements.

[0098] Dynamic compensation for metal multipath error: Scene matching: Based on the rough location (x_uwb, y_uwb, z_uwb) calculated in the current step, query the prior information of the garage constructed in step one to determine whether there are any metal objects nearby and their type and location.

[0099] Error prediction: Based on the geometric relationship (distance, angle of incidence) between the tag and the nearby metal surface, query the metal multipath error feature library to obtain the expected ranging error ε_pred,i and its uncertainty σ_ε,i for each base station link under the current environment.

[0100] Compensation in state estimation: In the state vector of the Extended Kalman Filter (EKF), an estimated state b_i for the ranging bias of each base station is introduced. In the observation update step of the EKF, the UWB measurement model is modified to: z_i = ||pos -anchor_i|| + b_i + v_i, where v_i is the measurement noise. Simultaneously, the predicted error ε_pred,i and its uncertainty σ_ε,i are used as the basis for adjusting the pseudo-observation or process noise of the bias state b_i, guiding the filter to quickly estimate and compensate for the actual bias.

[0101] Output: The UWB positioning result after compensation and robustness processing is P_uwb = (x, y, z); The corresponding covariance matrix Σ_uwb reflects the uncertainty of the current UWB positioning (including geometric configuration, NLOS case and compensation residual error). Health status identifiers (LOS / NLOS / Unstable) for each base station link.

[0102] This step processes the raw UWB measurements and performs robust processing and dynamic error compensation for the complex environment of the garage. Inputs include arrival time information from each base station received by the UWB tag, inertial measurement unit (IMU) data, and the aforementioned metal error feature library. The processing begins with preprocessing of the measurements and non-line-of-sight (NLS) propagation identification. First, the measurement sequence undergoes consistency verification, eliminating abnormal jump values. Then, the channel impulse response characteristics are jointly analyzed with the motion consistency provided by the IMU to determine the line-of-sight or NLS state for each base station link. Specifically, by calculating channel characteristic parameters such as first-path power ratio and delay spread, and comparing them with thresholds, obstructed NLS propagation paths can be effectively identified.

[0103] In the robust location calculation stage, the system adaptively selects the algorithm based on the number of base stations identified as line-of-sight. When there are sufficient line-of-sight base stations, a hyperbolic positioning model based on time difference of arrival is used, and an iterative algorithm is employed to solve for the initial estimate of the tag location. When there are insufficient line-of-sight base stations, a weighted least squares method is introduced to assign lower confidence weights to the measurements of base stations identified as non-line-of-sight, thereby suppressing their contamination of the positioning results and ensuring the stability of the calculation in partially obstructed environments.

[0104] The core component is the dynamic compensation of metal multipath errors. Based on the current approximate location, the system queries prior environmental information to determine the type and spatial relationship of nearby metal objects. Using this relationship, it queries a metal multipath error feature library to obtain the expected ranging error values ​​and their uncertainties for each base station link. Subsequently, within an extended Kalman filter framework, the ranging deviation of each link is estimated as an augmented state. The predicted error values ​​are used to apply soft constraints to the deviation state, guiding the filter to quickly and accurately estimate the true multipath error, which is then subtracted in real time during state updates. This step ultimately outputs the compensated and robustly processed UWB location coordinates and their covariance matrix, along with the communication quality status of each link, providing highly reliable distance observation information with corrected system errors for subsequent fusion.

[0105] Step 3: Visual positioning and map matching enter: The current frame image I_t captured by the vehicle-mounted camera; The IMU's pre-integration results from the previous frame to the current frame (relative rotation ΔR, translation Δp); Visual map database; The rough UWB location P_uwb from step two (used for initial search and scale recovery).

[0106] Processing procedure: Image feature extraction and description: A Gaussian pyramid is constructed for the image I_t. At each pyramid level, a feature detector is dynamically selected based on the average brightness L_avg of the current image. If L_avg > threshold_th (good lighting), use the FAST corner detector combined with the BRIEF descriptor (i.e., ORB features) to pursue real-time performance. If L_avg <= threshold_th (dim lighting), use SIFT feature detectors and descriptors to pursue robustness; Simultaneously, the Line Segment Detector (LSD) extracts long line segment features to match structured column edges.

[0107] Fast map matching based on location priors: Location retrieval: Using the rough UWB location P_uwb provided in step two, query the voxel index of the visual map database to quickly retrieve all map feature points M_local within a certain radius (e.g., 5 meters) around P_uwb.

[0108] Two-level feature matching: Coarse matching: For each feature in the current image feature point set F_t, its descriptor is compared with the descriptors of map points in M_local using Hamming distance (ORB) or Euclidean distance (SIFT) to find the nearest and second nearest neighbors. If the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than a preset ratio (e.g., 0.8), the matching pair is accepted, forming a preliminary matching pair set C_rough.

[0109] Geometric verification and fine matching: Using the RANSAC algorithm, the fundamental matrix or homography matrix between the current frame and the map (depending on the motion model) is estimated based on C_rough, and the reprojection error is calculated. Mismatches that do not meet geometric constraints are eliminated, resulting in a set of inliers, C_inliers. If the number of inliers is sufficient, the Iterative Closest Point (ICP) algorithm is used to perform fine registration on the 3D point cloud, solving for a more accurate camera pose T_cam_map (rotation matrix R and translation vector t).

[0110] Scale restoration and local optimization: Monocular vision suffers from scale ambiguity. Using the absolute position P_uwb provided by UWB, the scale factor s = ||P_uwb - P_uwb_prev|| / ||t|| of the visually estimated translation vector t is calculated, where P_uwb_prev is the UWB position at the previous time step. Low-pass filtering of s yields a stable scale factor s_fused, which in turn gives the absolute scale visual pose P_vis = s_fused * t.

[0111] Construct a local sliding window containing the five most recent keyframes. Perform local bundle adjustment within the window to optimize the pose of these keyframes and the 3D coordinates of their observed map points, minimizing the reprojection error of all feature points. This step further optimizes the local consistency of visual localization.

[0112] Output: The visual localization result is P_vis = (x, y, z, roll, pitch, yaw); The uncertainty of the covariance matrix Σ_vis of the localization result mainly comes from the number of feature matches and the reprojection error; Visual system health status: Good (interior points > threshold), Degraded (few interior points), Lost (no valid match).

[0113] This step utilizes the matching of images from the vehicle-mounted camera with a pre-built visual map to achieve high relative accuracy pose estimation and resolve the scale blur problem inherent in monocular vision. Inputs include the current frame image, pre-integrated data from the inertial measurement unit, a visual map database, and a coarse position prior provided by UWB. Processing begins with adaptive image feature extraction. The system dynamically selects a feature extraction strategy based on the average image brightness: fast ORB features are used in well-lit conditions to ensure real-time performance; in dim lighting, the system switches to the more robust SIFT features. Simultaneously, for the structured environment of a garage, additional line segment features such as column edges are extracted to form a multimodal feature description.

[0114] Subsequently, fast map matching based on location priors is performed. Utilizing the coarse location provided by UWB, all candidate map feature points within a specific radius around this location are quickly retrieved from the visual map database, greatly narrowing the matching search range. The matching process consists of two levels: first, preliminary association is performed through descriptor distance calculation to find nearest and second nearest neighbor feature points; then, a random sampling consensus algorithm and geometric verification are used to eliminate erroneous matches, and the iterative nearest point algorithm is used to solve for the optimal rigid body transformation from the current camera coordinate system to the map coordinate system, thereby obtaining the camera's fine pose.

[0115] Since monocular vision cannot perceive absolute scale, scale restoration is necessary. The scale factor is calculated by comparing the relative translation estimated by visual estimation in two consecutive frames with the absolute displacement provided by UWB at the corresponding time. This scale factor is then filtered and smoothed to obtain a stable value, thereby restoring the visual relative pose to absolute scale. Finally, local bundle adjustment is performed within a sliding window containing several recent keyframes to optimize the pose of these frames and the coordinates of observed map points, minimizing reprojection error and further improving the local consistency and accuracy of visual localization. This step outputs a six-DOF visual localization result with absolute scale, the corresponding uncertainty covariance, and the health status of feature tracking.

[0116] Step 4: Tightly Coupled Sensor Fusion and Positioning Output enter: UWB positioning results P_uwb and its covariance Σ_uwb, link status; Visual localization results P_vis and their covariance Σ_vis, health status; Raw IMU data (acceleration a, angular velocity ω) and pre-integration results.

[0117] Processing procedure: Spatiotemporal alignment: Time synchronization: All sensor data is timestamped based on a hardware clock and synchronized to a unified master clock (such as PTP) via software or hardware. Before fusion, high-frequency data from the IMU is used for interpolation to unify UWB and visual data to the same time t_k. Spatial calibration: During system initialization, the extrinsic parameters (rotation R_C^B, t_C^B; R_I^B, t_I^B; R_U^B, t_U^B) of the camera, IMU, and UWB tag in the vehicle coordinate system B are accurately measured using a hand-eye calibration method. All sensor observations are transformed to the vehicle coordinate system B and then correlated with the world coordinate system W.

[0118] Adaptive tightly coupled extended Kalman filter (EKF): State vector definition: The state vector X contains 15 dimensions: X = [p^T, v^T, q^T, b_a^T, b_g^T]^T. Where p is the vehicle's position in the world frame (3-dimensional), v is the velocity (3-dimensional), q is the attitude quaternion (4-dimensional), b_a is the accelerometer bias (3-dimensional), and b_g is the gyroscope bias (3-dimensional).

[0119] State prediction (IMU driven): The state is predicted using the IMU pre-integration results ΔR, Δv, Δp between two frames [t_k-1, t_k]. The state transition equation (discrete form) is as follows: p_k = p_k-1 + v_k-1 Δt + 0.5 (R_k-1 (a_m - b_a) + g) Δt^2; v_k = v_k-1 + (R_k-1 (a_m - b_a) + g) Δt; q_k = q_k-1 q(ω_m - b_g); b_a,k = b_a,k-1 (modeled as a random walk); b_g,k = b_g,k-1 (modeled as a random walk); Simultaneously, based on the IMU noise parameters, the predicted state covariance matrix P_k|k-1 is recursively calculated.

[0120] Observational Updates (UWB and Visual): UWB Observation Model: When the UWB ranging value z_uwb,i of the i-th base station is available, the observation equation is: h_uwb,i(X) = ||p + R * t_U^B - anchor_i|| + b_uwb,i, where b_uwb,i is the ranging bias of the link estimated in step two. Calculate the observation residual y_uwb,i = z_uwb,i - h_uwb,i(X) and its Jacobian matrix H_uwb,i. Determine the observation noise R_uwb,i based on Σ_uwb and the link state, and perform EKF standard updates.

[0121] Visual observation model: When the visual pose P_vis and covariance Σ_vis are available, they can be used as direct observations of the position p and orientation q in the state vector. The observation equations are: h_vis,p(X) = p, h_vis,q(X) = q. Calculate the residual y_vis = [P_vis.position; P_vis.orientation] - h_vis(X) and its Jacobian matrix H_vis. The observation noise R_vis is determined by Σ_vis. Perform EKF updates.

[0122] Adaptive fusion strategy: Dynamically adjust the weights of sensors in the fusion process based on their health status (i.e., adjust the observation noise R). If a UWB link is marked as NLOS, significantly increase its R_uwb,i and decrease its weight. If the visual status is Degraded, increase R_vis; if it is Lost, skip the current visual update.

[0123] Exception handling and smooth output: When both UWB and vision are unavailable for an extended period of time, the system degrades to pure inertial navigation and issues an external alarm.

[0124] The fused pose output from EKF is low-pass filtered or smoothed with fixed hysteresis to suppress high-frequency noise, and finally outputs a stable and smooth vehicle position, velocity and attitude estimate X_fused.

[0125] Output: The fused high-precision, high-frequency pose result: X_fused = (position, velocity, attitude); The covariance matrix Σ_fused of the fusion result represents the positioning accuracy; System operating status: Normal (multi-sensor fusion), Caution (partial sensor degradation), Degraded (mainly dependent on IMU).

[0126] This step is the core fusion process, integrating UWB, vision, and inertial measurement unit (IMU) data through tight coupling to output a stable, continuous, and high-precision final pose. Inputs include the processed UWB and vision positioning results and their uncertainties, as well as the raw inertial data. Processing first involves temporal synchronization and spatial alignment of the multi-sensor data, ensuring all observations are unified to the same timestamp and vehicle coordinate system.

[0127] The process is implemented within an adaptive extended Kalman filter framework. The filter state vector incorporates vehicle position, velocity, attitude, and zero bias from the inertial sensors. State prediction is driven by the inertial measurement unit (IMU), which recursively derives the state and covariance for the next time step using its pre-integration results. During the observation update phase, UWB and visual data are tightly coupled: UWB provides raw distance observations to each base station, and its observation model directly depends on the filter's position state; visual data provides reprojection error observations of feature points. This tight coupling approach allows for more fundamental utilization of raw sensor information and flexible handling of situations where some sensors fail.

[0128] The system dynamically adjusts the noise covariance of each observation based on UWB link quality indicators and visual health status to achieve adaptive fusion. When a UWB link is determined to be non-line-of-sight or has few visually matching feature points, its observation noise is amplified, thus automatically reducing its weight in the fusion process. When visual tracking is completely lost, the filter can temporarily rely solely on UWB and inertial data to maintain positioning. After filtering and updating, the optimal state estimate is obtained. Finally, the pose output by the filter is smoothed and then output. This step continuously outputs vehicle position, velocity, and attitude with centimeter-level accuracy. Its covariance matrix accurately reflects the reliability of the current positioning results. The system also outputs the overall positioning health status, providing a reliable basis for upper-level autonomous driving applications.

[0129] Processing Flow Description: This flow employs a dual-path parallel processing architecture. The UWB processing chain and the vision processing chain run independently before converging in the fusion engine. The fusion engine intelligently selects either EKF or factor graph optimization mode based on the sensor status, ensuring that the system can still maintain usable positioning even when some sensors fail. The output stage includes quality assessment and degradation mechanisms, forming a complete closed loop.

[0130] The present invention has the following beneficial effects: The technical solution provided by this invention brings a series of significant benefits. First, positioning accuracy achieves a substantial leap. In typical garage environments containing complex metal structures, through tight coupling and fusion of UWB and vision, an average positioning accuracy of centimeter level (e.g., 8.2 cm) can be achieved. Compared to solutions relying on a single sensor, the accuracy improvement can reach over 60%, providing a solid foundation for applications such as automatic parking and parking space-level navigation. Second, the system's environmental adaptability and robustness are comprehensively enhanced. The established metal multipath dynamic compensation model can specifically suppress signal interference caused by steel columns, pipes, etc., in garages, reducing positioning errors in densely metal areas by more than 60%. At the same time, the visual layered matching mechanism effectively addresses the challenges of lighting changes and texture repetition, ensuring stable operation in various harsh scenarios. Furthermore, the system achieves excellent fault tolerance through a dual-mode fusion framework and dynamic weight allocation. Even in extreme cases where some sensors temporarily fail or experience performance degradation, the system can still maintain continuous and reliable positioning output, improving overall availability to over 99.7%. Finally, the solution balances performance and deployment cost. The optimized UWB base station layout and the incrementally updatable visual map strategy reduce hardware deployment and subsequent maintenance costs by about 20% while ensuring performance, and have good prospects for commercial application.

[0131] The key inventive points of this invention are embodied in a series of closely related core technologies. The primary inventive point is the construction of a UWB multipath dynamic compensation model specifically for the characteristics of metal environments in parking garages. This model, through a pre-established metal reflection characteristic library, can predict and compensate for multipath ranging errors caused by specific geometries and surface materials in real time, improving the usability of UWB data from the source. Second, a hierarchical visual map construction and adaptive matching mechanism is designed. By fusing structural and texture features and dynamically switching feature extraction strategies according to lighting conditions, the matching success rate and positioning reliability of the vision system in low-texture and unevenly lit environments like parking garages are significantly improved. Third, an innovative dual-mode fusion framework with tight coupling between UWB, vision, and IMU is proposed. This framework uses extended Kalman filtering for efficient fusion under normal conditions and seamlessly switches to factor graph optimization mode when sensor anomalies occur, ensuring optimal state estimation and output continuity of the positioning system under any operating condition. Fourth, a dynamic weight allocation algorithm based on sensor confidence is embedded in the aforementioned fusion framework. This algorithm can quantitatively evaluate the quality of data from each sensor in real time and adaptively adjust its contribution weight in the final fusion result accordingly. This intelligently determines whether to trust absolute position information dominated by UWB or relative pose information dominated by vision, achieving optimal fusion of multi-source information. These inventions collectively constitute a complete innovation chain from environmental perception, error correction, data fusion to decision output.

[0132] like Figure 4As shown, the present invention also provides a garage positioning device 400 based on UWB and vision tight coupling, comprising: UWB positioning module 401 is used to determine coarse UWB positioning based on UWB base station network configuration file and metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU. The visual positioning module 402 is used to perform visual positioning and map matching based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, to obtain visual positioning. The fusion positioning module 403 is used to input the UWB coarse positioning and the visual positioning into the tightly coupled unit for fusion positioning to obtain the fused pose of the vehicle.

[0133] The garage positioning device based on UWB and vision tight coupling provided in the above embodiments can realize the technical solutions described in the above embodiments of the garage positioning method based on UWB and vision tight coupling. The specific implementation principles of each module or unit can be found in the corresponding content in the above embodiments of the garage positioning method based on UWB and vision tight coupling, which will not be repeated here.

[0134] like Figure 5 As shown, the present invention also provides an electronic device 500. The electronic device 500 includes a processor 501, a memory 502, and a display 503. Figure 5 Only some components of the electronic device 500 are shown, but it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead.

[0135] In some embodiments, memory 502 may be an internal storage unit of electronic device 500, such as a hard disk or memory of electronic device 500. In other embodiments, memory 502 may also be an external storage device of electronic device 500, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 500.

[0136] Furthermore, the memory 502 may include both internal storage units of the electronic device 500 and external storage devices. The memory 502 is used to store application software and various types of data installed on the electronic device 500.

[0137] In some embodiments, processor 501 may be a central processing unit (CPU), microprocessor, or other data processing chip for running program code stored in memory 502 or processing data, such as the garage positioning method based on UWB and vision tight coupling in this invention.

[0138] In some embodiments, display 503 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 503 is used to display information from electronic device 500 and to display a visual user interface. Components 501-503 of electronic device 500 communicate with each other via a system bus.

[0139] In some embodiments of the present invention, when processor 501 executes a UWB-based, vision-tightly coupled garage positioning program in memory 502, the following steps can be implemented: Based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, the coarse positioning of the UWB is determined. Based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning. The UWB coarse positioning and the visual positioning are input into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.

[0140] It should be understood that when the processor 501 executes the UWB-based and vision-coupled garage positioning program in the memory 502, in addition to the functions mentioned above, it can also perform other functions, as detailed in the description of the corresponding method embodiments above.

[0141] Furthermore, the embodiments of the present invention do not specifically limit the type of electronic device 500 mentioned. Electronic device 500 can be a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, laptop computer, or other portable electronic device. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic device can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, electronic device 500 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).

[0142] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the garage positioning method based on UWB and tightly coupled vision provided by the methods described above, the method comprising: Based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, the coarse positioning of the UWB is determined. Based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning. The UWB coarse positioning and the visual positioning are input into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.

[0143] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0144] The above provides a detailed description of the garage positioning method and device based on UWB and vision tight coupling provided by the present invention. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of ​​the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A garage positioning method based on UWB and visual tight coupling, characterized in that, include: Based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU, the coarse positioning of the UWB is determined. Based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning. The UWB coarse positioning and the visual positioning are input into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.

2. The UWB and vision tightly coupled based garage positioning method according to claim 1, wherein, Also includes: Images of the garage area and pose information of the vehicle-mounted camera are collected, and a dense 3D point cloud model of the garage is generated based on the images of the garage area and pose information of the vehicle-mounted camera. Structural and textural features are extracted from the dense 3D point cloud model; the structural features include planar features and linear features. The feature point indexes of the structural and texture features are stored to construct a visual map database file. 3.The UWB and vision tightly coupled based garage positioning method according to claim 1, wherein, Also includes: The garage environment information is imported into the ray tracing simulation software for coverage simulation and blind spot analysis. The constraint is to ensure that any target area maintains a line-of-sight connection with at least 3 UWB base stations. The objective function is to minimize the geometric dilution of the average positioning accuracy of all areas. The pose of the UWB base stations is optimized using the particle swarm optimization algorithm. Then, the coordinates of all UWB base stations are time-synchronized and calibrated to generate a UWB base station network configuration file.

4. The UWB and vision tightly coupled based garage positioning method according to claim 1, wherein, Also includes: Acquire channel impulse response data collected by a UWB channel sounder on the metal structure inside the garage; Based on the channel impulse response data, the statistical distribution of the potential error in ranging caused by metal multipath propagation is calculated. Based on the statistical distribution of the potential errors, a feature library of metal multipath errors is generated.

5. The UWB and vision tightly coupled based garage positioning method according to claim 1, wherein, Based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement time series of the UWB and the real-time motion data of the IMU, a rough UWB localization is determined, including: A chi-square test is performed on the UWB measurement time series to remove outliers and obtain the verified measurement sequence. The verified measurement sequences are subjected to NLOS identification to determine base station links with suspected NLOS and base station links with suspected LOS. When the number of base station links in LOS is greater than or equal to a preset link number threshold, multiple hyperbolic equations are constructed using one base station in the LOS base station link as a reference. The initial position estimate of the LOS base station link is obtained by solving the hyperbolic equations. When the number of base station links in LOS is less than a preset link number threshold, the measurement sequences of suspected base station links in NLOS are assigned weights and the weighted least squares method is used to calculate the initial location estimate of the suspected base station links in NLOS. Based on the initial location estimation of base station links using LOS and the initial location estimation of suspected base station links using NLOS, as well as the metal multipath error feature library, error prediction results for base station link locations are obtained. Based on the error prediction results, the initial location estimation of the base station link in LOS and the initial location estimation of the suspected base station link in NLOS are offset to obtain the coarse UWB positioning, and the covariance matrix of the coarse UWB positioning and the health status of the UWB are determined.

6. The garage positioning method based on UWB and vision tight coupling according to any one of claims 1-5, characterized in that, Based on the visual map database file, the current frame image captured by the vehicle-mounted camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, visual positioning and map matching are performed to obtain visual positioning, including: Feature extraction is performed on the current frame image captured by the vehicle-mounted camera to obtain image feature points; Based on the UWB coarse positioning, the visual map database file is queried to retrieve the map feature points of the area corresponding to the UWB coarse positioning. The image feature points are coarsely matched with the map feature points to form a preliminary set of matching pairs. Based on the preliminary matching pair set, the reprojection error between the current frame image of the vehicle camera and the visual map database file is calculated, and based on the reprojection error, mismatches that do not conform to geometric constraints are eliminated to obtain the interior point set. Based on the aforementioned set of interior points, and using the iterative nearest point algorithm to perform fine registration on the 3D point cloud, a more accurate vehicle camera pose can be obtained. Scale recovery is performed on the more accurate pose of the vehicle-mounted camera to obtain visual positioning, and the covariance matrix of the visual positioning and the health status of the vehicle-mounted camera are determined.

7. The garage positioning method based on UWB and vision tight coupling according to claim 6, characterized in that, Feature extraction is performed on the current frame image captured by the vehicle-mounted camera to obtain image feature points, including: A Gaussian pyramid is constructed for the current frame image captured by the vehicle-mounted camera, and the average brightness of the current frame image at each pyramid level is determined. When the average brightness of the current frame image is greater than a preset brightness threshold, the FAST corner detector combined with the BRIEF descriptor is used to extract and describe the features of the current frame image captured by the vehicle camera, and the running line segment detector is used to extract the long line segment features in the current frame image captured by the vehicle camera to obtain image feature points. When the average brightness of the current frame image is less than or equal to a preset brightness threshold, the SIFT feature detector and descriptor are used to extract and describe the features of the current frame image captured by the vehicle camera, and the running line segment detector is used to extract the long line segment features in the current frame image captured by the vehicle camera to obtain image feature points. 8.The UWB and vision tightly coupled based garage positioning method of claim 1, wherein, The UWB coarse positioning and the visual positioning are input into a tightly coupled unit for fusion positioning to obtain the vehicle's fused pose, including: The UWB coarse positioning and its covariance matrix, along with the UWB health status, the visual positioning and its covariance matrix, the vehicle camera health status, the IMU raw data, and the pre-integration results are spatiotemporally aligned and then input into a tightly coupled unit. The tightly coupled unit processes the data using an adaptive tightly coupled extended Kalman filter architecture to obtain the vehicle's fused pose.

9. The UWB and vision tightly coupled based garage positioning method according to claim 8, characterized in that, After spatiotemporally aligning the input parameters using tightly coupled sensors, an adaptive tightly coupled extended Kalman filter architecture is employed for processing to obtain the vehicle's fused pose, including: After the input parameters are spatiotemporally aligned by tightly coupled sensors, an adaptive tightly coupled extended Kalman filter architecture is used for processing to obtain the vehicle's preliminary fused pose. The preliminary fused pose is then subjected to low-pass filtering or fixed-hysteresis smoothing to obtain the final fused pose of the vehicle.

10. A garage positioning device based on UWB and visual tight coupling, characterized in that, include: The UWB positioning module is used to determine a rough UWB positioning based on the UWB base station network configuration file and the metal multipath error feature library, as well as the measurement timing of the UWB and the real-time motion data of the IMU. The visual positioning module is used to perform visual positioning and map matching based on the visual map database file, the current frame image captured by the vehicle camera, the pre-integration result of the IMU from the previous frame to the current frame, and the UWB coarse positioning, so as to obtain visual positioning. The fusion positioning module is used to input the UWB coarse positioning and the visual positioning into the tightly coupled unit for fusion positioning to obtain the vehicle's fused pose.