A monocular vision vehicle positioning system fusing uwb ranging information and method thereof

By integrating UWB ranging information into a monocular vision vehicle positioning system, and utilizing modules for measurement preprocessing, monocular vision odometry, initialization, non-line-of-sight recognition and elimination, and factor graph optimization, the system solves the problems of scale inconsistency and scale drift in monocular vision odometry, achieving high-precision global positioning in complex environments and reducing the number of UWB base stations.

CN116804557BActive Publication Date: 2026-07-24TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2023-06-01
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Monocular visual odometry cannot directly obtain the actual depth information of an object, resulting in an unobservable scale, scale drift, and an inability to estimate global pose. At the same time, UWB ranging accuracy decreases in complex environments, and occlusion leads to non-line-of-sight effects, increasing the cost of use.

Method used

A monocular vision vehicle localization system integrating UWB ranging information constructs a measurement preprocessing module, a monocular vision odometry module, an initialization module, a non-line-of-sight recognition and elimination module, and a factor graph optimization module by combining monocular vision and UWB information. The measurement preprocessing module receives images and UWB ranging information, and the factor graph optimization module estimates the global pose of the vehicle.

Benefits of technology

It can accurately estimate the global pose of a vehicle in real time under complex working conditions, solve the scale uncertainty and scale drift problems of monocular vision odometry, reduce the number of UWB base stations, improve positioning accuracy and real-time performance, and avoid the influence of non-line-of-sight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116804557B_ABST
    Figure CN116804557B_ABST
Patent Text Reader

Abstract

The application relates to a monocular vision vehicle positioning system fusing UWB ranging information and a method thereof, the system comprising: a measurement preprocessing module for carrying out distortion removal and photometric correction on a monocular camera original image, and interpolating and acquiring a synchronous distance measurement according to UWB label ranging information; a monocular vision odometer module for incrementally estimating a camera pose and constructing two kinds of visual relative pose measurements; a non-line-of-sight identification and elimination module for eliminating non-line-of-sight measurements from the synchronous distance measurement; an initialization module for estimating an initial global heading angle and a visual odometer initial scale; and a factor graph optimization module combining output data of the monocular vision odometer module, the non-line-of-sight identification and elimination module and the initialization module to estimate vehicle global pose information. Compared with the prior art, the application can simultaneously solve the problems existing in the monocular vision odometer, greatly reduce the required number of UWB base stations, and accurately estimate the vehicle global pose in real time under various complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and in particular to a monocular vision vehicle positioning system and method that integrates UWB ranging information. Background Technology

[0002] With the advancement of automotive intelligence and electrification, autonomous driving has become one of the mainstream trends in future automotive development. Vehicle localization is a key technology in the field of autonomous driving. As the input to downstream algorithms such as path planning and trajectory tracking, it directly affects the completion of autonomous vehicle driving tasks. Among the many localization methods currently available, monocular vision odometry localization, which uses temporal images from a monocular camera as input, has attracted widespread attention due to its simplicity, flexibility, low cost, and low power consumption.

[0003] However, several challenges still need to be addressed before monocular vision odometry can be practically applied to vehicle positioning:

[0004] Since the actual depth information of an object cannot be obtained directly from a monocular image, monocular visual odometry can only determine the camera's direction of motion and relative distance, which means there is a problem of scale obscurity.

[0005] As the vehicle moves and the scene changes, the problem of monocular scale obscurity will lead to inconsistency in the scale of monocular visual odometry, i.e., the scale drift problem.

[0006] Monocular visual odometry achieves continuous localization by constantly calculating the relative pose between adjacent frames. This incremental calculation principle will cause the localization error to accumulate continuously as the vehicle travels. Furthermore, the pose estimated by monocular visual odometry is relative to the pose of the first frame, rather than the pose of the vehicle in a global coordinate system. In other words, monocular visual odometry can only estimate the relative pose and cannot know the global pose.

[0007] Furthermore, existing technologies consider that UWB (Ultra-Wideband) ranging technology is a short-range ranging technology based on radio waves, which can achieve high-precision distance measurement. Currently, this technology is widely used in indoor positioning, autonomous driving, smart homes and other fields, and has advantages such as low power consumption and high reliability. Therefore, UWB ranging technology can be applied to the positioning of autonomous vehicles. However, in practical applications, non-line-of-sight effects and multipath effects caused by occlusion will seriously reduce the accuracy of UWB ranging and make it impossible to accurately estimate the global pose of the vehicle under complex conditions. In order to effectively use UWB positioning in complex environments with a lot of occlusion, it is often necessary to carefully design the layout scheme of UWB base stations according to the actual scenario to ensure that there are a sufficient number of line-of-sight base stations in different areas and to switch the base stations according to the area. This will inevitably greatly increase the actual cost of using UWB positioning technology. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art by providing a monocular vision vehicle positioning system and method that integrates UWB ranging information. By fully combining monocular vision positioning information and UWB ranging information, it can simultaneously solve the problems of monocular vision odometry and greatly reduce the number of UWB base stations required. It can also accurately estimate the global pose of the vehicle in real time under various complex working conditions.

[0009] The objective of this invention can be achieved through the following technical solution: a monocular vision vehicle positioning system integrating UWB ranging information, comprising a measurement preprocessing module, a monocular vision odometer module, an initialization module, a non-line-of-sight recognition and elimination module, and a factor graph optimization module. The measurement preprocessing module is connected to the monocular vision odometer module, the non-line-of-sight recognition and elimination module, and the initialization module, respectively. The monocular vision odometer module is connected to the initialization module and the non-line-of-sight recognition and elimination module, respectively. The monocular vision odometer module, the non-line-of-sight recognition and elimination module, and the initialization module are all connected to the factor graph optimization module.

[0010] The measurement preprocessing module is used to perform distortion correction and photometric correction on the raw images acquired by the monocular camera, and to interpolate and obtain the synchronized distance measurement based on the UWB tag ranging information.

[0011] The monocular visual odometry module is used to incrementally estimate the camera pose and construct two visual relative pose measurements based on the monocular camera image after distortion correction and photometric correction.

[0012] The non-line-of-sight identification and rejection module is used to filter the distance measurements after synchronization in order to reject non-line-of-sight measurements.

[0013] The initialization module is used to estimate the initial global heading angle and the initial visual odometry scale based on the distance measurement and camera pose estimation results after synchronization.

[0014] The factor graph optimization module is used to combine the output data of the monocular vision odometry module, the non-line-of-sight recognition and elimination module, and the initialization module to estimate the global pose information of the vehicle.

[0015] A monocular vision vehicle localization method incorporating UWB ranging information includes the following steps:

[0016] S1. Acquire the raw images captured by the monocular camera and the ranging information from the UWB tag;

[0017] S2. Perform distortion correction and photometric correction on the monocular camera images sequentially;

[0018] Based on the ranging information of UWB tags, the quality of the corresponding distance measurement after synchronization is determined by judging the quality of the corresponding distance measurement before synchronization.

[0019] S3. Based on the monocular camera images after distortion correction and photometric correction, the camera pose is estimated incrementally using a sliding window optimization method. The sliding window optimization results are used to construct two visual relative pose measurements.

[0020] S4. Perform non-line-of-sight identification and rejection processing on the distance measurement after synchronization;

[0021] S5. Combine the distance measurement after synchronization and the camera pose estimation results to estimate the initial global heading angle and the initial scale of the visual odometry.

[0022] S6. Construct an initial pose factor map using the initial global heading angle and the initial visual odometry scale. Construct visual relative pose constraints using two types of visual relative pose measurements. Combine distance measurements eliminated by non-line-of-sight and the set historical prior constraints. Use a method of alternating distance noise model estimation and factor map optimization to estimate the key frame frequency pose, which is the global pose of the vehicle.

[0023] Furthermore, the specific process of step S2 is as follows:

[0024] S21. For monocular camera images, the image generation process is simulated using both a geometric camera model and a photometric camera model. The geometric camera model adopts a traditional pinhole camera model. The camera intrinsic parameter matrix and distortion coefficients are obtained through camera intrinsic parameter calibration, and radial and tangential distortions of the original image are removed.

[0025] The photometric camera model is modeled using a nonlinear response function, a nonparametric vignetting plot, and exposure time.

[0026] Photometric correction is performed on the distortion-removed image based on the camera photometric calibration results;

[0027] S22. For the ranging message of the UWB tag, linear interpolation is used to obtain the distance measurement synchronized with the image frame timestamp based on the timestamp. First, based on the signal strength information in the original message, the quality of the corresponding distance measurement before synchronization is determined. When the quality of the nearest distance measurement on both sides of the image frame timestamp is good, the quality of the corresponding distance measurement after synchronization is set to good; otherwise, the quality of the corresponding distance measurement after synchronization is set to bad.

[0028] Furthermore, in step S22, the quality of the corresponding distance measurement before synchronization is specifically judged using the following formula:

[0029]

[0030] Where rx_rssi and fp_rssi are the total received signal strength indicator and the first path signal strength indicator, respectively, θ thrs To set the filtering threshold.

[0031] Furthermore, the two types of visual relative pose measurement in step S3 are specifically adjacent relative pose measurement and co-visual relative pose measurement.

[0032] Furthermore, the specific process of step S3 is as follows:

[0033] S31. First, the distortion-corrected and photometric-corrected monocular camera image is divided into key frames and non-key frames. Only key frames can be added to the sliding window and participate in the sliding window optimization. For key frames, their pose and the depth of their high gradient points will undergo multiple sliding window optimizations until they are marginalized out of the sliding window.

[0034] For non-key frames, align them directly with the latest key frame to obtain the pose.

[0035] S32. Subsequently, two visual relative pose measurements were constructed using a visual odometry sliding window. After each sliding window optimization, the latest keyframe pose T was calculated. k and the new keyframe pose T k-1 The relative transformation between them is called adjacent relative pose measurement;

[0036] Before each edge-mapping keyframe, calculate the pose T of the latest keyframe. k and the oldest keyframe pose T k-7 The relative transformation between them is called common-view relative pose measurement;

[0037] Wherein, any two frames pose T i and T j Relative pose between The calculation formula is:

[0038]

[0039] In the formula, and Representing the frames F obtained by sliding window optimization estimation respectively i and frame F j The pose, since the pose defined in the subsequent factor map optimization is in the carrier coordinate system, is obtained through the extrinsic parameters between the carrier coordinate system and the camera coordinate system. Relative transformation between camera frames This is further transformed into a relative transformation between carrier frames.

[0040]

[0041] In the formula, This refers to the relative pose between camera frames.

[0042] Furthermore, step S4 involves further screening of UWB distance measurements with good quality after synchronization. The specific process includes:

[0043] Combined with the latest adjacent relative pose transformation of monocular visual odometry pose at the previous time step in the sum factor graph Predict the vehicle's attitude and position at the current moment:

[0044]

[0045]

[0046] Where R and t represent the rotation matrix and translation vector corresponding to the pose matrix T, respectively. This indicates that the translation vector is at a monocular scale, and it needs to be multiplied by the scale to be transformed to the true scale. The current time scale is used for prediction.

[0047] Based on this, we can further predict the current UWB tag position.

[0048]

[0049] in, The external parameters between the pre-calibrated UWB tags and the camera;

[0050] Finally, for all distance measurements at the current moment, the distance between the predicted tag location and the corresponding UWB base station is calculated. If the measured distance and the predicted distance exceed the set threshold, the measurement is determined to be a non-line-of-sight measurement and is directly removed.

[0051] Furthermore, the specific process of step S5 is as follows:

[0052] First, the position of the UWB tag at the corresponding time is determined by minimizing the error between the distance measured from the vehicle-mounted UWB tag to the known location UWB base station and the actual distance measurement.

[0053]

[0054] in, and These represent the UWB label and the base station with sequence number i in world frame F, respectively. w In the global location, Ω is the set of base station indices with good distance measurement quality after synchronization at the corresponding time, and r iThe measured UWB distance from the tag to the base station with serial number i is assumed to be that the vehicle moves in the same horizontal plane during the initialization phase, and the z coordinate of the tag is fixed to the known height.

[0055] Then, before the visual odometry window edge-out for the first time, the co-view relative pose between the latest and oldest keyframes is calculated, ignoring the vehicle's vertical displacement, and the initial global yaw angle is calculated. init and initial scale s init They are respectively:

[0056]

[0057]

[0058] in, The UWB translation vector is a visual translation vector composed of the vehicle's lateral and longitudinal displacements. This represents the difference between the x and y coordinates of the UWB label at the corresponding time.

[0059] Furthermore, in step S6, the factor graph contains two types of state nodes, six types of factors, and a distance noise model. The two types of state nodes include pose state nodes and scale state nodes. The pose state node is the global pose of the vehicle carrier frame, which contains a unit quaternion q representing the pose and a translation vector t representing the position. The scale state node is used to recover the monocular scale.

[0060] The six types of factors include neighbor factors, co-view factors, distance factors, scale smoothing factors, height prior factors, and marginalization factors. Among them, neighbor factors and co-view factors are collectively referred to as monocular visual odometry factors. Once a new visual relative pose measurement is obtained, the factor graph optimization module will be triggered to create the latest keyframe timestamp corresponding state node. The pose nodes and scale nodes newly added to the factor graph will be connected to the corresponding old pose nodes through the monocular visual odometry factors.

[0061] The distance factor is added to the factor graph in the form of all distance measurements that have been eliminated through non-line-of-sight elimination.

[0062] The scale smoothing factor is added to each newly added scale state node;

[0063] The height prior factor is added to each newly added pose state node;

[0064] The marginalization factor is obtained by converting the monocular visual odometry factor through Shure complement and then connecting it to the corresponding node.

[0065] Furthermore, the specific process of constructing the initial pose factor map using the initial global heading angle and the initial visual odometry scale in step S6 is as follows:

[0066] Based on the initial scale of visual odometry, a scale prior factor is added during initialization. This scale prior factor is used to ensure that the state variables at each scale are consistent with the initial scale s. init There will be no significant deviation between them;

[0067] Based on the initial global heading angle, a heading angle prior factor is added during initialization. This prior factor is used to ensure that the initial pose heading angle does not deviate from the estimated initial global heading angle yaw. init The heading angle prior factor is added only once when initialization is complete.

[0068] Compared with the prior art, the present invention has the following advantages:

[0069] I. This invention constructs a measurement preprocessing module, a monocular visual odometry module, an initialization module, a non-line-of-sight recognition and elimination module, and a factor graph optimization module. The measurement preprocessing module receives raw images from a monocular camera and ranging messages from UWB tags, providing distortion-corrected and photometrically corrected images and synchronized distance measurements for subsequent steps. The monocular visual odometry module uses the corrected temporal images as input to optimize incrementally estimating the camera pose, and uses the estimation results to construct visual relative pose constraints in the subsequent factor graph optimization process. The initialization module estimates the initial global heading angle and the initial scale of the visual odometry; the resulting initial values ​​are used to construct the initial pose factor graph, ensuring the smooth progress of the subsequent factor graph optimization process. The non-line-of-sight recognition and elimination module further filters the synchronized distance measurements, thereby avoiding erroneous ranging values ​​from affecting the accuracy of subsequent positioning. The factor graph optimization module jointly optimizes the constraints from the monocular visual odometry, UWB distance measurements, and scene prior knowledge, thereby estimating the vehicle's global pose. By fully combining monocular visual positioning information and UWB ranging information, it can not only solve the problems of scale uncertainty, scale drift and inability to estimate global pose in monocular visual odometry, but also achieve real-time and accurate estimation of vehicle global pose under various complex working conditions without the need to deploy multiple UWB base stations.

[0070] II. This invention, by designing a factor graph containing two types of state nodes, six types of factors, and a distance noise model, can fully optimize and integrate UWB ranging information and visual positioning information. On the one hand, it utilizes the metric distance measurement from the base station to the vehicle-mounted UWB tag to provide metric scale information for monocular visual odometry, estimating and recovering the monocular scale, thus solving the problems of monocular scale uncertainty and scale drift. On the other hand, it uses the distance measurement between the vehicle-mounted UWB tag and the base station with known coordinates to correct the cumulative drift of the visual odometry in real time, and converts the relative pose estimation result of the monocular visual odometry into a global pose. This achieves a real-time global vehicle positioning scheme that couples UWB distance measurement and visual relative pose measurement, exhibiting excellent performance with high accuracy, no drift, and low latency.

[0071] Third, this invention further identifies and eliminates non-line-of-sight (LOS) distance measurements after synchronization. By combining the latest relative pose transformation of the monocular visual odometry with the pose of the previous moment in the factor graph, the current UWB tag position is predicted. Then, for all distance measurements at the current moment, the distance between the predicted tag position and the corresponding UWB base station is calculated. When the measured distance and the predicted distance exceed a set threshold, the measurement is determined to be a non-line-of-sight measurement. This allows for accurate identification and direct elimination of non-line-of-sight measurements, effectively avoiding the impact of erroneous distance measurements on the accuracy of subsequent positioning.

[0072] Third, this invention utilizes monocular visual odometry to construct visual relative pose constraints through two visual relative pose measurements. By designing co-view factors and adjacent factors, it achieves the fusion of the visual odometry sliding window and the factor graph. Under the premise of ensuring real-time performance, the visual constraints in the sliding window are introduced into the factor graph, which can effectively address problems such as discontinuous and unsmooth pure UWB positioning trajectories, abnormal positioning under non-line-of-sight signals, and inability to position due to insufficient base stations. Compared with pure UWB positioning schemes, only a small number of base stations need to be simply deployed. Attached Figure Description

[0073] Figure 1 This is a schematic diagram of the system structure of the present invention;

[0074] Figure 2 This is a schematic diagram of the method flow of the present invention;

[0075] Figure 3 This is a schematic diagram illustrating the application process of an example.

[0076] Figure 4 This is a schematic diagram of the visual odometry sliding window and two visual relative pose measurements in the embodiment.

[0077] Figure 5 This is the initial pose factor diagram in the embodiment;

[0078] Figure 6 The factor diagram in the example;

[0079] Figure 7 This is a schematic diagram comparing the positioning trajectories in the XY plane and XZ plane in the embodiment;

[0080] The markings in the diagram are as follows: 1. Measurement preprocessing module; 2. Monocular visual odometry module; 3. Initialization module; 4. Non-line-of-sight recognition and elimination module; 5. Factor graph optimization module. Detailed Implementation

[0081] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0082] Example

[0083] like Figure 1 As shown, a monocular vision vehicle positioning system integrating UWB ranging information includes a measurement preprocessing module 1, a monocular vision odometer module 2, an initialization module 3, a non-line-of-sight recognition and elimination module 4, and a factor graph optimization module 5. The measurement preprocessing module 1 is connected to the monocular vision odometer module 2, the non-line-of-sight recognition and elimination module 4, and the initialization module 3, respectively. The monocular vision odometer module 2 is connected to the initialization module 3 and the non-line-of-sight recognition and elimination module 4, respectively. The monocular vision odometer module 2, the non-line-of-sight recognition and elimination module 4, and the initialization module 3 are all connected to the factor graph optimization module 5.

[0084] The measurement preprocessing module 1 is used to perform distortion correction and photometric correction on the raw images acquired by the monocular camera, and to interpolate and obtain the synchronized distance measurement based on the UWB tag ranging information.

[0085] The monocular visual odometry module 2 uses the distortion-reduced and photometrically corrected monocular camera images to incrementally estimate the camera pose and construct two visual relative pose measurements.

[0086] The non-line-of-sight identification and rejection module 4 is used to filter the distance measurements after synchronization in order to reject non-line-of-sight measurements.

[0087] The initialization module 5 is used to estimate the initial global heading angle and the initial scale of the visual odometry based on the distance measurement and camera pose estimation results after synchronization.

[0088] The factor graph optimization module 6 is used to combine the output data of the monocular vision odometry module, the non-line-of-sight recognition and elimination module, and the initialization module to estimate the global pose information of the vehicle.

[0089] The above system is applied in practice to realize a monocular vision vehicle localization method that integrates UWB ranging information, such as... Figure 2 As shown, it includes the following steps:

[0090] S1. Acquire the raw images captured by the monocular camera and the ranging information from the UWB tag;

[0091] S2. Perform distortion correction and photometric correction on the monocular camera images sequentially;

[0092] Based on the ranging information of UWB tags, the quality of the corresponding distance measurement after synchronization is determined by judging the quality of the corresponding distance measurement before synchronization.

[0093] S3. Based on the monocular camera images after distortion correction and photometric correction, the camera pose is estimated incrementally using a sliding window optimization method. The sliding window optimization results are used to construct two visual relative pose measurements.

[0094] S4. Perform non-line-of-sight identification and rejection processing on the distance measurement after synchronization;

[0095] S5. Combine the distance measurement after synchronization and the camera pose estimation results to estimate the initial global heading angle and the initial scale of the visual odometry.

[0096] S6. Construct an initial pose factor map using the initial global heading angle and the initial visual odometry scale. Construct visual relative pose constraints using two types of visual relative pose measurements. Combine distance measurements eliminated by non-line-of-sight and the set historical prior constraints. Use a method of alternating distance noise model estimation and factor map optimization to estimate the key frame frequency pose, which is the global pose of the vehicle.

[0097] This embodiment applies the above-described solution, such as Figure 3 As shown, a measurement preprocessing module, a monocular visual odometry module, an initialization module, a non-line-of-sight recognition and elimination module, and a factor graph optimization module are constructed. The input terminals of the measurement preprocessing module are connected to the monocular camera and the UWB base station, respectively. The working principle of each functional module is explained in detail below:

[0098] I. Measurement Preprocessing Module

[0099] The measurement preprocessing module receives raw images from a monocular camera and ranging messages from UWB tags, providing subsequent steps with distortion-corrected and photometrically corrected images and synchronized distance measurements.

[0100] For monocular images, this scheme uses both a geometric camera model and a photometric camera model to simulate the image generation process. The geometric camera model adopts a traditional pinhole camera model, obtaining the camera intrinsic parameter matrix and distortion coefficients through camera intrinsic parameter calibration, and removing radial and tangential distortions from the original image. The photometric camera model is modeled using a nonlinear response function, a nonparametric vignetting map, and exposure time. Photometric correction is then performed on the distortion-removed image based on the camera photometric calibration results.

[0101] For the original ranging message, this scheme uses linear interpolation based on the timestamp to obtain the distance measurement synchronized with the image frame timestamp. Based on the signal strength information in the original message, the quality of the corresponding distance measurement before synchronization is evaluated.

[0102]

[0103] Where rx_rssi and fp_rssi are the total received signal strength indicator and the first path signal strength indicator, respectively, θ thrs To set the filtering threshold.

[0104] The quality of the corresponding synchronized distance measurement will only be set to "good" if the quality of the nearest distance measurement on both sides of the image frame timestamp is "good". Otherwise, its quality will be "bad".

[0105] II. Monocular Vision Odometry Module

[0106] The monocular visual odometry module uses the corrected temporal image as input to incrementally estimate the camera pose. This scheme uses it as the front end, and its estimation results are used to construct visual relative pose constraints in the subsequent factor map optimization process.

[0107] This module solves for camera pose by minimizing photometric error. It divides the input image into keyframes and non-keyframes; only keyframes are added to the sliding window and participate in sliding window optimization. For non-keyframes, their pose is directly aligned with the latest keyframe. For keyframes, their pose and the depth of their high-gradient points undergo multiple sliding window optimizations until they are marginalized out of the window. The sliding window length is set to 8.

[0108] This scheme utilizes a visual odometry sliding window to construct two types of visual relative pose measurements. For example... Figure 4 As shown, after each sliding window optimization, the latest keyframe pose T is calculated. k and the new keyframe pose T k-1 The relative transformation between them is considered as adjacent relative pose measurements. Before each edge-mapping keyframe, the pose T of the latest keyframe is calculated. k and the oldest keyframe pose T k-7 The relative transformation between them is considered as a co-view relative pose measurement. Any two frames of pose T... i and T j The relative pose T between i j The calculation formula is:

[0109]

[0110] in, and Representing the frames F obtained by sliding window optimization estimation respectively i and frame F j The position.

[0111] Since the pose defined in the subsequent factor map optimization is in the carrier coordinate system, it is necessary to use the extrinsic parameters between the carrier coordinate system and the camera coordinate system. Relative transformation between camera frames Transformed into relative transformations between carrier frames

[0112]

[0113] in, This is the relative pose between camera frames calculated using the previous formula.

[0114] III. Initialization Module

[0115] The initialization module is used to estimate the initial global heading angle and the initial visual odometry scale. The obtained initial values ​​are used to construct the initial pose factor map, ensuring the smooth progress of subsequent factor map optimization.

[0116] First, the position of the UWB tag at the corresponding time is determined by minimizing the error between the distance measured from the vehicle-mounted UWB tag to the known location UWB base station and the actual distance measurement.

[0117]

[0118] in, and These represent the UWB label and the base station with sequence number i in world frame f, respectively. w In the global location, Ω is the set of base station indices with good distance measurement quality after synchronization at the corresponding time, and r i Let be the measured UWB distance from the tag to the base station with sequence number i. During the initialization phase, the vehicle can be approximated as moving within the same horizontal plane, and the z-coordinate of the tag is fixed at a known height.

[0119] Then, before the visual odometry window edge-out for the first time, the co-view relative pose between the latest and oldest keyframes is calculated. Ignoring the vehicle's vertical displacement, the visual translation vector consisting of the vehicle's lateral and longitudinal displacements is extracted. UWB translation vector This represents the difference between the x and y coordinates of the UWB tag at the corresponding moment. The initial global heading angle is yaw. init and initial scale s init The calculation formulas are as follows:

[0120]

[0121]

[0122] Using the initial value estimation results described above, prior factors are constructed and generated. Figure 5 The initial pose factor graph is shown. The neighbor factor, common-view factor, and distance factor will be explained in detail in the factor graph optimization module. In this stage, the adjacent relative pose measurements corresponding to the neighbor factors are calculated within a sliding window before the first marginalization. The scale prior factor is used to ensure that the state quantities at each scale are consistent with the initial scale s. init There will be no large deviation between them. The heading angle prior factor can ensure that the initial pose heading angle will not deviate from the estimated initial global heading angle yaw. init .

[0123] After initial pose factor map optimization, the accurate global pose of the first 8 carrier frames can be obtained.

[0124] IV. Non-line-of-sight recognition and rejection module

[0125] The non-line-of-sight identification and rejection module further filters UWB distance measurements that are of good quality after synchronization, thereby avoiding damage to the positioning system caused by erroneous distance measurements.

[0126] Prediction scale at the current moment The mean value of the states at the previous 7 time steps in the factor graph is taken. This is combined with the latest neighboring relative pose transformation from monocular visual odometry. pose at the previous time step in the sum factor graph Predict the vehicle's attitude and position at the current moment:

[0127]

[0128]

[0129] Where R and t represent the rotation matrix and translation vector corresponding to the pose matrix T, respectively. This indicates that the translation vector is at a monocular scale, and it needs to be multiplied by the scale to be transformed to the true scale.

[0130] Based on this, the current UWB tag position can be predicted using the following formula.

[0131]

[0132] in, This refers to the external parameters between the pre-calibrated UWB tags and the camera.

[0133] Finally, for all distance measurements at the current moment, the distance between the predicted tag location and the corresponding UWB base station is calculated. Measurements whose measured or predicted distances exceed a set threshold will be classified as non-line-of-sight measurements and directly discarded.

[0134] V. Factor Plot Optimization Module

[0135] The factor graph optimization module jointly optimizes constraints derived from monocular visual odometry, UWB distance measurement, and scene prior knowledge. For example... Figure 6 As shown, the factor graph designed in this scheme contains two types of state nodes, six types of factors, and a distance noise model. Each time step corresponds to two state nodes: a pose state node and a scale state node. The former is the global pose of the vehicle carrier frame, which contains a unit quaternion q representing the pose and a translation vector t representing the position; the latter is used to recover the monocular scale.

[0136] 1. Monocular visual odometry factor: The adjacent factor and the co-view factor are collectively referred to as the monocular visual odometry factor. Each visual relative pose measurement triggers the factor graph optimization module to create the latest keyframe timestamp-corresponding state node. Newly added pose and scale nodes to the factor graph are connected to their corresponding old pose nodes via the monocular visual odometry factor. The monocular visual odometry factor error term is modeled as follows:

[0137] r vision =[r q r p ]

[0138]

[0139]

[0140] Here, j and i correspond to the newly added state node and the old state node, respectively. This represents the quaternion multiplication operation. and R represents the rotation quaternion corresponding to visual relative pose measurement and the displacement vector at the monocular scale, where R is the rotation matrix corresponding to q.

[0141] 2. Distance Factor: All distance measurements that have been eliminated through non-line-of-sight elimination are added to the factor graph as distance factors. Since the pose state in the factor graph describes the global pose of the carrier frame in the world frame, while UWB measures the distance from the tag to the base station, the distance factor error term r is constructed accordingly. range First, we need to use extrinsic parameters to obtain the pose of the label at the current time step:

[0142]

[0143] in, The UWB ranging value from the tag to base station number i. This represents the known location of UWB within the carrier frame. The position of base station number i in the world frame is measured by a total station during the base station deployment phase.

[0144] A Gaussian mixture model is used to more accurately model the distance noise distribution, with the number of Gaussian components set to 2. The entire factor graph optimization process consists of three steps: first, after adding all current state nodes and factors, a factor graph optimization is performed to calculate the residual terms; then, the Gaussian mixture model parameters are estimated based on these residual terms using the expectation-maximization algorithm; finally, the optimized mixture model is used to perform another factor graph optimization to obtain the optimal estimate of the target state.

[0145] 3. Scale Smoothing Factor: A scale smoothing factor is added to each newly added scale state node. The average scale state value of the previous 7 time steps is used as the scale prediction value at the current time moment, and it is added to the factor graph as a prior constraint.

[0146] 4. Height Prior Factor: Add a height prior factor to each newly added pose state node. Since the vehicle height does not change during normal driving, the vehicle height at the previous moment is used as the height prior for the current moment.

[0147] 5. Marginalization Factor: As the number of state nodes increases, the time required for distance noise model estimation and factor graph optimization will also increase. To meet the real-time requirements of the positioning system, only nodes and factors within a certain time window are retained, and state nodes exceeding the time window are marginalized. The monocular visual odometry factors are converted into marginalization factors using Shur complement and then connected to the corresponding nodes.

[0148] To verify the effectiveness of this technical solution, this embodiment takes an underground parking lot as an example and fully simulates the entire process of a vehicle entering the underground parking lot via a ramp, finding a parking space, parking, driving out of the parking space, and finally leaving the underground parking lot via a ramp. Figure 7 The comparison between the real-time positioning trajectory obtained using this method and the actual ground truth positioning trajectory is shown. The ground truth is provided by the LVI-SAM visual laser inertial positioning method. Figure 7 It can be seen that the positioning trajectory of this scheme and the true trajectory highly overlap. The quantitative error analysis results show that the root mean square value of the global position error of this scheme is 0.175m, the maximum error is 0.344m, and the 95th percentile of error is 0.288m. Combined with the real-time statistical results, the average positioning time delay of this scheme is 88.48ms.

[0149] In summary, this solution can fully combine monocular visual information and UWB ranging information, while solving a series of problems existing in monocular visual odometry (such as non-objective scale, scale drift, cumulative positioning error, and inability to achieve global positioning), and greatly reducing the number of UWB base stations required (by utilizing the relative pose constraints between frames of monocular visual odometry, it effectively addresses problems such as discontinuous and unsmooth pure UWB positioning trajectories, abnormal positioning under non-line-of-sight signals, and inability to position due to insufficient base stations). It can accurately estimate the global pose of the vehicle in real time under various complex working conditions.

Claims

1. A monocular vision vehicle localization method fusing UWB ranging information, applied to a monocular vision vehicle localization system fusing UWB ranging information, characterized in that, The system includes a measurement preprocessing module (1), a monocular visual odometry module (2), an initialization module (3), a non-line-of-sight recognition and elimination module (4), and a factor graph optimization module (5). The measurement preprocessing module (1) is connected to the monocular visual odometry module (2), the non-line-of-sight recognition and elimination module (4), and the initialization module (3), respectively. The monocular visual odometry module (2) is connected to the initialization module (3) and the non-line-of-sight recognition and elimination module (4), respectively. The monocular visual odometry module (2), the non-line-of-sight recognition and elimination module (4), and the initialization module (3) are connected to the factor graph optimization module (5), respectively. The measurement preprocessing module (1) is used to perform distortion correction and photometric correction on the original images acquired by the monocular camera, and to interpolate the distance measurement after synchronization based on the UWB tag ranging information. The monocular visual odometry module (2) is used to incrementally estimate the camera pose and construct two visual relative pose measurements based on the monocular camera image after distortion correction and photometric correction. The non-line-of-sight identification and elimination module (4) is used to screen the distance measurements after synchronization in order to eliminate non-line-of-sight measurements; The initialization module (3) is used to estimate the initial global heading angle and the initial scale of the visual odometry based on the distance measurement and camera pose estimation results after synchronization. The factor graph optimization module (5) is used to combine the output data of the monocular vision odometer module (2), the non-line-of-sight recognition and elimination module (4), and the initialization module (3) to estimate the global pose information of the vehicle. The monocular vision vehicle localization method that integrates UWB ranging information includes the following steps: S1. Acquire the raw images captured by the monocular camera and the ranging information from the UWB tag; S2. Perform distortion correction and photometric correction on the monocular camera images sequentially; Based on the ranging information of UWB tags, the quality of the corresponding distance measurement after synchronization is determined by judging the quality of the corresponding distance measurement before synchronization. S3. Based on the monocular camera images after distortion correction and photometric correction, the camera pose is estimated incrementally using a sliding window optimization method. The sliding window optimization results are used to construct two visual relative pose measurements. S4. Perform non-line-of-sight identification and rejection processing on the distance measurement after synchronization; S5. Combine the distance measurement after synchronization and the camera pose estimation results to estimate the initial global heading angle and the initial scale of the visual odometry. S6. Construct an initial pose factor map using the initial global heading angle and the initial visual odometry scale. Construct visual relative pose constraints using two types of visual relative pose measurements. Combine distance measurements eliminated by non-line-of-sight and the set historical prior constraints. Use the method of alternating distance noise model estimation and factor map optimization to estimate the key frame frequency pose, that is, obtain the vehicle global pose. The two types of visual relative pose measurement in step S3 are adjacent relative pose measurement and co-visual relative pose measurement. The specific process of step S3 is as follows: S31. First, the distortion-corrected and photometric-corrected monocular camera image is divided into key frames and non-key frames. Only key frames can be added to the sliding window and participate in the sliding window optimization. For key frames, their pose and the depth of their high gradient points will undergo multiple sliding window optimizations until they are marginalized out of the sliding window. For non-key frames, align them directly with the latest key frame to obtain the pose. S32. Subsequently, two visual relative pose measurements were constructed using a visual odometry sliding window. After each sliding window optimization, the latest keyframe pose was calculated. and the next new keyframe pose The relative transformation between them is called adjacent relative pose measurement; Before each edge-mapping keyframe, calculate the pose of the latest keyframe. and the oldest keyframe pose The relative transformation between them is called common-view relative pose measurement; Among them, any two frames pose and Relative pose between The calculation formula is: In the formula, and These represent the frames obtained from the sliding window optimization estimation. and frame The pose, since the pose defined in the subsequent factor map optimization is in the carrier coordinate system, is obtained through the extrinsic parameters between the carrier coordinate system and the camera coordinate system. Relative transformation between camera frames This is further transformed into a relative transformation between carrier frames. : In the formula, This refers to the relative pose between camera frames.

2. The monocular vision vehicle localization method fusing UWB ranging information according to claim 1, characterized in that, The specific process of step S2 is as follows: S21. For monocular camera images, the image generation process is simulated using both a geometric camera model and a photometric camera model. The geometric camera model adopts a traditional pinhole camera model. The camera intrinsic parameter matrix and distortion coefficients are obtained through camera intrinsic parameter calibration, and radial and tangential distortions of the original image are removed. The photometric camera model is modeled using a nonlinear response function, a nonparametric vignetting plot, and exposure time. Photometric correction is performed on the distortion-removed image based on the camera photometric calibration results; S22. For the ranging message of the UWB tag, linear interpolation is used to obtain the distance measurement synchronized with the image frame timestamp based on the timestamp. First, the signal strength information in the original message is used to determine the quality of the corresponding distance measurement before synchronization. When the quality of the nearest distance measurement on both sides of the image frame timestamp is the same When the time is right, the quality of the distance measurement after synchronization is set to [value]. Otherwise, the quality of the distance measurement after synchronization is set to... .

3. The monocular vision vehicle localization method fusing UWB ranging information according to claim 2, characterized in that, In step S22, the quality of the corresponding distance measurement before synchronization is specifically determined using the following formula. Make a judgment: in, and These are the total received signal strength indicator and the first path signal strength indicator, respectively. To set the filtering threshold.

4. The monocular vision vehicle localization method fusing UWB ranging information according to claim 3, characterized in that, Step S4 is to determine the quality after synchronization. Further screening using UWB distance measurements, specifically including: Combined with the latest adjacent relative pose transformation of monocular visual odometry pose at the previous time step in the sum factor graph Predict the vehicle's attitude and position at the current moment: in, and Representing the pose matrix respectively The corresponding rotation matrix and translation vector, This indicates that the translation vector is at a monocular scale, and it needs to be multiplied by the scale to be transformed to the true scale. The current time scale is used for prediction. Based on this, we can further predict the current UWB tag position. : in, The external parameters between the pre-calibrated UWB tags and the camera; Finally, for all distance measurements at the current moment, the distance between the predicted tag location and the corresponding UWB base station is calculated. If the measured distance and the predicted distance exceed the set threshold, the measurement is determined to be a non-line-of-sight measurement and is directly removed.

5. The monocular vision vehicle localization method fusing UWB ranging information according to claim 3, characterized in that, The specific process of step S5 is as follows: First, the position of the UWB tag at the corresponding time is determined by minimizing the error between the distance measured from the vehicle-mounted UWB tag to the known location UWB base station and the actual distance measurement. in, and These represent the UWB label and the serial number, respectively. Base stations in the world frame Global position in The distance measurement quality after synchronization at the corresponding time is The set of base station serial numbers For the label to the serial number The actual UWB distance measured by the base station is assumed during the initialization phase to be that the vehicle is moving within the same horizontal plane, and the tag is... The coordinates are fixed at a known height; Then, before the visual odometry window edge-out for the first time, the co-view relative pose between the latest and oldest keyframes is calculated, ignoring the vehicle's vertical displacement, to obtain the initial global heading angle. and initial scale They are respectively: in, The UWB translation vector is a visual translation vector composed of the vehicle's lateral and longitudinal displacements. UWB label for the corresponding time moment coordinates and The difference in coordinates.

6. The monocular vision vehicle localization method fusing UWB ranging information according to claim 5, characterized in that, In step S6, the factor graph contains two types of state nodes, six types of factors, and a distance noise model. The two types of state nodes include pose state nodes and scale state nodes. The pose state node represents the global pose of the vehicle carrier frame and contains a single unit quaternion representing the pose. and a translation vector representing position The scale state node is used to recover the monocular scale. The six types of factors include neighbor factors, co-view factors, distance factors, scale smoothing factors, height prior factors, and marginalization factors. Among them, neighbor factors and co-view factors are collectively referred to as monocular visual odometry factors. Once a new visual relative pose measurement is obtained, the factor graph optimization module will be triggered to create the latest keyframe timestamp corresponding state node. The pose nodes and scale nodes newly added to the factor graph will be connected to the corresponding old pose nodes through the monocular visual odometry factors. The distance factor is added to the factor graph in the form of all distance measurements that have been eliminated through non-line-of-sight elimination. The scale smoothing factor is added to each newly added scale state node; The height prior factor is added to each newly added pose state node; The marginalization factor is obtained by converting the monocular visual odometry factor through Shure complement and then connecting it to the corresponding node.

7. The monocular vision vehicle localization method fusing UWB ranging information according to claim 6, characterized in that, The specific process of constructing the initial pose factor map using the initial global heading angle and the initial visual odometry scale in step S6 is as follows: Based on the initial scale of visual odometry, a scale prior factor is added during initialization. This scale prior factor is used to ensure that the state variables at each scale are consistent with the initial scale. There will be no significant deviation between them; Based on the initial global heading angle, a heading angle prior factor is added during initialization. This prior factor is used to ensure that the initial pose heading angle does not deviate from the estimated initial global heading angle. The heading angle prior factor is added only once when initialization is complete.

Citation Information

Patent Citations

  • Spatial positioning method and system

    CN112837374A

  • Positioning method and system of UWB and visual tight coupling SLAM algorithm

    CN116027266A