Multi-view vision and inertial navigation fusion-based control point-free rapid mapping method and system
By constructing a rigid combined measurement system of an industrial camera and an inertial measurement unit (IMU) with a non-coplanar layout, and combining infinity plane homography constraints and temperature compensation, the problem of online spatiotemporal calibration of multi-view vision and inertial navigation fusion in complex environments was solved, achieving high-precision, robust, and rapid mapping, and outputting three-dimensional mapping results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HENAN JINGCHENG SURVEY PLANNING DESIGN CO LTD
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-03
Smart Images

Figure CN121932968B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of surveying and mapping and visual-inertial fusion technology, specifically to a rapid surveying method and system for multi-view vision and inertial navigation fusion without control points. Background Technology
[0002] Surveying and mapping refers to the technical process of acquiring spatial location, shape, size, and related attribute information of a target area to generate map data, 3D point clouds, 3D models, or other spatial information results. The multi-view vision and inertial navigation fusion-based rapid surveying and mapping method and system without control points acquires image information of the scene under test through multiple visual sensors with different perspectives, and combines this with platform motion information acquired by the inertial measurement unit to perform scene positioning, mapping, and surveying processing, thereby achieving rapid surveying and mapping of the target area without the need to establish ground control points.
[0003] In existing surveying methods, ground control points are typically established and combined with visual or inertial measurement data for positioning and mapping to ensure accuracy. However, in complex environments or large survey areas, the establishment of control points is labor-intensive and inefficient, making it difficult to meet the application requirements of rapid surveying. Therefore, some technical solutions attempt to use a fusion of visual and inertial information to complete positioning and reconstruction in scenarios with few or no control points.
[0004] However, existing vision-inertial navigation fusion methods still have technical shortcomings in online spatiotemporal calibration. Traditional multi-view vision and IMU fusion methods typically assume that the camera optical axes are approximately coplanar or that the motion is mainly planar when performing online calibration. This leads to degradation or accuracy reduction in calibration methods based on the infinity plane during complex 3D motion. In addition, existing methods often ignore the impact of temperature changes on the micro-deformation of rigid systems and sensor timing, resulting in severe accuracy degradation under long-term operation.
[0005] Therefore, how to achieve highly robust online spatiotemporal calibration under complex motion and temperature variations, and maintain accuracy stability during long-term operation, has become a technical challenge that urgently needs to be solved in this field. Summary of the Invention
[0006] To address the issues of easy degradation in calibration under complex 3D motion and accuracy decay due to temperature drift during long-term operation in existing technologies, this invention provides a rapid mapping method and system without control points that integrates multi-view vision and inertial navigation. This method enables high-precision and robust rapid mapping without the need to establish ground control points.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a rapid mapping method and system for multi-view vision and inertial navigation fusion without control points, comprising:
[0008] S1. Construct a rigid combined measurement system including at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and collect multi-view image data and corresponding inertial data of the scene under test.
[0009] S2. Based on the homography constraint of the infinity plane, online spatiotemporal calibration is performed on the multi-view image data and inertial data to determine the time offset parameters and external parameters between each industrial camera and the inertial measurement unit (IMU), and to compensate for parameter changes caused by temperature drift.
[0010] S3. Perform joint extraction and matching of point features and line features on the spatiotemporally calibrated multi-view image data to obtain multi-view feature observation information;
[0011] S4. Construct a multi-view geometric constraint factor based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information. Combine this with the IMU pre-integration obtained from the inertial data to construct a nonlinear optimization model that tightly couples vision and inertial navigation.
[0012] S5. Based on the nonlinear optimization model, perform sliding window local optimization and global closed-loop detection to solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test.
[0013] S6. Generate the 3D mapping results of the scene to be tested based on the corrected pose parameters and 3D structural parameters.
[0014] Preferably, step S1 specifically includes:
[0015] At least three industrial cameras and inertial measurement units (IMUs) are mounted on the same rigid carrier to form a rigid combined measurement system;
[0016] Adjust the mounting position and orientation of at least three industrial cameras to ensure that the at least three industrial cameras are arranged in a non-coplanar layout;
[0017] The rigid combined measurement system is used to acquire multi-view image data of the scene under test and inertial data corresponding to the multi-view image data;
[0018] Among the at least three industrial cameras, at least one industrial camera is arranged in a forward direction, and the other two industrial cameras are arranged in different tilt directions, and the optical axes of each industrial camera are not parallel to each other and are not in the same plane.
[0019] Preferably, step S2 specifically includes:
[0020] Extract distant features from multi-view image data, and construct an infinity homography matrix based on the distant features;
[0021] The infinity plane homography matrix is correlated with the inertial data in time to establish a spatiotemporal constraint relationship between the industrial camera and the inertial measurement unit (IMU).
[0022] The time offset parameters and extrinsic parameters between each industrial camera and the inertial measurement unit (IMU) are iteratively solved based on the spatiotemporal constraint relationship, and the time offset parameters and extrinsic parameters are corrected in combination with temperature changes.
[0023] Preferably, the time offset parameter is the time delay parameter between each industrial camera and the inertial measurement unit (IMU);
[0024] The extrinsic parameters include rotational and displacement extrinsic parameters between each industrial camera and the inertial measurement unit (IMU). During the mapping process, the time offset parameters and displacement extrinsic parameters are updated online according to temperature changes. The rotational extrinsic parameters are optimized online based on the homography constraint of the infinity plane.
[0025] Preferably, step S3 specifically includes:
[0026] Extract point and line features from multi-view image data after spatiotemporal calibration;
[0027] The point features and line features are matched separately, and consistency filtering is performed based on the matching results;
[0028] Based on the selected point feature matching results and line feature matching results, multi-view feature observation information is obtained.
[0029] Preferably, step S4 specifically includes:
[0030] Based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information, a multi-view geometric constraint factor is constructed.
[0031] A multi-view reprojection error term is established based on the multi-view geometric constraint factor, and the IMU pre-integration is calculated based on the inertial data;
[0032] A nonlinear optimization model tightly coupled with vision and inertial navigation is constructed using the multi-view reprojection error term and the IMU pre-integral quantity.
[0033] Preferably, the nonlinear optimization model includes multi-view geometric constraints and IMU pre-integration constraints, and uses the pose parameters, velocity parameters, IMU bias parameters, and three-dimensional structural parameters of the scene as optimization variables.
[0034] Preferably, S5 specifically includes:
[0035] The nonlinear optimization model is locally optimized within a sliding window, and historical states that exceed the sliding window are marginalized.
[0036] The pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test are solved based on the local optimization results.
[0037] When a closed loop is detected, global consistency correction is performed on the pose parameters and the three-dimensional structural parameters.
[0038] Preferably, S6 specifically includes:
[0039] Three-dimensional reconstruction is performed based on the corrected pose parameters and three-dimensional structural parameters;
[0040] Output the three-dimensional mapping results of the scene to be tested, wherein the three-dimensional mapping results include at least one of three-dimensional point cloud, three-dimensional model and map data.
[0041] A multi-view vision and inertial navigation fusion-based rapid mapping system without control points, the system comprising:
[0042] The data acquisition module constructs a rigid combined measurement system comprising at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and acquires multi-view image data and corresponding inertial data of the scene under test.
[0043] The spatiotemporal calibration module, based on the homography constraint of the infinity plane, performs online spatiotemporal calibration on the multi-view image data and inertial data, determines the time offset parameters and external parameters between each industrial camera and the inertial measurement unit (IMU), and compensates for parameter changes caused by temperature drift.
[0044] The feature processing module performs joint extraction and matching of point and line features on multi-view image data after spatiotemporal calibration to obtain multi-view feature observation information.
[0045] The fusion modeling module constructs multi-view geometric constraint factors based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information. Combined with the IMU pre-integration obtained from the inertial data, it constructs a nonlinear optimization model that tightly couples vision and inertial navigation.
[0046] The optimization and correction module performs sliding window local optimization and global closed-loop detection based on the nonlinear optimization model to solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test.
[0047] The results output module generates 3D mapping results of the scene under test based on the corrected pose parameters and 3D structural parameters.
[0048] This invention provides a rapid, control-point-free mapping method and system that integrates multi-view vision and inertial navigation. It offers the following advantages:
[0049] 1. This invention constructs a rigid combined measurement system consisting of at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and combines multi-view geometric constraints and IMU pre-integration to construct a tightly coupled optimization model, achieving rapid mapping without the need to set up ground control points. The non-coplanar layout effectively enhances the system's adaptability to complex three-dimensional motion and can support high-precision three-dimensional mapping processing.
[0050] 2. This invention performs online spatiotemporal calibration based on homography constraints at infinity and monitors ambient temperature in real time. It dynamically compensates and updates time offset parameters and external parameters, which can maintain the stability of the spatiotemporal correspondence between the industrial camera and the inertial measurement unit (IMU) during long-term operation. This effectively suppresses the cumulative error caused by temperature drift and significantly improves the mapping accuracy in complex environments.
[0051] 3. This invention, by jointly extracting and matching point and line features from multi-view image data after spatiotemporal calibration, can simultaneously utilize texture information and structural edge information in the scene, thereby enhancing the integrity of multi-view feature observation information and providing reliable input for subsequent geometric constraint construction and fusion optimization.
[0052] 4. This invention combines local optimization via sliding window with global closed-loop detection to solve and correct the pose parameters of a rigid combined measurement system and the three-dimensional structural parameters of the scene under test. This can reduce the cumulative deviation in long-distance mapping and output mapping results such as three-dimensional point clouds, three-dimensional models, or map data. Attached Figure Description
[0053] Figure 1 The flowchart shows the multi-view vision and inertial navigation fusion method for rapid mapping without control points according to the present invention.
[0054] Figure 2 This is an architecture diagram of the multi-view vision and inertial navigation fusion control point-free rapid mapping system of the present invention. Detailed Implementation
[0055] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] Please see the appendix Figure 1This invention provides a method and system for rapid mapping without control points using multi-view vision and inertial navigation fusion, comprising:
[0057] S1. Construct a rigid combined measurement system including at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and collect multi-view image data and corresponding inertial data of the scene under test.
[0058] Furthermore, step S1 specifically includes:
[0059] At least three industrial cameras and inertial measurement units (IMUs) are mounted on the same rigid carrier to form a rigid combined measurement system;
[0060] Adjust the mounting position and orientation of at least three industrial cameras to ensure that the at least three industrial cameras are arranged in a non-coplanar layout;
[0061] A rigid combined measurement system is used to acquire multi-view image data of the scene under test and corresponding inertial data.
[0062] Among the at least three industrial cameras, at least one industrial camera is set in the forward direction, and the other two industrial cameras are set in different tilt directions, and the optical axes of each industrial camera are not parallel to each other and are not in the same plane.
[0063] Specifically, at least three industrial cameras and inertial measurement units (IMUs) are mounted on the same rigid carrier to form a rigid combined measurement system. The rigid carrier can be a metal mounting frame, an integrated support platform, or other mounting structures with sufficient structural strength. Each industrial camera and IMU is fixed to the rigid carrier by means of screws, clips, or flanges to ensure that the relative positional relationship between each industrial camera and the IMU remains stable during the measurement process, providing a stable hardware foundation for establishing the spatiotemporal correspondence between the industrial cameras and the IMUs.
[0064] After installation, the installation positions and orientations of at least three industrial cameras are adjusted to form a non-coplanar layout. Specifically, at least one industrial camera is set forward, and the other two industrial cameras are set along different tilt directions. The optical axes of each industrial camera are not parallel to each other and are not in the same plane. Through the above arrangement, different industrial cameras can observe the scene under test from different directions, thereby forming a spatially intersecting multi-view observation structure, which provides a basis for the subsequent construction of multi-view geometric constraints. In one embodiment, one industrial camera can be set as a forward observation camera, and the other two industrial cameras can be set as a left front-down tilt observation camera and a right front-down tilt observation camera, respectively. The inertial measurement unit (IMU) is installed in the middle of the rigid carrier.
[0065] After the rigid combined measurement system is constructed, each industrial camera acquires multi-view image data of the scene under test according to a preset sampling frequency. The inertial measurement unit (IMU) synchronously acquires angular velocity and acceleration data at the corresponding time. The multi-view image data and inertial data can be correlated through a unified time reference, timestamp correspondence, or trigger signal correspondence. The acquired multi-view image data is used to characterize the image information of the scene under test under different views, and the acquired inertial data is used to characterize the attitude and motion state information of the measurement system during the motion process. Both serve as input data for subsequent online spatiotemporal calibration and fusion optimization processing.
[0066] S2. Based on the homography constraint of the infinity plane, online spatiotemporal calibration is performed on multi-view image data and inertial data to determine the time offset parameters and external parameters between each industrial camera and the inertial measurement unit (IMU), and to compensate for parameter changes caused by temperature drift.
[0067] Furthermore, step S2 specifically includes:
[0068] Extract distant features from multi-view image data and construct an infinity homography matrix based on the distant features;
[0069] The infinity plane homography matrix is mapped to inertial data in time to establish a spatiotemporal constraint relationship between the industrial camera and the inertial measurement unit (IMU).
[0070] The time offset parameters and extrinsic parameters between each industrial camera and the inertial measurement unit (IMU) are iteratively solved based on the spatiotemporal constraints, and the time offset parameters and extrinsic parameters are corrected by taking into account temperature changes.
[0071] Specifically, distant features are extracted from multi-view image data, and an infinity homography matrix is constructed based on the extracted distant features. Specifically, feature detection is first performed on the image data acquired by each industrial camera to select distant feature points located in distant background regions that remain stable in consecutive image frames. Then, an infinity homography constraint is established based on the correspondence of the same distant feature point in consecutive image frames. For a distant feature point, its correspondence in adjacent images can be expressed as:
[0072] ;
[0073] in: ;
[0074] In the formula, This represents the homogeneous coordinates of distant feature points in the image at the previous time step. This represents the homogeneous coordinates of the corresponding distant feature point in the image at the next time step. Represents the homography matrix at infinity. This represents the intrinsic parameter matrix of an industrial camera. The rotation matrix represents the change in pose of the industrial camera at adjacent time points, denoted by "". "" indicates that the two differ by a non-zero scaling factor. By constructing the homography matrix at infinity, the attitude change of the industrial camera can be described using distant features, providing a basis for establishing the spatiotemporal constraint relationship between the industrial camera and the inertial measurement unit (IMU) in the future.
[0075] After obtaining the homography matrix at infinity, it is temporally mapped to the inertial data to establish a spatiotemporal constraint relationship between the industrial camera and the inertial measurement unit (IMU). Image frames and inertial data are paired based on the image acquisition time and the sampling time of the IMU's output angular velocity and acceleration data. Combined with the rigid connection between the industrial camera and the IMU, the attitude changes represented by the homography matrix at infinity in the image domain are mapped to the motion state changes represented by the inertial data. The spatiotemporal relationship between the industrial camera and the IMU can be expressed by the following formula:
[0076] ;
[0077] In the formula, This represents the extrinsic transformation matrix from the inertial measurement unit (IMU) coordinate system to the industrial camera coordinate system. Indicates the rotational extrinsic parameter. Indicates the displacement extrinsic parameter. Indicates the time of image acquisition. Indicates the sampling time of the inertial measurement unit (IMU). The time offset parameter between the industrial camera and the inertial measurement unit (IMU) is represented by the above processing. Visual observation and inertial observation can be unified into the same spatiotemporal framework, providing a basis for subsequent online solution of time offset parameters and extrinsic parameters.
[0078] By using the online calibration method based on homography of the infinite plane, combined with the multi-angle observation information provided by the non-coplanar multi-camera layout, the solution of the homography matrix still maintains good constraints under complex three-dimensional motion conditions, effectively overcoming the calibration failure problem caused by motion degradation under the traditional coplanar layout.
[0079] After establishing spatiotemporal constraints, the time offset parameters, rotational extrinsic parameters, and displacement extrinsic parameters between each industrial camera and the inertial measurement unit (IMU) are iteratively solved based on these constraints. The online update of the rotational extrinsic parameters is based on the homography constraint at infinity. Iterative optimization is performed by minimizing the error between the observed distant features in consecutive image frames and the corresponding inertial data, achieving real-time accurate estimation of the rotational extrinsic parameters. The displacement extrinsic parameters and time offset parameters are solved using a combination of iterative solutions and temperature compensation. During the iterative update process, the system's operating environment temperature is simultaneously acquired, and the displacement extrinsic parameters and time offset parameters are compensated and corrected based on temperature changes to reduce the impact of temperature drift on the calibration results.
[0080] Specifically, this invention achieves temperature drift compensation in the following way: temperature compensation of the displacement extrinsic parameter adopts a combination of offline calibration and online correction, and the temperature drift coefficient of the displacement extrinsic parameter is pre-calibrated at different temperatures. In the actual surveying process, the temperature values collected in real time by the temperature sensor are used... Compensation is performed on the displacement extrinsic parameters. Time offset parameter Also affected by temperature, its compensation coefficient Obtained through offline calibration. The temperature compensation relationship between the displacement extrinsic parameter and the time offset parameter can be expressed as:
[0081] , ;
[0082] In the formula, Indicates the current temperature. Indicates reference temperature. This represents the displacement extrinsic parameter at the reference temperature. This represents the time offset parameter at the reference temperature. This represents the temperature compensation coefficient for the displacement external parameter. This represents the temperature compensation coefficient for the time offset parameter. Indicates temperature as Time-corrected displacement extrinsic parameters Indicates temperature as The corrected time offset parameters. Since the temperature drift of the displacement extrinsic mainly affects the baseline length in the reprojection error, the residual error can be further absorbed through multi-view reprojection error in subsequent tight-coupled optimization, thereby ensuring the overall mapping accuracy.
[0083] Through iterative solving and temperature compensation, the calibration relationship between the industrial camera and the inertial measurement unit (IMU) can be continuously updated during the mapping process, thus providing an accurate spatiotemporal alignment basis for subsequent joint extraction of point and line features, construction of multi-view geometric constraints, and tight coupling optimization.
[0084] In one embodiment, a series of consecutive frames of images taken during the measurement platform's forward movement can be selected as the calibration image sequence. The outline of building tops, structural edges near the distant horizon, or other stable features in the distant landscape are extracted from the image sequence to construct the corresponding infinity homography matrix. Simultaneously, the angular velocity and acceleration data output by the inertial measurement unit (IMU) within the corresponding time period are read to establish the temporal correspondence between the image frames and the inertial data. The time offset parameter, rotation extrinsic parameter, and displacement extrinsic parameter between the industrial camera and the IMU are iteratively solved. When the ambient temperature changes, the displacement extrinsic parameter and time offset parameter are corrected based on the current temperature. This implementation method enables online spatiotemporal calibration without the need for an additional dedicated calibration field, and the obtained calibration parameters can be directly used for subsequent fusion mapping processing.
[0085] Furthermore, the time offset parameter, which is the time delay between each industrial camera and the inertial measurement unit (IMU), is compensated using the aforementioned offline calibration combined with online correction. The rotational extrinsic parameter is updated online based on the homography constraint of the infinity plane, and the displacement extrinsic parameter is corrected using the offline calibrated temperature drift coefficient combined with online temperature measurement. During the mapping process, the time offset parameter and displacement extrinsic parameter are updated online according to temperature changes to reduce the impact of sensor installation relationship drift and sampling timing drift caused by temperature changes on the calibration results.
[0086] Specifically, the time offset parameter is the time delay between each industrial camera and the inertial measurement unit (IMU). It maps the image acquisition time of each industrial camera to the inertial data sampling time of the IMU, and uses the time difference between the two as the time delay parameter for determination and updating. Since there may be deviations in the data acquisition links and triggering sequences of different sensors, modeling and correcting the time delay parameter ensures that multi-view image data and inertial data remain consistent on the time axis, thus providing an accurate time basis for subsequent joint processing of visual and inertial observations.
[0087] In this embodiment, the extrinsic parameters include rotational and displacement extrinsic parameters between each industrial camera and the inertial measurement unit (IMU). The rotational extrinsic parameters characterize the attitude correspondence between the industrial camera coordinate system and the IMU coordinate system, while the displacement extrinsic parameters characterize the positional offset between them. During the mapping process, the rotational extrinsic parameters are optimized online in real-time using infinity plane homography constraints, while the displacement extrinsic parameters are corrected online based on temperature changes to reduce the impact of installation relationship drift caused by temperature variations on the calibration results. In one embodiment, when the measurement platform moves from an indoor environment to an outdoor environment, the time delay parameter and displacement extrinsic parameters can be corrected online based on the current temperature value collected by the temperature sensor to ensure that subsequent multi-view image data and inertial data can participate in fusion processing based on a consistent spatiotemporal relationship.
[0088] By introducing a temperature compensation mechanism, the time offset parameter and displacement extrinsic parameter are updated online as dynamic variables that change with temperature. At the same time, the rotation extrinsic parameter is optimized online based on the alignment of vision and inertial navigation. This enables the system to maintain the spatiotemporal alignment accuracy between the camera and IMU under varying temperature conditions such as day-night temperature differences or indoor-outdoor environment switching, thus solving the problem of cumulative error caused by temperature drift during long-term operation.
[0089] S3. Perform joint extraction and matching of point features and line features on the spatiotemporally calibrated multi-view image data to obtain multi-view feature observation information;
[0090] Furthermore, step S3 specifically includes:
[0091] Extract point and line features from multi-view image data after spatiotemporal calibration;
[0092] Match point features and line features separately, and perform consistency filtering based on the matching results;
[0093] Based on the selected point feature matching results and line feature matching results, multi-view feature observation information is obtained.
[0094] Specifically, point features and line features are extracted from multi-view image data after spatiotemporal calibration. Image points with obvious grayscale changes or texture features, as well as image line segments representing scene edges, contours or structural boundaries, can be detected in image frames acquired by each industrial camera. The extracted point features and line features are used as the visual feature information of the current image frame. By extracting point features and line features simultaneously, both texture and structural information in the image can be utilized, providing a more complete observation basis for subsequent feature matching and construction of multi-view feature observation information.
[0095] This involves matching point features and line features separately, and then performing consistency filtering based on the matching results. Point feature correspondences and line feature correspondences are established between synchronized images from different industrial cameras or between consecutive image frames from the same industrial camera. The matching results are then checked for consistency, and abnormal matches that do not meet the requirements of geometric correspondence or matching stability are eliminated. By performing consistency filtering on the point feature matching results and line feature matching results, visual features with reliable correspondences in multi-view images can be retained, thereby reducing the impact of incorrect matching on the subsequent construction of multi-view geometric constraints. In one embodiment, during the surveying of the building facade, point features can be extracted from the wall texture area, and line features can be extracted from the edges of doors and windows and the roof outline. Then, the point features and line features obtained from different perspectives are matched and filtered separately to obtain stable scene feature correspondences.
[0096] In this embodiment, multi-view feature observation information is obtained based on the filtered point feature matching results and line feature matching results. The point feature correspondence and line feature correspondence retained after consistency screening are uniformly organized so that they represent the observation results of the same scene features under different industrial camera perspectives or different acquisition times. The multi-view feature observation information can be used as input data for subsequent construction of multi-view geometric constraint factors and establishment of a vision-inertial navigation tightly coupled nonlinear optimization model, thereby providing visual observation basis for subsequent solution of pose parameters and three-dimensional structural parameters.
[0097] By jointly extracting and matching point and line features, even in areas with scarce texture (such as walls and floors), sufficient feature observations can be maintained by relying on structural edge information, providing a stable data foundation for the subsequent construction of multi-view geometric constraints.
[0098] S4. Construct multi-view geometric constraint factors based on the fixed baseline relationship between each industrial camera and multi-view feature observation information. Combine the IMU pre-integration obtained from the inertial data to construct a nonlinear optimization model that tightly couples vision and inertial navigation.
[0099] Furthermore, step S4 specifically includes:
[0100] Based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information, a multi-view geometric constraint factor is constructed.
[0101] A multi-view reprojection error term is established based on the multi-view geometric constraint factor, and the IMU pre-integration quantity is calculated based on the inertial data.
[0102] A nonlinear optimization model tightly coupled with vision and inertial navigation is constructed by using multi-view reprojection error terms and IMU pre-integration.
[0103] Specifically, a multi-view geometric constraint factor is constructed based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information. First, based on the fixed installation relationship of each industrial camera in the rigid combined measurement system, the relative pose relationship between any two industrial cameras is determined. Then, combined with the observation results of the same scene feature under different industrial camera views, the geometric correspondence of the scene feature under multiple views is established. Since the baseline relationship between each industrial camera remains unchanged during the measurement process, this fixed baseline relationship can be used to uniformly constrain the multi-view feature observation information, so that the observation results under different views are expressed under the same geometric relationship, thereby providing a basis for the subsequent establishment of a multi-view reprojection error model. In this embodiment, for the mapping scene of buildings, roads, or terrain areas, forward-facing industrial cameras, left-tilted industrial cameras, and right-tilted industrial cameras can be used to observe the same target area. Based on the corresponding positions of the same scene in the images of different industrial cameras and the fixed relative positional relationship between the industrial cameras, a multi-view geometric constraint factor is constructed.
[0104] A multi-view reprojection error model is established based on multi-view geometric constraint factors, and the IMU pre-integration is calculated based on inertial data. Specifically, for point features, the deviation between the predicted projection position and the actual observation position of the scene's 3D points under the views of each industrial camera can be used as the point feature reprojection error term, which can be expressed as:
[0105] ;
[0106] In the formula, Indicates the first The scene's 3D points are at the... Reprojection error from the perspective of an industrial camera Indicates the first The scene's 3D points are at the... The actual observation point coordinates in an industrial camera Represents the projection function. Indicates the first Pose transformation matrix of an industrial camera Indicates the first A scene of three-dimensional points.
[0107] For line features, the distance from a point to the projected line can be constructed as the geometric error term for the line feature. Let a spatial line be represented by two endpoints or a parametric equation, and its projection onto the image plane be a line segment. The line feature error is defined as the distance from the endpoint of the extracted line segment to the corresponding spatial line projection, which can be expressed as:
[0108] ;
[0109] In the formula, Indicates the first The spatial straight line in the th Line feature error from the perspective of an industrial camera The distance function from the point to the projected line is represented by the following: Indicates the first The spatial straight line in the th The endpoints of line segments extracted from an industrial camera image (or using the overall error of the line segment). Parametric representation of a straight line in space.
[0110] Meanwhile, the motion state between adjacent moments can be integrated based on the angular velocity and acceleration data output by the inertial measurement unit (IMU) to obtain the IMU pre-integral quantity, which characterizes the attitude change, velocity change and displacement change of the measurement system between adjacent moments. By establishing a multi-view reprojection error term and calculating the IMU pre-integral quantity, the motion state and scene structure information of the rigid combined measurement system can be described from both visual observation and inertial observation perspectives.
[0111] In this embodiment, a nonlinear optimization model tightly coupled with vision and inertial navigation is constructed using the multi-view reprojection error term and the IMU pre-integration. Specifically, the multi-view reprojection error term and the IMU pre-integration can be incorporated into a unified objective function to form an optimization model under joint constraints of vision and inertial navigation, which can be expressed as:
[0112] ;
[0113] In the formula, The nonlinear optimization objective function represents the tight coupling between vision and inertial navigation. The sum of squares of the feature reprojection errors for all points. This is the sum of squares of the geometric errors of all line features. This represents the sum of squares of the IMU pre-integration residuals. By minimizing this objective function, we can simultaneously utilize multi-view image observation information and inertial data within the same optimization framework to jointly solve for the pose parameters of the measurement system and the 3D structural parameters of the scene, thus providing a unified model foundation for subsequent sliding window local optimization and global loop closure detection.
[0114] Furthermore, the nonlinear optimization model includes multi-view geometric constraints and IMU pre-integration constraints, and uses the pose parameters, velocity parameters, IMU bias parameters, and three-dimensional structural parameters of the scene as optimization variables.
[0115] Specifically, the nonlinear optimization model includes multi-view geometric constraints and IMU pre-integration constraints. The multi-view geometric constraints characterize the consistency of scene features observed from different industrial camera perspectives, and are composed of the aforementioned point feature reprojection error term and line feature geometric error term. The IMU pre-integration constraints characterize the motion continuity of the rigid combined measurement system between adjacent time points, and are composed of the residual term formed by pre-integrating the angular velocity and acceleration data output by the inertial measurement unit (IMU). By incorporating the multi-view geometric constraints and IMU pre-integration constraints into the same nonlinear optimization model, visual and inertial observation information can be used simultaneously within a unified framework to jointly constrain the motion state of the measurement system and the scene structure, thus providing a consistent optimization basis for subsequent pose solving and 3D structure reconstruction.
[0116] In this embodiment, the nonlinear optimization model uses the pose parameters, velocity parameters, IMU bias parameters, and 3D structural parameters of the scene as optimization variables for the rigid combined measurement system. Specifically, the pose parameters characterize the position and attitude of the rigid combined measurement system at each sampling time; the velocity parameters characterize the motion velocity of the rigid combined measurement system at each sampling time; the IMU bias parameters characterize the gyroscope and accelerometer biases of the IMU during the measurement process; and the 3D structural parameters of the scene characterize the 3D spatial positions of each feature point in the scene under test. By using the above parameters together as optimization variables, the motion state of the measurement system, inertial measurement errors, and scene structure estimation errors can be simultaneously corrected during the optimization process, allowing visual observation and inertial observation to complement and constrain each other in the same solution process.
[0117] In one embodiment, the pose parameters, velocity parameters, IMU bias parameters, and 3D coordinates of corresponding scene feature points of the rigid combined measurement system at each moment within a certain time window can be uniformly included in the set of optimization variables. Combined with multi-view geometric constraints and IMU pre-integration constraints, the optimized parameter results can be directly used as the basic data for subsequent sliding window local optimization, loop closure detection, and 3D mapping result generation, thereby ensuring the consistency between pose estimation and scene structure restoration during the mapping process.
[0118] By combining local optimization with global loop closure detection through sliding window, cumulative errors can be effectively suppressed during long-distance mapping, ensuring the consistency of the final mapping results across the entire scope.
[0119] Through the tight coupling optimization of the aforementioned multi-view geometric constraints and IMU pre-integration, the various technical features of this invention form a complete synergistic effect: the non-coplanar layout provides a multi-angle observation basis for infinity plane calibration, accurate spatiotemporal calibration ensures the reliability of point and line feature extraction, rich feature observation information enhances the stability of multi-view geometric constraints, and dynamic temperature compensation ensures the parameter consistency of the entire system during long-term operations. This systematic design enables this invention to maintain high-precision mapping capabilities even under complex motion and temperature variations, solving engineering problems that traditional methods struggle to address simultaneously.
[0120] S5. Based on the nonlinear optimization model, perform sliding window local optimization and global closed-loop detection to solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test.
[0121] Furthermore, S5 specifically includes:
[0122] The nonlinear optimization model is locally optimized within a sliding window, and historical states that exceed the sliding window are marginalized.
[0123] The pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test are solved based on the local optimization results.
[0124] When a closed loop is detected, global consistency correction is performed on the pose parameters and 3D structural parameters.
[0125] Specifically, the nonlinear optimization model is locally optimized within a sliding window, and historical states that exceed the sliding window are marginalized. The sliding window can be formed by selecting the state variables of the current time and several previous time points. Within this window, the aforementioned multi-view geometric constraints and IMU pre-integration constraints are jointly optimized. For historical states that exceed the sliding window, their impact on the current estimation result is retained as prior information by marginalization to avoid the continuous increase in the number of optimization variables. Through the above processing, while ensuring that the computational scale is controlled, the visual observation information and inertial observation information of the most recent time period can be continuously utilized to provide a foundation for the real-time solution of subsequent pose parameters and three-dimensional structural parameters.
[0126] In this embodiment, the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test are solved based on the local optimization results. Specifically, the position and attitude information of the rigid combined measurement system at each sampling time can be obtained based on the optimized state variables within the sliding window, and the three-dimensional spatial coordinate information corresponding to the scene feature points can be obtained simultaneously. By jointly solving the pose parameters and the three-dimensional structural parameters, the motion state estimation results of the measurement system can be kept consistent with the scene structure restoration results, thereby providing reliable parameter inputs for subsequent closed-loop correction and the generation of three-dimensional mapping results.
[0127] Furthermore, upon detecting a loop closure, global consistency correction is performed on the pose parameters and 3D structural parameters. The current frame or current keyframe can be matched with historical keyframes. When it is determined that the current observation area overlaps with the historical observation area, a loop closure is identified. The pose parameters and 3D structural parameters obtained from the aforementioned local optimization are then adjusted globally to eliminate the deviations accumulated during long-term measurement. In one embodiment, when the measurement platform completes a full circle along the measurement area and re-passes through previously acquired building facades or road intersections, a loop closure can be detected based on the matching relationship between the current image and historical images. The measurement trajectory and the corresponding scene 3D structure are then uniformly corrected, thereby ensuring that the entire mapping result remains consistent globally.
[0128] S6. Generate the 3D mapping results of the scene to be tested based on the corrected pose parameters and 3D structural parameters.
[0129] Furthermore, S6 specifically includes:
[0130] Three-dimensional reconstruction is performed based on the corrected pose parameters and three-dimensional structural parameters;
[0131] Output the 3D mapping results of the scene to be tested. The 3D mapping results include at least one of 3D point cloud, 3D model and map data.
[0132] Specifically, based on the corrected pose parameters and 3D structural parameters, 3D reconstruction can be performed. The pose parameters of the rigid combined measurement system after global consistency correction can be used as the spatial pose basis of the images from each viewpoint, and the 3D structural parameters of the corresponding scene feature points can be used as the input of scene spatial geometric information. The scene to be measured can be 3D restored according to the observation relationship of each industrial camera. Through the above processing, the parameter results obtained by the aforementioned fusion optimization and closed-loop correction can be transformed into the spatial structural expression of the scene to be measured, thus providing a basis for the subsequent mapping results output.
[0133] In this embodiment, the three-dimensional mapping results of the scene to be measured are output. The three-dimensional mapping results include at least one of three-dimensional point cloud, three-dimensional model and map data. Specifically, the three-dimensional point cloud data of the scene can be directly output based on the three-dimensional reconstruction results, or a surface model, solid model or other form of three-dimensional model can be further generated based on the three-dimensional point cloud. The corresponding map data can also be formed based on the pose parameters and scene structure information. In one embodiment, for a building area mapping scene, the three-dimensional point cloud data of the building area can be output; for a terrain area or road area mapping scene, the corresponding three-dimensional model or map data can be further generated to meet the data usage needs of subsequent mapping applications.
[0134] Please see the appendix Figure 2A multi-view vision and inertial navigation fusion control point-free rapid mapping system, the system includes:
[0135] The data acquisition module constructs a rigid combined measurement system comprising at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and acquires multi-view image data and corresponding inertial data of the scene under test.
[0136] The spatiotemporal calibration module, based on the homography constraint of the infinity plane, performs online spatiotemporal calibration on multi-view image data and inertial data, determines the time offset parameters and external parameters between each industrial camera and the inertial measurement unit (IMU), and compensates for parameter changes caused by temperature drift.
[0137] The feature processing module performs joint extraction and matching of point and line features on multi-view image data after spatiotemporal calibration to obtain multi-view feature observation information.
[0138] The fusion modeling module constructs multi-view geometric constraint factors based on the fixed baseline relationship between each industrial camera and multi-view feature observation information. Combined with the IMU pre-integration obtained from inertial data, it constructs a nonlinear optimization model that tightly couples vision and inertial navigation.
[0139] The optimization and correction module performs sliding window local optimization and global closed-loop detection based on a nonlinear optimization model to solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test.
[0140] The results output module generates 3D mapping results of the scene under test based on the corrected pose parameters and 3D structural parameters.
[0141] Specifically, the data acquisition module is used to construct a rigid combined measurement system comprising at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar manner, and to acquire multi-view image data and corresponding inertial data of the scene under test. The spatiotemporal calibration module is used to perform online spatiotemporal calibration of the multi-view image data and inertial data based on the homography constraint of the infinity plane, determining the time offset parameters and extrinsic parameters between each industrial camera and the IMU, and compensating for parameter changes caused by temperature drift. The feature processing module is used to jointly extract and match point and line features from the spatiotemporally calibrated multi-view image data to obtain multi-view feature observation information. The fusion modeling module is used for… Based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information, a multi-view geometric constraint factor is constructed. Combined with the IMU pre-integration obtained from the inertial data, a nonlinear optimization model tightly coupled with vision and inertial navigation is constructed. The optimization and correction module is used to perform sliding window local optimization and global closed-loop detection based on the nonlinear optimization model, and solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test. The result output module is used to generate the three-dimensional mapping results of the scene under test based on the corrected pose parameters and three-dimensional structural parameters. Through the coordinated cooperation of the above modules, the multi-view vision and inertial navigation fusion mapping processing of the scene under test can be completed without the need to set up ground control points.
[0142] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for multi-view vision and inertial navigation fusion without control points for rapid mapping, characterized in that, include: S1. Construct a rigid combined measurement system including at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and collect multi-view image data and corresponding inertial data of the scene under test. S2. Based on the homography constraint of the infinity plane, online spatiotemporal calibration is performed on the multi-view image data and inertial data to determine the time offset parameters and external parameters between each industrial camera and the inertial measurement unit (IMU), and to compensate for parameter changes caused by temperature drift. S3. Perform joint extraction and matching of point features and line features on the spatiotemporally calibrated multi-view image data to obtain multi-view feature observation information; S4. Construct multi-view geometric constraint factors based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information, and combine the IMU pre-integration obtained from the inertial data to construct a nonlinear optimization model that tightly couples vision and inertial navigation. S5. Based on the nonlinear optimization model, perform sliding window local optimization and global closed-loop detection to solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test. S6. Generate the 3D mapping results of the scene to be measured based on the corrected pose parameters and 3D structural parameters; The S1 step specifically includes: At least three industrial cameras and inertial measurement units (IMUs) are mounted on the same rigid carrier to form a rigid combined measurement system; Adjust the mounting position and orientation of at least three industrial cameras to ensure that the at least three industrial cameras are arranged in a non-coplanar layout; The rigid combined measurement system is used to acquire multi-view image data of the scene under test and inertial data corresponding to the multi-view image data; Among the at least three industrial cameras, at least one industrial camera is arranged in a forward direction, and the other two industrial cameras are arranged in different tilt directions, and the optical axes of each industrial camera are not parallel to each other and are not in the same plane. The S4 step specifically includes: Based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information, a multi-view geometric constraint factor is constructed. A multi-view reprojection error term is established based on the multi-view geometric constraint factor, and the IMU pre-integration is calculated based on the inertial data; A nonlinear optimization model tightly coupled with vision and inertial navigation is constructed using the multi-view reprojection error term and the IMU pre-integral quantity. The nonlinear optimization model includes multi-view geometric constraints and IMU pre-integration constraints, and uses the pose parameters, velocity parameters, IMU bias parameters, and three-dimensional structural parameters of the scene as optimization variables. S5 specifically includes: The nonlinear optimization model is locally optimized within a sliding window, and historical states that exceed the sliding window are marginalized. The pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test are solved based on the local optimization results. When a closed loop is detected, global consistency correction is performed on the pose parameters and the three-dimensional structural parameters.
2. The multi-view vision and inertial navigation fusion's control point free rapid mapping method according to claim 1, characterized in that, The S2 step specifically includes: Extract distant features from multi-view image data, and construct an infinity plane homography matrix based on the distant features; The infinity plane homography matrix is correlated with the inertial data in time to establish a spatiotemporal constraint relationship between the industrial camera and the inertial measurement unit (IMU). The time offset parameters and extrinsic parameters between each industrial camera and the inertial measurement unit (IMU) are iteratively solved based on the spatiotemporal constraints, and the time offset parameters and extrinsic parameters are corrected in combination with temperature changes.
3. The multi-view vision and inertial navigation fusion-based rapid mapping method without control points according to claim 2, characterized in that, The time offset parameter is the time delay parameter between each industrial camera and the inertial measurement unit (IMU); The extrinsic parameters include rotational and displacement extrinsic parameters between each industrial camera and the inertial measurement unit (IMU). During the mapping process, the time offset parameters and displacement extrinsic parameters are updated online according to temperature changes. The rotational extrinsic parameters are optimized online based on the homography constraint of the infinity plane.
4. The multi-view vision and inertial navigation fusion-based rapid mapping method without control points according to claim 1, characterized in that, The S3 step specifically includes: Extract point and line features from multi-view image data after spatiotemporal calibration; The point features and line features are matched separately, and consistency filtering is performed based on the matching results; Based on the selected point feature matching results and line feature matching results, multi-view feature observation information is obtained.
5. The multi-view vision and inertial navigation fusion-based rapid mapping method without control points according to claim 1, characterized in that, S6 specifically includes: Three-dimensional reconstruction is performed based on the corrected pose parameters and three-dimensional structural parameters; Output the three-dimensional mapping results of the scene to be tested, wherein the three-dimensional mapping results include at least one of three-dimensional point cloud, three-dimensional model and map data.
6. A rapid mapping system without control points that integrates multi-view vision and inertial navigation, characterized in that, The system for the multi-view vision and inertial navigation fusion rapid mapping method without control points as described in any one of claims 1-5, the system comprising: The data acquisition module constructs a rigid combined measurement system comprising at least three industrial cameras and inertial measurement units (IMUs) arranged in a non-coplanar layout, and acquires multi-view image data and corresponding inertial data of the scene under test. The spatiotemporal calibration module, based on the homography constraint of the infinity plane, performs online spatiotemporal calibration on the multi-view image data and inertial data, determines the time offset parameters and external parameters between each industrial camera and the inertial measurement unit (IMU), and compensates for parameter changes caused by temperature drift. The feature processing module performs joint extraction and matching of point and line features on multi-view image data after spatiotemporal calibration to obtain multi-view feature observation information. The fusion modeling module constructs multi-view geometric constraint factors based on the fixed baseline relationship between each industrial camera and the multi-view feature observation information. Combined with the IMU pre-integration obtained from the inertial data, it constructs a nonlinear optimization model that tightly couples vision and inertial navigation. The optimization and correction module performs sliding window local optimization and global closed-loop detection based on the nonlinear optimization model to solve and correct the pose parameters of the rigid combined measurement system and the three-dimensional structural parameters of the scene under test. The results output module generates 3D mapping results of the scene under test based on the corrected pose parameters and 3D structural parameters.
Citation Information
Patent Citations
High-precision topographic surveying and mapping system and method based on unmanned aerial vehicle
CN120121039A
Method and device for positioning and mapping industrial automobile crane based on multi-sensing fusion
CN121346767A