Multi-camera radar-inertia-visual odometer method fusing sonar image

By employing a multi-camera radar-inertial-visual odometry method and fusing multimodal sensor data using a unified nonlinear optimization framework, the problem of low positioning accuracy of traditional sensors in complex environments is solved, enabling high-precision and robust navigation for cross-medium aircraft.

CN120949245APending Publication Date: 2025-11-14BEIHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511269710.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional single-sensor solutions have low positioning accuracy in complex environments, and sensor performance fluctuates across different media, lacking adaptive adjustment capabilities, making it difficult to achieve high-precision and robust positioning and navigation.

Method used

A multi-camera radar-inertial-visual odometry method is adopted, which integrates visual, sonar, lidar and inertial data through a unified nonlinear optimization framework, and combines medium detection and mode selection to achieve tightly coupled multimodal information fusion.

Benefits of technology

Achieving high-precision and stable positioning and navigation in different media environments enhances the environmental adaptability of cross-media aircraft, strengthens the system's robustness and computational efficiency, and is suitable for autonomous navigation in a variety of complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949245A_ABST
    Figure CN120949245A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-camera radar-inertia-visual odometer method fusing a sonar image, and belongs to the technical field of robot perception and navigation. According to the method, environment perception is realized through cooperative work of multiple sensors; in the air, a multi-camera system and a conventional laser radar are taken as main modes; in water, an active imaging sonar and an underwater laser radar are used as main modes. The system detects the medium of the aircraft according to the resistivity, dynamically switches the main mode and fuses the auxiliary mode information. Through a tight coupling optimization framework, observation data of vision, sonar, laser radar and an inertial measurement unit are uniformly modeled as residual constraints, a nonlinear least square problem is constructed and solved, and efficient fusion of multi-source information is realized. The method has robustness in low-texture, low-illumination and underwater turbid environments, and the positioning precision and stability in a cross-medium scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot perception and navigation technology, specifically relating to a multi-camera radar-inertial-visual odometry method that fuses sonar images. Background Technology

[0002] With the development of underwater robots, unmanned underwater vehicles (AUVs / ROVs), and cross-medium aircraft, the demand for high-precision positioning and navigation technologies in complex environments is increasing. However, traditional single-sensor solutions have significant limitations: pure visual odometry is prone to failure in low-texture, low-light, or turbid underwater environments; while inertial measurement units (IMUs) can provide short-term high-frequency attitude information, they suffer from cumulative drift; sonar can provide distance and orientation information in underwater environments, but its accuracy is limited by beamwidth and sound wave attenuation; lidar has excellent geometric perception capabilities in air, but it is susceptible to refraction and scattering interference underwater.

[0003] Existing technologies typically employ loosely coupled methods to fuse multi-sensor data, such as combining observations from different modalities through filtering or simple weighting. This approach struggles to fully utilize the complementarity between sensors and is prone to decreased positioning accuracy due to information redundancy or conflicts. Furthermore, in cross-media scenarios, sensor performance fluctuates with environmental changes, and traditional methods lack adaptive adjustment capabilities, making it difficult to guarantee system robustness.

[0004] Therefore, there is an urgent need for a tightly coupled multimodal fusion method that can dynamically select the main sensor mode and integrate visual, sonar, lidar and inertial data based on a unified optimization framework, so as to achieve high-precision and robust positioning and navigation in air, underwater and cross-medium environments. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a multi-camera radar-inertial-visual odometry method that integrates sonar images. Through a unified nonlinear optimization framework, visual, IMU, sonar, and lidar observation information are modeled in the same state estimation problem, and sliding window optimization is used to achieve high-precision, robust positioning across media.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A multi-camera radar-inertial-visual odometry method fusing sonar images, the method comprising:

[0008] Step 1, Multimodal sensor layout and calibration: Fix the multi-camera array, IMU, sonar and lidar on the same carrier according to the actual application scenario, and calibrate the external parameters and time synchronization of each sensor;

[0009] Step 2, Medium Detection and Mode Selection: The medium in which the aircraft is located is determined by measuring the resistance. If it is air, the visible light imaging mode of the multi-camera array is the main mode; if it is water, the ultrasonic imaging mode of the sonar system is the main mode.

[0010] Step 3, Multimodal residual modeling: Construct visible light imaging visual residuals or ultrasonic imaging residuals based on the selected primary modality, and filter auxiliary modal data based on information gain and redundancy, and calculate the fusion weights of the auxiliary modal data;

[0011] Step 4, Tightly Coupled Optimization Solution: Unify the observation constraints of the multimodal sensors into nonlinear least squares optimization, iteratively update the state variables of the multimodal sensors, and realize cross-medium tightly coupled fusion positioning.

[0012] In a second aspect, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned multi-camera radar-inertial-visual odometry method for fusing sonar images.

[0013] Thirdly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned multi-camera radar-inertial-visual odometry method that fuses sonar images.

[0014] The beneficial effects of this invention are as follows:

[0015] Cross-medium adaptive positioning: By automatically switching the main sensor mode through resistivity detection (mainly camera and conventional lidar in air, and sonar and underwater lidar underwater), it can ensure stable operation in different media and significantly improve the environmental adaptability of cross-medium aircraft.

[0016] High-precision tightly coupled fusion: A unified nonlinear optimization framework is adopted to model the observation data of vision, sonar, lidar and IMU as tightly coupled residual constraints. The efficient fusion of multi-source information is achieved through sliding window optimization, which suppresses the drift problem of single sensor and the positioning accuracy can reach the centimeter level.

[0017] Enhanced robustness in complex environments: Under harsh conditions such as low texture, low light, or underwater turbidity, geometric constraints are provided by sonar and lidar to compensate for visual failures and ensure the stability and reliability of the system in diverse scenarios.

[0018] Computational efficiency optimization: Auxiliary modal data is dynamically filtered based on information gain and redundancy to avoid invalid computation. At the same time, sparse optimization methods are used to reduce computing power consumption, so as to achieve efficient state estimation under real-time requirements.

[0019] Versatility and scalability: The modular design supports flexible configuration of different sensors and can be adapted to various application scenarios such as underwater robots, drones and cross-media platforms, providing a universal solution for autonomous navigation in complex environments. Attached Figure Description

[0020] Figure 1 This is a flowchart of a multi-camera radar-inertial-visual odometry method that integrates sonar images according to the present invention. Detailed Implementation

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] like Figure 1 As shown, this invention proposes a multi-camera-LiDAR-inertial tightly coupled odometry method that fuses sonar images. The system includes a multi-camera array, an inertial measurement unit (IMU), sonar sensors, and a LiDAR. Each sensor uses a unified clock synchronization module for data time alignment and an extrinsic parameter calibration module to obtain relative pose relationships. During system operation, the data acquired by the multi-modal sensors is input into a tightly coupled optimization module. Within a sliding window, each state variable is simultaneously estimated, and each modal observation is modeled as a corresponding constraint factor, unified into a nonlinear optimization problem for solution. Specifically, the method includes:

[0023] Step 1, Multimodal sensor layout and calibration: Fix the multi-camera array, IMU, sonar and lidar on the same carrier according to the actual application scenario, and calibrate the external parameters and time synchronization of each sensor;

[0024] Step 2, Medium Detection and Mode Selection: The medium in which the aircraft is located is determined by measuring the resistance. If it is air, the visible light imaging mode of the multi-camera array is the main mode; if it is water, the ultrasonic imaging mode of the sonar system is the main mode.

[0025] Step 3, Multimodal residual modeling: Construct a visible light imaging visual residual or ultrasonic imaging residual model based on the selected primary mode, and filter auxiliary modal data based on information gain and redundancy, and calculate the fusion weight of the auxiliary modal data;

[0026] Step 4, Tightly Coupled Optimization Solution: Unify the observation constraints of the multimodal sensors into nonlinear least squares optimization, iteratively update the state variables of the multimodal sensors, and realize cross-medium tightly coupled fusion positioning.

[0027] The steps described above are explained in detail below.

[0028] In step 1, the multi-camera array, inertial measurement unit (IMU), sonar sensor, and lidar are mounted on the same carrier platform. To ensure accurate spatial relationships among the sensors during information fusion, external parameter calibration and time synchronization are required after installation.

[0029] For extrinsic parameter calibration, a visual-inertial joint calibration method based on checkerboard calibration and time synchronization is adopted between the multi-camera array and the IMU to obtain the rotation matrix and translation vector of the camera coordinate system relative to the IMU coordinate system. For the sonar sensor and the IMU, a specific underwater calibration target and pose control platform are used to solve the extrinsic parameters from the sonar sensor to the IMU by minimizing the registration error between the sonar observation point and the known geometric target point. For the lidar and the IMU, the extrinsic parameters from the lidar to the IMU are obtained by using a multi-view point cloud registration and inertial trajectory alignment method.

[0030] In terms of time synchronization, a unified clock synchronization module is adopted to unify the data collected by multiple camera arrays, IMU, sonar sensors and lidar under the same timestamp, so as to ensure that the data of each mode corresponds to the same state time when fusion calculation.

[0031] This step ensures precise spatial and temporal alignment of the multimodal sensors, providing accurate observational geometry for subsequent tightly coupled information fusion, thereby improving the accuracy and consistency of the optimization results.

[0032] In step 2, during multimodal information fusion, to avoid invalid or redundant data interfering with the optimization results, it is necessary to select the primary and auxiliary modes for the multimodal sensors. Utilizing the difference in resistivity between water and air, a reference resistance value is set. The resistance of the medium in which the aircraft is located is measured through a pair of metal contacts located on the aircraft fuselage and related circuitry. If the measured resistance value is greater than the reference resistance value, the aircraft is determined to be in an air medium, and the visible light imaging mode of the multi-camera system is selected as the primary mode; if it is less than the reference resistance value, the aircraft is determined to be in a water medium, and the ultrasonic imaging mode of the sonar system is selected as the primary mode. Resistivity monitoring offers the following advantages: 1. High accuracy in medium identification. The resistivity difference between water and air is significant, and stable identification can be achieved through metal contacts arranged on the cross-medium aircraft, independent of external lighting. 2. Strong anti-interference capability. Compared to image or acoustic discrimination, resistivity is less affected by environmental texture, lighting, and noise interference, making it more reliable. It can operate normally even in turbid water or low-light environments. 3. Fast real-time performance and response speed. Resistance measurement is a low-latency, low-computation physical detection method that can quickly switch the main mode of an aircraft across media, avoiding the adaptive lag problem of vision and sonar-related mode selection algorithms.

[0033] In step 3, constructing the ultrasonic imaging residual of an active imaging sonar system or the visible light imaging visual residual of a multi-camera system includes:

[0034] For feature points observed by the camera If its corresponding three-dimensional point is Camera extrinsic parameters are ,in: Let be the rotation matrix of the camera coordinate system relative to the body coordinate system. Let be the translation matrix of the camera coordinate system relative to the body coordinate system. The pose is... , respectively represent: Indicates the first The rotation matrix of the machine's coordinate system relative to the world coordinate system at any given time. Indicates the first The translation matrix of the body coordinate system relative to the world coordinate system at any given time is used to first translate the feature points. Transformation from world coordinates to camera coordinates:

[0035] ,

[0036] Project and calculate the reprojection error:

[0037] ,

[0038] in For camera projection function (considering intrinsic parameters and distortion). For the first The first frame observed by the camera The actual pixel coordinates of a 3D point The required reprojection error is the difference between the "actual observed pixel position" and the "predicted pixel projection position".

[0039] Structured / unstructured visual factors are used (Structureless BA is optional to save variables), and reprojection errors are considered. Apply information matrix This yields the visual residual of visible light imaging.

[0040] Active imaging sonar typically uses polar coordinates (range measurement) With angle This can be represented as a two-dimensional intensity map. For sonar observations, a geometric measurement model (echo point mapping) is used as a residual model: feature points (such as obvious boundaries or echo peaks) are extracted from the sonar image and mapped to physical direction angles. Distance from the ray Polar coordinate error can be constructed:

[0041] ,

[0042] in For a point in a given state Predicted polar coordinates , This indicates the actual measured distance and the corresponding actual azimuth angle corresponding to the peak value (or feature point) of the sonar echo.

[0043] The process involves determining which auxiliary information in the auxiliary modes is non-redundant and usable for assisted localization, and establishing the fusion weights for the auxiliary information. When the visible light imaging mode of a multi-camera system is selected as the primary mode, the IMU, sonar, and underwater lidar are auxiliary modes; when the ultrasonic imaging mode of a sonar system is selected as the primary mode, the multi-camera array, IMU, sonar, and conventional lidar are auxiliary modes. The selection of auxiliary information specifically includes using an information gain index to quantify the degree to which observations reduce the uncertainty of state estimation. The calculation formula is as follows:

[0044] ,

[0045] Where P is the state covariance matrix of the current mode. Let be the Jacobian matrix corresponding to the observation of the i-th sensor. Let be the observation noise covariance matrix of the i-th sensor. This indicates the matrix trace operation, and the superscript T indicates transpose.

[0046] Simultaneously, the observation redundancy between the new mode and the current mode is calculated:

[0047] ,

[0048] If information gain Above the threshold And redundancy Below the threshold If the new mode is selected as the auxiliary mode for the current optimization, it will replace the current mode. To balance the contributions of different modes in the optimization, the Softmax normalization method is used to calculate the fusion weights.

[0049] ,

[0050] The observation noise covariance is then weighted and adjusted.

[0051] ,

[0052] In this way, modes with high information gain have a larger weight in optimization, while the weight of modes with redundant or unstable information is automatically reduced.

[0053] This step enables the adaptive selection of auxiliary modes that are valuable to the current localization task, effectively suppresses the interference of redundant information on optimization, and enhances the robustness of the system in different environments.

[0054] In step 4, the state variables are estimated based on the tightly coupled odometry, and the observation constraints of multimodal sensors (vision, inertial measurement unit, sonar, and lidar) are unified into the same optimization problem to construct a multi-source information tightly coupled fusion model. The specific steps include the following:

[0055] Step 4.1, State Variable Definition: Let the state variable of the k-th keyframe in the sliding window be:

[0056] ,

[0057] in ∈SO(3) is a rotation matrix ∈R³ is a translation vector. ∈R³ represents velocity. The zero biases are for the accelerometer and gyroscope, respectively.

[0058] Step 4.2: IMU pre-integration factor construction. Within the sliding window, considering the IMU measurements between keyframes n and m, a Lie group-based pre-integration method is used to integrate the raw accelerometer and gyroscope data to construct the IMU residuals.

[0059] ,

[0060] in: , These are the velocities of the nth and mth keyframes in the world coordinate system, respectively. These are the position vectors of the nth and mth keyframes in the world coordinate system, respectively. The translation increment obtained by IMU pre-integration represents the predicted displacement change after removing the zero bias effect; The velocity increment obtained by IMU pre-integration represents the predicted velocity change. The rotation increment obtained from pre-integration represents the predicted attitude change, and g is the gravitational acceleration vector. This represents the time interval between keyframes n and m. is a logarithmic mapping operator on the SO(3) group, used to transform the rotation error matrix into a three-dimensional vector.

[0061] Considering the visual reprojection factor, for keyframes The first in The observation residuals for each feature point are:

[0062] ,

[0063] in For pixel coordinate observation, For the first The three-dimensional position of each feature point in the world coordinate system. The position of keyframe k in the world coordinate system; For camera projection model, This is the modal selection coefficient, which is 1 when the camera is the dominant mode and 0 when it is not the dominant mode.

[0064] Considering sonar observation factors, the sonar polar coordinate observations are... Switch to the world system and match the map points Alignment, residual is:

[0065] ,

[0066] in, The echo distance measured by sonar. , where is the predicted target distance; For three-dimensional points Two-dimensional coordinates projected onto the sonar imaging plane. This is the mode selection factor, which is 1 when sonar is the dominant mode and 0 when it is not the dominant mode.

[0067] Considering lidar, the LiDAR geometric factor adopts a point-to-plane error model. For the th laser points The residual is:

[0068] ,

[0069] in For the first The three-dimensional coordinates of the laser point in the world coordinate system. Match it with a reference point on the plane. Let be the normal vector of the plane. , For keyframes The rotation matrix and position vector are calculated. When the aircraft is underwater, underwater lidar data is used; when it is in the air, conventional lidar data is used.

[0070] Step 4.3, Problem Construction: Unify all the above factors into the same nonlinear least squares optimization, with the objective function as follows:

[0071] ,

[0072] in Represents the corresponding set of observations. To observe the covariance matrix, (·) represents the robust kernel function. Where, These are the IMU pre-integration residual, visual reprojection residual, sonar polar coordinate residual, and lidar point-area residual, respectively. These are the corresponding observation noise covariance matrices.

[0073] Step 4.4, Solving and Updating: Iteratively solve using the sparse Gauss-Newton method or the Levenberg-Marquardt method, updating the state variables of all keyframes. and map points This enables tight coupling and fusion positioning of multimodal sensors.

[0074] In a second aspect, the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned multi-camera radar-inertial-visual odometry method for fusing sonar images.

[0075] Thirdly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned multi-camera radar-inertial-visual odometry method that fuses sonar images.

[0076] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-camera radar-inertial-visual odometry method that fuses sonar images, characterized in that, The method includes: Step 1, Multimodal sensor layout and calibration: Fix the multi-camera array, IMU, sonar and lidar on the same carrier according to the actual application scenario, and calibrate the external parameters and time synchronization of each sensor; Step 2, Medium Detection and Mode Selection: The medium in which the aircraft is located is determined by measuring the resistance. If it is air, the visible light imaging mode of the multi-camera array is the main mode; if it is water, the ultrasonic imaging mode of the sonar system is the main mode. Step 3, Multimodal residual modeling: Construct visible light imaging visual residuals or ultrasonic imaging residuals based on the selected primary modality, and filter auxiliary modal data based on information gain and redundancy, and calculate the fusion weights of the auxiliary modal data; Step 4, Tightly Coupled Optimization Solution: Unify the observation constraints of the multimodal sensors into nonlinear least squares optimization, iteratively update the state variables of the multimodal sensors, and realize cross-medium tightly coupled fusion positioning.

2. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, In step 1, regarding extrinsic parameter calibration, for the multi-camera array and IMU, a visual-inertial joint calibration method based on checkerboard calibration and time synchronization is adopted to obtain the rotation matrix and translation vector of the camera coordinate system relative to the IMU coordinate system; for the sonar sensor and IMU, using a specific underwater calibration target and pose control platform, the extrinsic parameters from the sonar sensor to the IMU are solved by minimizing the registration error between the sonar observation point and the known geometric target point; for the lidar and IMU, the extrinsic parameters from the lidar to the IMU are obtained using a multi-view point cloud registration and inertial trajectory alignment method.

3. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, In step 1, regarding time synchronization, a unified clock synchronization module is used to unify the data collected by the multi-camera array, IMU, sonar sensor and lidar under the same timestamp, so as to ensure that the data of each mode corresponds to the same state time when fusion calculation.

4. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, In step 2, a reference resistance value is set by utilizing the difference in resistivity between water and air. The resistance of the medium in which the aircraft is located is measured by a pair of metal contacts located on the aircraft body. If the measured resistance value is greater than the reference resistance value, it is determined that the aircraft is in an air medium. If it is less than the reference resistance value, it is determined that the aircraft is in a water medium.

5. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, In step 3, when constructing the ultrasonic imaging residual, the distance and angle errors are calculated based on the polar coordinates of the feature points extracted by sonar and the predicted values; when constructing the visible light imaging residual, the visual residual is calculated based on the reprojection error of the feature points observed by the camera, and information matrix weighting is applied respectively.

6. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, In step 3, the information gain is determined by calculating the trace change of the observation with respect to the state covariance matrix, the redundancy is evaluated by the correlation between observations from different sensors, and the fusion weight is calculated by using the Softmax normalization method to calculate the exponentially weighted normalized value of the information gain of each mode.

7. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, Step 4 includes defining the keyframe state variables within the sliding window; establishing four residual models: IMU pre-integration, visual reprojection, sonar polar coordinates, and lidar; and constructing a unified nonlinear least squares optimization objective function for state estimation and map point optimization.

8. The multi-camera radar-inertial-visual odometry method according to claim 1, characterized in that, In step 4, the nonlinear least squares optimization is solved iteratively using the sparse Gauss-Newton method or the Levenberg-Marquardt method.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the multi-camera radar-inertial-visual odometry method for fusing sonar images as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, enable the processor to implement the multi-camera radar-inertial-visual odometry method according to any one of claims 1-8.

Citation Information

Cited By

  • Underwater positioning system of underwater robot based on man-machine interaction

    CN120949243A

  • An underwater positioning system for underwater robots based on human-machine interaction

    CN120949243B

  • Tight coupling visual inertial odometer method and system based on infrared and visible light fusion image and uncertainty information thereof

    CN121855551A