Dual-view cooperative perception and adaptive weighted positioning method and system for underwater robots in confined spaces
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA YANGTZE POWER
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-04
AI Technical Summary
(1)GNSS信号完全拒止:高频电磁波在水体中衰减极快,导致全球导航卫星系统(GNSS)在机器人入水瞬间即完全失效,机器人失去绝对地理坐标参考,无法进行初始化定位
1、基于简化的Jaffe-McGlamery水下成像物理模型,通过红通道补偿纠正水下图像颜色失真问题,结合限制对比度自适应直方图均衡化(CLAHE)提升图像局部对比度,有效抑制水中悬浮颗粒散射带来的雾状模糊与背景噪声;同时通过计算对比度增强率量化图像质量,为后续融合提供精准依据,显著提高了视觉标记特征检测的成功率,适配水下弱纹理、低光照的恶劣光学环境。
Smart Images

Figure CN122510340A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of underwater robot control technology, specifically, it relates to a dual-field-of-view cooperative perception and adaptive weighted localization method and system for underwater robots in confined spaces. Background Technology
[0002] Large-scale hydropower infrastructures typically possess enormous physical dimensions (hundreds of meters in height) and massive reservoir capacities. Traditionally, structural integrity assessments of dam surfaces (such as crack detection and wear analysis) have relied primarily on professional divers using tactile sensing or handheld camera techniques. However, this manual approach faces significant challenges in extreme underwater environments: divers not only endure immense water pressure and the risk of decompression sickness but are also limited by physiological constraints, resulting in short operation times, low coverage, and a lack of standardized data, making it difficult to meet the demands of digital operation and maintenance throughout the entire lifecycle.
[0003] Therefore, a consensus has been reached within the industry to shift from manual intervention to robotic operations. ROVs or AUVs equipped with high-definition optical cameras, multi-beam sonar, and laser scanners promise to achieve intelligent, comprehensive defect detection. However, for robots to perform sophisticated inspections on dam surfaces, the primary prerequisite is centimeter-level precision positioning, which remains a global challenge that has not yet been fully resolved in confined underwater environments.
[0004] The so-called "restricted underwater environment" refers not only to the narrowness of the physical space, but also to the dual limitations in the dimensions of "signal and perception." Existing conventional underwater positioning technologies all have significant shortcomings in such scenarios: (1) Complete rejection of GNSS signals: High-frequency electromagnetic waves attenuate very quickly in water, causing the Global Navigation Satellite System (GNSS) to completely fail the moment the robot enters the water. The robot loses its absolute geographic coordinate reference and cannot perform initial positioning.
[0005] (2) Limitations of acoustic positioning systems: Ultra-short baseline (USBL) failure: USBL systems rely on acoustic phase difference for positioning. When working on the dam facade, the huge concrete wall forms a strong acoustic reflection interface, resulting in severe multipath effects. The superposition of direct waves and reflected waves leads to phase ambiguity, and the positioning error can reach several meters or even diverge.
[0006] (3) Doppler Velocimetry (DVL) blind zone and loss of lock: DVL calculates velocity by measuring the Doppler frequency shift of seabed echoes. However, when the robot works close to the wall to obtain high-resolution images, it often enters the minimum measurement distance blind zone of DVL, resulting in loss of lock and inability to provide velocity updates.
[0007] (4) Strong interference with the magnetic compass: The dense steel mesh inside the dam and the strong electromagnetic field generated by the operation of the large hydro-turbine generator unit will produce complex hard iron and soft iron effects. This makes the reading of the electronic compass based on the fluxgate principle extremely unreliable, and the heading angle deviation may exceed tens of degrees, causing the dead reckoning to diverge rapidly.
[0008] (5) Degradation of visual perception: Although visual positioning is theoretically highly accurate, the underwater optical environment is extremely harsh: Illumination attenuation and color distortion: Water absorption of light is wavelength selective, with red light attenuating the fastest, resulting in images appearing in a blue-green hue and losing color features. Contrast reduction caused by scattering: Backscattering caused by suspended particles in the water forms background noise similar to "fog," significantly reducing image contrast; forward scattering leads to blurred details. Weak texture features: The surface of the dam is mostly repetitive concrete texture, lacking significant corner or edge features. Traditional sparse feature methods are prone to feature tracking failure and scale drift in such weak texture areas.
[0009] To address the aforementioned issues, existing research mainly focuses on monocular vision localization or simple multi-sensor fusion, but the following technical bottlenecks still exist: (1) Planar pose ambiguity: Existing technologies mostly use iterative PnP (Perspective-n-Point) solvers (such as the Levenberg-Marquardt algorithm) to calculate camera pose. However, when observing planar targets (such as visual markers on the surface of a dam), the PnP problem has multiple solutions. Especially when underwater refraction causes the light path to bend, the iterative method is prone to getting trapped in local minima, resulting in a "flip error" in the calculated pose, that is, the surface normal vector is reversed by 180 degrees, which seriously affects the control stability.
[0010] (2) The contradiction between field of view and accuracy: Monocular systems have difficulty in balancing long-distance guidance (requiring narrow field of view telephoto lenses) and close-range operations (requiring wide field of view wide-angle lenses). During the process of approaching the dam, the target is often lost or the whole picture cannot be observed due to the distance being too close.
[0011] (3) Limitations of the static noise assumption: Traditional Kalman filtering (KF) or extended Kalman filtering (EKF) frameworks typically assume that the observation noise covariance matrix is a static constant. However, the turbidity and illumination conditions of the underwater environment are dynamic, and the reliability of visual observations fluctuates drastically. The static matrix cannot adapt to this change, resulting in the filter being unable to cut off erroneous observations in time when vision fails, or failing to converge quickly when vision recovers.
[0012] In summary, there is an urgent need to develop an adaptive localization framework that can integrate heterogeneous dual-view information, possess physical perception image enhancement capabilities, and be robust to planar pose calculation. Summary of the Invention
[0013] The technical problem to be solved by the present invention is to provide a dual-field-of-view cooperative perception and adaptive weighted localization method and system for underwater robots in confined spaces. By combining image enhancement, analytical pose estimation and adaptive quality weighted fusion algorithm, it achieves high robustness and high precision localization in complex underwater environments.
[0014] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces, comprising the following steps: S1. Data Acquisition: Environmental images, including main field-of-view images and auxiliary field-of-view images, are acquired using a heterogeneous dual-field-of-view camera system mounted on an underwater robot and pre-calibrated with external parameters. S2. Image Enhancement: Based on the underwater imaging physical model, the acquired raw images are enhanced, and then the image quality index is calculated. The enhancement process includes red channel compensation and contrast-limited adaptive histogram equalization. S3, Pose Calculation: Using an analytical pose estimation algorithm based on a small plane, visual markers in the enhanced main and secondary field-of-view images are detected and their poses are calculated to obtain independent pose observations of the visual markers from the two cameras. S4. Weight Calculation: Based on the image quality index obtained in step S2 and the independent pose observations of the two cameras obtained in step S3, the observation distance of the visual marker is calculated by taking the Euclidean norm of the three spatial coordinate components of the translation vector, and the fusion weights are calculated for the pose observations of the two cameras respectively. S5. Adaptive fusion: The pose observations of the two cameras are weighted and fused using fusion weights to obtain the visual pose estimate of the robot's body center. S6. State Filtering: Construct an error state Kalman filter, take the visual pose estimate obtained in step S5 as the observation input, and dynamically adjust the observation noise covariance matrix of the filter according to the fusion weights to output the robot's final position and posture.
[0015] In a preferred embodiment, in step S2, the underwater imaging physical model is a simplified Jaffe-McGlamery model.
[0016] In a preferred embodiment, in step S2, the image quality index used is the contrast enhancement rate, and the formula for calculating the contrast enhancement rate is: ; in, The contrast of the original image. To enhance the contrast of the resulting image, the contrast is defined as the standard deviation of the image grayscale or the mean of the local gradient.
[0017] In the preferred embodiment, the specific process of the analytical pose estimation algorithm based on the microplane in step S3 is as follows: S31. Extract the pixel coordinates of the four corner points of the visual marker on the image plane; S32. Using the camera intrinsic parameter matrix, calculate the homography matrix between the image plane and the marker physical plane; S33. Perform singular value decomposition on the homography matrix and solve analytically to obtain two sets of possible rotation matrix and translation vector assumptions; S34. Calculate the reprojection error of the two sets of assumptions on the image plane respectively, and select the set with the smallest reprojection error and satisfying the positive physical depth constraint as the optimal pose solution, and use it as the camera pose observation value.
[0018] In a preferred embodiment, in step S4, the fusion weight is the product of the image quality factor and the geometric uncertainty factor, and the fusion weight is calculated using the following formula: ; in, , L Indicates the left camera. R Indicates the right camera; Indicates camera i The fusion weights; Image quality factor This represents the geometric uncertainty factor.
[0019] In a preferred embodiment, the image quality factor reflects the contribution of the image's inherent sharpness to the solution accuracy, based on the contrast enhancement rate in step S4. The calculation formula is as follows: ; in, It is the adjustment coefficient.
[0020] In the preferred embodiment, the geometric uncertainty factor reflects the impact of the observation geometry on accuracy, and the calculation formula is as follows: ; in, Let Euclidean distance be the distance from the camera to the visual marker. It is a non-zero small constant.
[0021] In the preferred embodiment, step S5 specifically involves the following operations: For the depth coordinates of the robot's center Weighted fusion is performed to obtain the weighted fused depth coordinates. The calculation formula is: ; Among them, subscript L , R They represent the left and right cameras, respectively. and These represent the depth observation values from the left and right cameras, respectively; For the horizontal coordinate of the robot's center and By employing geometric midpoint calculation based on rigid body geometric constraints, the optimal estimate of the horizontal coordinates of the robot's body center after binocular data fusion was obtained. : ; in, , These represent the horizontal plane observation coordinate vectors extracted from the pose observations of the left and right cameras after being transformed into a unified robot body center coordinate system.
[0022] In a preferred embodiment, the method for dynamically adjusting the observation noise covariance matrix of the filter in step S6 is as follows: ; in: The nominal observation noise covariance matrix is preset. This is the scaling factor; , These represent the fusion weights for the left and right cameras, respectively. This is the adjusted observation noise covariance matrix.
[0023] In a preferred embodiment, the method is applied to underwater environments where GNSS signals are blocked, strong acoustic multipath effects exist, and weak texture features are present, including but not limited to the water-facing surface of hydraulic dams, diversion tunnels, water diversion tunnels, and the surfaces of deep-sea engineering structures.
[0024] This invention also provides a dual-field-of-view cooperative perception and adaptive weighted localization system for underwater robots in confined spaces, used to perform the above-described method, comprising: The perception module includes a wide-angle camera and a narrow-field-of-view gimbal camera with pre-calibrated extrinsic parameters to cover the field of view for short-range operations and remote guidance; as well as an inertial measurement unit and a depth sensor; The enhancement module is configured to perform Jaffe-McGlamery model parameter estimation and CLAHE algorithm to improve underwater image contrast and calculate contrast enhancement rate; The solution module is configured to run the IPPE algorithm to resolve the six-DOF pose of visually labeled images from the augmented images and eliminate pose blur. The fusion module is configured to execute the AQWF algorithm and dynamically calculate the binocular fusion weights based on the image contrast enhancement rate and observation distance. The navigation module is configured to run the R-adaptive ESKF algorithm, which fuses visual pose and inertial data to output the robot's state.
[0025] In a preferred embodiment, the visual marker is an AprilTag QR code with known geometric dimensions, and the system is also equipped with an illumination unit to provide auxiliary light sources in deep-water environments.
[0026] The present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and running on the memory, wherein the processor executes the computer program to implement the above-described method for dual-field-of-view cooperative perception and adaptive weighted localization of underwater robots in confined spaces.
[0027] The present invention also provides a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-described method for dual-field-of-view cooperative perception and adaptive weighted localization of underwater robots in confined spaces.
[0028] This invention provides a dual-field-of-view cooperative perception and adaptive weighted localization method and system for underwater robots in confined spaces, which has the following beneficial effects: 1. Based on a simplified Jaffe-McGlamery underwater imaging physical model, the system corrects color distortion in underwater images through red channel compensation and enhances local image contrast by combining Limiting Contrast Adaptive Histogram Equalization (CLAHE), effectively suppressing fogging and background noise caused by scattering from suspended particles in the water. At the same time, the system quantifies image quality by calculating the contrast enhancement rate, providing accurate basis for subsequent fusion and significantly improving the success rate of visual marker feature detection, making it suitable for harsh optical environments with weak textures and low light levels underwater.
[0029] 2. The analytical pose estimation algorithm based on microplanes (IPPE) is used to solve the camera pose. The pose assumption is obtained directly by singular value decomposition of homography matrix. The optimal solution is selected by combining reprojection error and physical depth positive value constraint. Compared with the traditional iterative PnP algorithm, it fundamentally avoids the problem of getting trapped in local minima during the iteration process. It completely solves the normal flipping error of planar target pose calculation in underwater refraction environment, ensures the stability of robot control, and realizes robust calculation of six degrees of freedom pose.
[0030] 3. An adaptive quality-weighted fusion algorithm (AQWF) is proposed, which integrates image quality factors and geometric uncertainty factors. The algorithm quantifies the contribution of image quality based on the image contrast enhancement rate and combines the Euclidean distance from the camera to the visual marker to quantify the influence of geometric configuration. The algorithm dynamically assigns fusion weights to the pose observations of the dual-view field of view. The algorithm performs weighted fusion of depth coordinates and fusion of horizontal coordinates based on rigid body geometric constraints. It can automatically suppress the influence of poor observations (such as occlusion, distant turbidity, and uneven illumination) and effectively smooth binocular parallax noise. Compared with simple average fusion, it has better estimation performance.
[0031] 4. Construct an Error State Kalman Filter (ESKF) that uses the fused visual pose as the observation input and dynamically adjusts the observation noise covariance matrix of the filter based on the dual-field-of-view fusion weights: when the visual observation is clear and reliable at close range, the reliability of the visual observation is increased; when the visual observation degrades, the visual weights are automatically reduced and IMU prediction is relied upon. This breaks the limitations of the traditional static noise assumption of filtering, realizes the adaptive fusion of visual and inertial data, ensures that the system smoothly transitions when the turbidity and illumination of the underwater environment change dynamically, and avoids positioning divergence.
[0032] 5. Experiments have shown that the system can maintain centimeter-level positioning accuracy within a range of 0.5 meters (close-range wall contact) to 20 meters (long-range search), with an average Z-axis error of less than 5 cm, overcoming the defects of DVL near-range blind zone and acoustic positioning limited space multipath effect. Attached Figure Description
[0033] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a block diagram of the overall structure of the dual-field-of-view cooperative perception and positioning system for underwater robots in confined spaces provided in this embodiment of the invention. Figure 2 This is a flowchart of the image enhancement algorithm based on the Jaffe-McGlamery model in an embodiment of the present invention; Figure 3 This is a flowchart illustrating the pose calculation and disambiguation process based on the IPPE algorithm in this embodiment of the invention. Figure 4 This is a schematic diagram of the adaptive quality-weighted fusion algorithm in this embodiment of the invention. Figure 5 This is a flowchart illustrating the overall execution process of the positioning method provided in this embodiment of the invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.
[0035] Example 1: This embodiment is based on a work-class six-degree-of-freedom (6-DoF) remotely operated vehicle (ROV) specifically designed for dam inspection, such as... Figure 1 As shown, the core hardware of the system includes: (1) Heterogeneous dual-field vision subsystem Main camera (gimbal camera, defined as the right camera): A zoom camera with high resolution (e.g., 1080P) and narrow field of view (FOV < 60°) is selected and mounted on the front gimbal of the ROV, with tilt adjustment capability. Its main function is to detect visual markers at long distances (> 5 meters) and provide high-precision remote guidance.
[0036] Auxiliary camera (wide-angle camera, defined as left camera): A fixed-focus camera with an ultra-wide field of view (FOV>120°, such as the DJIAction series modification) is selected and mounted on a fixed bracket at the front of the ROV. Its main function is to provide a wide range of environmental perception at close range (<3 meters) to prevent visual targets from moving out of the field of view when maneuvering at large angles or working close to walls.
[0037] Dual-camera calibration: In a clean water environment, the intrinsic parameters (focal length) of two cameras are calibrated using a checkerboard calibration board. Main point distortion coefficient ) and extrinsic parameter calibration (rotation matrix of the right camera relative to the left camera) Translation vector ).
[0038] (2) Auxiliary navigation sensor Inertial Measurement Unit (IMU): Provides high-frequency (>100Hz) triaxial angular velocity and triaxial acceleration data.
[0039] Depth sensor: Provides high-precision depth (pressure) measurement.
[0040] Doppler velocity measurement (DVL): Provides three-axis velocity when bottom locking conditions are met, but in the close-range operation mode of this invention, DVL is often in the blind zone, and the system needs to automatically switch to vision-driven mode.
[0041] Computing platform: Equipped with an NVIDIA Jetson AGX Xavier or similar high-performance embedded computer, running the Linux operating system and ROS (Robot Operating System) middleware, responsible for all image processing and algorithm calculation.
[0042] A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces, such as... Figure 5 As shown, it includes the following steps: S1. Data Acquisition: Environmental images, including main field-of-view images and auxiliary field-of-view images, are acquired using a heterogeneous dual-field-of-view camera system mounted on an underwater robot and pre-calibrated with external parameters.
[0043] The heterogeneous dual-field-of-view camera system includes a main field-of-view camera and an auxiliary field-of-view camera, which have overlapping fields of view and pre-calibrated extrinsic parameters.
[0044] S2. Image Enhancement: Based on the underwater imaging physical model, the acquired raw images are enhanced, and then the image quality index is calculated. The enhancement process includes red channel compensation and contrast-limited adaptive histogram equalization.
[0045] The degradation of underwater images is mainly due to light absorption and scattering. This embodiment uses a simplified Jaffe-McGlamery model for image restoration, aiming to recover the structural information used for feature extraction.
[0046] In this embodiment, as Figure 2 As shown, it includes the following steps: Step S201: Optical Model Building Total irradiance received by the camera It can be represented as a linear superposition of three components: ; direct components The portion of light reflected from the target surface that reaches the camera directly without scattering contains the main image details. Its attenuation follows the Beer-Lambert law: ,in It is the emissivity of an object's surface. It is the attenuation coefficient. It's depth of field.
[0047] Forward scattering Light is scattered at small angles during propagation, causing image details to become blurred.
[0048] Backscattering Suspended particles in the water reflect ambient light into the camera, creating background noise (Veiling Light) and reducing contrast.
[0049] Step S202: Red Channel Compensation Strategy Water absorbs red light (wavelength ~600-700nm) much more readily than blue and green light. To correct the reduced contrast between the black border and white background of the AprilTag due to the lack of red light, the following compensation is implemented: ; in, For the normalized red and green channel strength, The channel mean. To compensate for the gain coefficient, this step significantly restores the color balance of the image.
[0050] Step S203: Limit contrast adaptive histogram equalization (CLAHE) To address the hazy blurring caused by scattering, the CLAHE algorithm is used to enhance local contrast. (1) Segment the image into Tiles are non-overlapping sub-blocks.
[0051] (2) Calculate the grayscale histogram of each sub-block.
[0052] (3) Set a cutoff limit to cut off the frequency of the histogram that exceeds the limit and distribute it evenly to the other parts of the histogram to limit the excessive amplification of noise.
[0053] (4) Equalize the processed histogram.
[0054] (5) Use bilinear interpolation to eliminate the block effect at the sub-block boundary.
[0055] Step S204: Image quality measurement For the needs of subsequent fusion algorithms, a contrast enhancement rate is defined. As a quantitative indicator of image quality: ; ; Experiments show that this processing improves image contrast by an average of 2.3% to 15.8%, significantly increasing the success rate of feature detection.
[0056] Step S205: Output the enhanced image and quality factor.
[0057] S3. Pose calculation: Using the analytical pose estimation (IPPE) algorithm based on microplanes, visual markers in the enhanced main and secondary field-of-view images are detected and their poses are calculated to obtain two independent camera pose observations.
[0058] After acquiring the enhanced image, the system needs to calculate the 6-DoF pose of the camera relative to the AprilTag, such as... Figure 3 As shown, the specific steps include: Step S301: Feature Detection The AprilTag algorithm library is used to detect tag corners in enhanced images. Sub-pixel coordinates of the four corner points are extracted. .
[0059] Step S302: Homography Matrix Decomposition For planar targets, image points with world coordinates Through homography matrix Related: ; in It is the camera intrinsic parameter matrix.
[0060] Step S303: The IPPE algorithm utilizes SVD decomposition. Matrix. Calculation Using polar decomposition or SVD to Decompose into rotation matrices Translation vector .
[0061] Step S304: Analytical decomposition usually produces two physically possible solutions, which is determined by the properties of perspective projection.
[0062] Step S305: Disambiguation and Optimization of the Solution Calculate the reprojection error of the two hypothetical solutions on the image plane: ; Step S306: Select Minimum and satisfying depth The solution is taken as the final pose.
[0063] Compared to iterative methods, IPPE is non-iterative and does not suffer from the problem of getting trapped in local minima, thus completely solving the "normal inversion" problem.
[0064] S4. Weight Calculation: Based on the image quality index obtained in step S2 and the pose observation values from the camera to the visual marker obtained in step S3, the observation distance of the visual marker is calculated by taking the Euclidean norm of the three spatial coordinate components of the translation vector, and the fusion weights are calculated for the pose observation values of the two cameras respectively.
[0065] The AQWF algorithm is the core of this invention. The principle and logic of the adaptive quality-weighted fusion algorithm are as follows: Figure 4 As shown, this is used to fuse the independent observations from the left and right cameras. and Specifically, it includes the following steps: (1) Coordinate System One Using calibrated external parameters Transform the observations from the right camera to the coordinate system of the left camera (or the robot itself): ; Theoretically, if there is no error, Should equal to However, in reality, due to noise and calibration errors, there is a discrepancy between the two.
[0066] (2) Calculation of weighting factors Define fusion weights ,in .
[0067] Image quality factor This reflects the contribution of the image's inherent sharpness to the solution accuracy, based on the contrast enhancement rate in step S204. Perform the calculation: ; in, It is the adjustment coefficient.
[0068] The physical meaning of this formula is: the greater the contrast enhancement (meaning the original image is more murky but effectively enhanced, or the edges are clearer), the higher the relative confidence of the observation is assigned by the system (Note: the logic here can also be interpreted as the enhanced image having a higher signal-to-noise ratio. The specific positive and negative correlation can be set according to actual tuning. This embodiment adopts positive correlation logic, that is, it is believed that the image with significant enhancement effect contains more effective information for restoration).
[0069] Geometric uncertainty factor This reflects the impact of the observation geometry on accuracy. Based on the covariance propagation theory of pose calculation, depth error... With distance The confidence level is directly proportional to the square of the distance. Therefore, the confidence level (weight) should be inversely proportional to the square of the distance. ; in, This is to prevent division by zero by minima. This means that the closer the distance, the higher the weight.
[0070] S5. Adaptive Fusion: The pose observations of the two cameras are weighted and fused using fusion weights to obtain the visual pose estimate of the robot's body center.
[0071] Weighted fusion includes deep fusion and horizontal fusion: Deep fusion (Z-axis): ; Horizontal fusion (XY axis): Based on rigid body constraints, take the geometric midpoint or weighted midpoint.
[0072] .
[0073] S6. State Filtering: Construct an Error State Kalman Filter (ESKF), take the visual pose estimate obtained in step S5 as the observation input, dynamically adjust the observation noise covariance matrix of the filter according to the fusion weight, and output the robot's final position and attitude by combining the data from the inertial measurement unit (IMU).
[0074] The fused pose is used as a measurement input to the ESKF and is loosely or tightly coupled with the IMU.
[0075] The standard Kalman filter observation update formula is: ; ; in: It is the observation noise covariance matrix. This invention uses the total weights of AQWF... Introduction Update: ; in, It is the nominal noise matrix. It is the scaling factor.
[0076] Scene 1: Clear vision and close distance great extremely small Increase Highly reliable visual observation.
[0077] Scenario 2: Monocular obstruction or cloudiness at a distance Decrease Increase Decrease Relying on IMU prediction.
[0078] This mechanism ensures a smooth transition of the system during visual degradation, avoiding localization divergence.
[0079] Example 2: To verify the effectiveness of this invention, a multi-condition experiment was conducted in a deep water pool. Nine typical locations (Pos1-9) were set up, covering different distances from 2.7 meters to 7.1 meters and different viewing angles. The following are the depth measurement data and fusion algorithm performance analysis for each test point: Pos1: Short-range, high signal-to-noise ratio operating conditions data: Depth of left camera Right camera depth , integration depth Analysis: Under this condition, the two-phase images are clear and the weight distribution is balanced. The AQWF fusion result effectively smooths out the binocular parallax noise.
[0080] Pos2: Heterogeneous mass extreme condition data: Left camera depth Right camera depth , integration depth Analysis: Excessive contrast adjustment on the left camera led to increased noise and a lower quality factor. The algorithm automatically reduces the weight of the left camera, resulting in a bias towards the better-quality right camera, thus avoiding simple averaging errors.
[0081] Pos3: Data under uneven lighting conditions: Depth of left camera Right camera depth , integration depth Analysis: The right camera is in a dark area and requires significant enhancement, resulting in a low signal-to-noise ratio. The left camera (closer and brighter) dominates the fusion (weight > 51.5%), correcting the positive drift of the right camera.
[0082] Pos4: Large-angle side-view data: Left camera depth Right camera depth , integration depth Analysis: The IPPE algorithm exhibits excellent solution stability under slant conditions, and no flipping errors occurred.
[0083] Pos5: Mid-to-Long Distance Consistency Test Data: Left Camera Depth Right camera depth , integration depth Analysis: The binocular data were highly consistent at a distance of 6 meters, verifying the accuracy of the external parameter calibration.
[0084] Pos6: Mid-range smoothing data: Left camera depth Right camera depth , integration depth Analysis: Under typical medium-range operating conditions, the fused output is smooth and stable.
[0085] Pos7: Long-distance side-view data: Left camera depth Right camera depth , integration depth Analysis: At long distances and large angles, geometric uncertainty factors play a dominant role, limiting the jitter in the depth direction.
[0086] Pos8: Close-range consistency test data: Left camera depth Right camera depth , integration depth Analysis: This verifies the accuracy of near-field manipulation.
[0087] Pos9: Balanced output of operating condition data: Left camera depth Right camera depth , integration depth Analysis: The image quality is similar, and the algorithm degrades to an approximate average, which is in line with expectations.
[0088] The experiment leads to the following conclusion: (1) At all test points, the deviation between the position after AQWF fusion and the true value of the laser scanner is in the centimeter range, with an average Z-axis error of <5cm.
[0089] (2) Under extreme heterogeneous quality conditions such as Pos2 and Pos3, AQWF showed better estimation performance than simple averaging and successfully suppressed the influence of poor observations.
[0090] Example 3: This embodiment provides a dual-field-of-view cooperative perception and adaptive weighted localization system for underwater robots in confined spaces, used in the method described in Embodiment 1, including: The perception module includes a wide-angle camera and a narrow-field-of-view gimbal camera with pre-calibrated extrinsic parameters to cover the field of view for short-range operations and remote guidance; as well as an inertial measurement unit and a depth sensor; The enhancement module is configured to perform Jaffe-McGlamery model parameter estimation and CLAHE algorithm to improve underwater image contrast and calculate contrast enhancement rate; The solution module is configured to run the IPPE algorithm to parse the six-DOF pose of the visual markers from the enhanced image and eliminate pose blur. The visual markers are AprilTag QR codes with known geometric dimensions.
[0091] The fusion module is configured to execute the AQWF algorithm and dynamically calculate the binocular fusion weights based on the image contrast enhancement rate and observation distance. The navigation module is configured to run the R-adaptive ESKF algorithm, which fuses visual pose and inertial data to output the robot's state.
[0092] It is also equipped with a lighting unit to provide auxiliary light sources in deep water environments.
[0093] Example 4: This embodiment discloses an electronic device, including: a memory, a processor, and a computer program stored in the memory and running on the memory. When the processor executes the computer program, it implements a dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces as described in Embodiment 1.
[0094] Example 5: A storage medium, which is a computer-readable storage medium, stores a computer program that, when executed by a processor, implements a dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces as described in Embodiment 1.
[0095] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces, characterized in that, Includes the following steps: S1. Data Acquisition: Environmental images, including main field-of-view images and auxiliary field-of-view images, are acquired using a heterogeneous dual-field-of-view camera system mounted on an underwater robot and pre-calibrated with external parameters. S2. Image Enhancement: Based on the underwater imaging physical model, the acquired raw images are enhanced, and then the image quality index is calculated. The enhancement process includes red channel compensation and contrast-limited adaptive histogram equalization. S3, Pose Calculation: Using an analytical pose estimation algorithm based on a small plane, visual markers in the enhanced main and secondary field-of-view images are detected and their poses are calculated to obtain independent pose observations of the visual markers from the two cameras. S4. Weight Calculation: Based on the image quality index obtained in step S2 and the independent pose observations of the two cameras obtained in step S3, the observation distance of the visual marker is calculated by taking the Euclidean norm of the three spatial coordinate components of the translation vector, and the fusion weights are calculated for the pose observations of the two cameras respectively. S5. Adaptive fusion: The pose observations of the two cameras are weighted and fused using fusion weights to obtain the visual pose estimate of the robot's body center. S6. State Filtering: Construct an error state Kalman filter, take the visual pose estimate obtained in step S5 as the observation input, and dynamically adjust the observation noise covariance matrix of the filter according to the fusion weights to output the robot's final position and posture.
2. The dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, In step S2, the simplified Jaffe-McGlamery model is selected as the underwater imaging physical model.
3. The dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, In step S2, the image quality index selected is the contrast enhancement rate, and the formula for calculating the contrast enhancement rate is: ; in, The contrast of the original image. To enhance the contrast of the resulting image, the contrast is defined as the standard deviation of the image gray levels or the mean of the local gradient.
4. The dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, In step S3, the specific process of the analytical pose estimation algorithm based on the microplane is as follows: S31. Extract the pixel coordinates of the four corner points of the visual marker on the image plane; S32. Using the camera intrinsic parameter matrix, calculate the homography matrix between the image plane and the marker physical plane; S33. Perform singular value decomposition on the homography matrix and solve analytically to obtain two sets of possible rotation matrix and translation vector assumptions; S34. Calculate the reprojection error of the two sets of assumptions on the image plane respectively, and select the set with the smallest reprojection error and satisfying the positive physical depth constraint as the optimal pose solution, and use it as the camera pose observation value.
5. The dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, In step S4, the fusion weight is the product of the image quality factor and the geometric uncertainty factor, and the fusion weight calculation formula is as follows: ; in, , L Indicates the left camera. R Indicates the right camera; Indicates camera i The fusion weights; Image quality factor This represents the geometric uncertainty factor.
6. A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 5, characterized in that, The image quality factor reflects the contribution of the image's inherent sharpness to the solution accuracy, based on the contrast enhancement rate in step S4. The calculation formula is as follows: ; in, It is the adjustment coefficient.
7. A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 5, characterized in that, The geometric uncertainty factor reflects the impact of the observation geometry on accuracy, and its calculation formula is as follows: ; in, Let Euclidean distance be the distance from the camera to the visual marker. It is a non-zero small constant.
8. A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, In step S5, the specific operation is as follows: For the depth coordinates of the robot's center Weighted fusion is performed to obtain the weighted fused depth coordinates. The calculation formula is: ; Among them, subscript L , R They represent the left and right cameras, respectively. and These represent the depth observation values from the left and right cameras, respectively; For the horizontal coordinate of the robot's center and By employing geometric midpoint calculation based on rigid body geometric constraints, the optimal estimate of the horizontal coordinate of the robot's body center after binocular data fusion was obtained. : ; in, , These represent the horizontal plane observation coordinate vectors extracted from the pose observations of the left and right cameras after being transformed into a unified robot body center coordinate system.
9. A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, In step S6, the method for dynamically adjusting the observation noise covariance matrix of the filter is as follows: ; in: The nominal observation noise covariance matrix is preset. This is the scaling factor; , These represent the fusion weights for the left and right cameras, respectively. This is the adjusted observation noise covariance matrix.
10. A dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces according to claim 1, characterized in that, It is applicable to underwater environments where GNSS signals are blocked, strong acoustic multipath effects exist, and weak texture features are present, including but not limited to the water-facing surface of hydraulic dams, diversion tunnels, water diversion tunnels, and the surface of deep-sea engineering structures.
11. A dual-field-of-view cooperative perception and adaptive weighted localization system for underwater robots in confined spaces, characterized in that, For performing the method according to any one of claims 1 to 10, comprising: The perception module includes a wide-angle camera and a narrow-field-of-view gimbal camera with pre-calibrated extrinsic parameters to cover the field of view for short-range operations and remote guidance; as well as an inertial measurement unit and a depth sensor; The enhancement module is configured to perform Jaffe-McGlamery model parameter estimation and CLAHE algorithm to improve underwater image contrast and calculate contrast enhancement rate; The solution module is configured to run the IPPE algorithm to resolve the six-DOF pose of visually labeled images from the augmented images and eliminate pose blur. The fusion module is configured to execute the AQWF algorithm and dynamically calculate the binocular fusion weights based on the image contrast enhancement rate and observation distance. The navigation module is configured to run the R-adaptive ESKF algorithm, which fuses visual pose and inertial data to output the robot's state.
12. A dual-field-of-view cooperative perception and adaptive weighted localization system for underwater robots in confined spaces according to claim 11, characterized in that, The visual marker is an AprilTag QR code with known geometric dimensions, and the system is also equipped with an illumination unit to provide auxiliary light sources in deep-water environments.
13. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored on the memory and running thereon, wherein the processor executes the computer program to implement a dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces as described in any one of claims 1 to 10.
14. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements a dual-field-of-view cooperative perception and adaptive weighted localization method for underwater robots in confined spaces as described in any one of claims 1 to 10.