Underwater robot monocular vision positioning method for complex environment

By integrating optical flow networks and deep networks on the underwater robots, distortion correction and dense correspondence extraction are performed, the robustness and accuracy of monocular visual positioning in the underwater environment are solved, and low-cost and high-precision underwater robot positioning is achieved.

CN120070275AInactive Publication Date: 2025-05-30DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510527265.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Underwater robots are difficult to position in complex environments, the existing technology is costly, cumbersome to install, and is affected by water flow and tides. Monocular visual positioning has problems such as sparse image features, poor descriptor differences and dynamic object interference in underwater environments.

Method used

A monocular visual positioning method for underwater robots is adopted for complex environments. Images are acquired through a monocular camera equipped by underwater robots, distortion correction is performed, and dense correspondence between images is extracted using an optical flow network, and scene depth map is predicted in combination with a depth network to realize underwater monocular visual positioning.

Benefits of technology

It effectively solves the robustness of monocular visual positioning in underwater environments, improves positioning accuracy, reduces costs, and is suitable for autonomous navigation and autonomous operations in complex marine environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070275A_ABST
    Figure CN120070275A_ABST
Patent Text Reader

Abstract

The invention provides an underwater robot monocular vision positioning method for a complex environment, and belongs to the technical field of underwater image processing. Comprising the following steps: acquiring an underwater image through a monocular camera carried by an underwater robot, and performing distortion correction on the underwater image to obtain an underwater image after distortion correction; extracting a dense corresponding relation between two adjacent frames of images through an optical flow network, and filtering abnormal optical flow values based on a bidirectional consistency screening strategy; a scene depth map is predicted by using the deep network, and the optical flow network and the deep network are jointly trained by minimizing the target function mean value of each pixel in the whole image, so that underwater monocular vision positioning is realized. According to the underwater robot monocular vision autonomous positioning method, dependence on a large scale of underwater data sets containing labels is effectively avoided through a self-supervised learning mode, and the high-precision underwater robot monocular vision autonomous positioning method is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater image processing, and more particularly, to a monocular vision positioning method for an underwater robot facing a complex environment. Background Art

[0002] In recent years, the demand for the wide application of underwater robots has been increasing day by day, such as real-time seabed mapping, pipeline detection, mine clearance, and daily monitoring of seafood. Especially in the process of deep-sea and far-sea resource exploration and scientific research, a large number of tasks need to rely on the assistance of underwater robots to complete. For complex ocean scenes, the positioning function of underwater robots is the basis for realizing the above applications. The positioning ability can significantly improve the autonomy of underwater robots when performing tasks in unknown scenes. Many underwater positioning methods rely on high-precision underwater sensors, such as sonar and Doppler velocimeters. These sensors can provide good positioning results, but the high cost limits their wide application. The multi-node positioning method based on an underwater acoustic sensor network has the advantages of a wide detection range and strong adaptability. However, this method requires cumbersome and time-consuming installation before use and is not suitable for application in unknown ocean scenes. In addition, sensor nodes are usually affected by water flow or tides and generate passive movement, resulting in deviation of the reference point position, thereby affecting the positioning accuracy. To achieve the autonomous navigation and autonomous operation of underwater robots in complex environments, low-cost positioning technology has become an urgent need.

[0003] Monocular vision sensors have broad application prospects in the field of positioning. However, due to the complex underwater environment, many problems still need to be solved urgently. For example, there are problems such as scarce image features, poor descriptor differences, and interference from dynamic objects in the underwater environment, which pose challenges to the robustness of visual positioning. The camera carried by an underwater robot is usually placed in a waterproof compartment. When the underwater camera images, light needs to pass through multiple media such as water bodies, waterproof compartment glass, and lenses, causing the propagation direction of light to change during the imaging process, resulting in pillow distortion or barrel distortion. In addition, during the installation of the underwater camera, the lens and the photosensitive plane are not strictly parallel, resulting in tangential distortion. The above problems will cause the underwater monocular vision positioning system to be unable to obtain an accurate scene structure, thereby affecting the accuracy of the underwater vision positioning method. Summary of the Invention

[0004] In view of the technical problems mentioned in the above background art, a monocular vision positioning method for an underwater robot facing a complex environment is provided.

[0005] The technical means adopted by the present invention are as follows: A monocular vision positioning method for an underwater robot facing a complex environment, characterized by comprising the following steps: Step 1: Obtain underwater images through the monocular camera carried by the underwater robot, perform distortion correction on the underwater images, and obtain the distortion-corrected underwater images; Step 2: Define two adjacent frames of the distortion-corrected underwater images as ( , ). Extract the dense correspondence between the two adjacent frames of images through an optical flow network, and filter out abnormal optical flow values based on a bidirectional consistency screening strategy; in the optical flow network, first extract pyramid features from different levels of the image through convolution; then, infer the optical flow field according to the high-level features of image and image and , and gradually refine the estimated value of the optical flow field through iteration to achieve the estimation of the camera's pose parameters; Step 3: Input two adjacent frames of images ( , ) into a depth network, use the depth network to predict the scene depth map, and jointly train the optical flow network and the depth network by minimizing the mean of the objective function of each pixel in the image to achieve underwater monocular vision positioning. Furthermore, the distortion correction of the underwater images in Step 1 includes the following steps: Step 11: Project the 3D points in space into the camera coordinate system, and define the projected coordinates as ; since there is a refractive index difference between the air medium and the water when light propagates in the underwater scene, the path will change non-linearly, so three parameters , and are used to correct the radial distortion of the camera lens. At the same time, parameters and are used to correct the tangential distortion of the camera lens to obtain the distortion-corrected coordinates; Step 12: Define the upper left corner of the image as the origin of the pixel coordinate system. The pixel coordinates are scaled by a scale on the axis, and scaled by a scale on the axis, and translated by from the origin; there is a scaling and a translation of the origin between the pixel coordinate system and the imaging plane. The calculation formula for projecting a three-dimensional space point onto the pixel plane is: ; where represents the pixel coordinates, represents the focal length Furthermore, the distortion-corrected coordinates are: ; Among them, represents the coordinates of the projection plane points; represents the coordinates of the distortion points; represents any point on the plane the distance between and the origin of the coordinate system; , and both represent the radial distortion coefficients, and both represent the tangential distortion coefficients.

[0006] Furthermore, step 2 includes the following steps: Step 21: Let the high-level feature representations of image and image be and , for any sub-pixel displacement : ; Among them, represents the spatial resolution factor, represents the source coordinates of the defined sampling points in the high-level feature , represents 's four adjacent pixels; Step 22: Establish the point correspondence between the corrected images and ; During the process of descriptor matching, establish the point correspondence between images and by calculating the correlation of the high-level feature vectors in the single pyramid feature and ; Aggregate the matching cost of each point in the high-level feature map together to construct the total cost value ; Step 23: Calculate the optical flow field of the current pyramid level: ; Among them, represents the descriptor matching unit; represents the screening of the matching cost value; represents the pixel points at the spatial resolution of the upper layer pyramid feature; represents upsampling; represents the amplitude; Since the cost volume function in the descriptor matching unit is aggregated by measuring the correlation pixel by pixel, in order to prevent false optical flow from being amplified by upsampling and passed to the next pyramid level, the pixel-level optical flow field is refined to sub-pixel accuracy, that is, through optical flow estimation to project as ; the sub-pixel refinement unit calculates the residual flow to minimize the feature space distance between and to obtain a more accurate optical flow field , and the calculation formula is: ; Step 24: Calculate the photometric error to obtain the difference between the pixels of image and the image after projective transformation : ; ; where SSIM represents structural similarity; τ represents a hyperparameter, τ = 0 . 85 to balance the image similarity error and the color intensity error; ω represents the projection function; represents reprojection of the pixel coordinates of the th image to the th image; represents the camera internal parameters; represents the predicted depth map of the th image; represents the relative pose between two images; then the reprojection function x of pixel is: ; Step 25: Calculate the average value of the loss function for each pixel, define the forward optical flow as , and the backward optical flow as : The optical flow network is trained by minimizing the average value of the loss function for each pixel in the entire image, and the loss function is expressed as: ; where , represents the edge-aware depth smoothing error; represents the bidirectional prediction consistency of the forward optical flow and the backward optical flow; Indicates the correspondence between the th image established through the optical flow field and the th image; Step 26: The monocular vision positioning system performs 2D-2D matching through the epipolar geometry method to perform the motion estimation of the underwater robot.

[0007] Furthermore, the step 26 includes the following steps: Step 261: For a certain pixel point in the figure , predict its corresponding position in the figure according to the forward optical flow, and then predict its pixel position in the figure of the predicted point according to the backward optical flow ; after calculating the consistency of the forward optical flow and the backward optical flow , select the one with the smallest optical flow inconsistency to form the best matches; that is, the projection process at the pixel point is expressed as:

[0008] ; wherein, represents the coordinate index value of the pixel; indicates the correspondence between the th image and the th image established through the optical flow field ( ); Step 262: After selecting the 2D-2D correspondence ( ), use the epipolar geometry to solve the fundamental matrix or the essential matrix ; Step 263: Obtain the motion estimation of the underwater robot by decomposing or : :

[0009] wherein, , , represents the skew-symmetric matrix, represents the camera internal parameter matrix.

[0010] Furthermore, by calculating the correlation of the high-level feature vectors in the single pyramid feature and to establish the image and The calculation formula for the point correspondence relationship between is: where represents the point in and the point in the matching cost between, represents the displacement vector of represents the length of the feature vector.

[0011] Furthermore, step 3 includes the following steps: Step 31: Jointly train the network by minimizing the mean of the objective function for each pixel in the entire image, then the regularization form of the edge-aware depth smoothing term is expressed as: ; where and represent the gradients in the horizontal and vertical directions respectively; the depth consistency error is calculated as: ; Take the photometric loss function, smoothness loss function, and depth consistency loss function as the final loss function , and jointly train the network by minimizing the mean of the objective function for each pixel in the entire image. The loss function is expressed as: ; where represents the photometric error, represents the depth smoothing error, represents the depth consistency error, and represent the loss weights; Step 32: Use the perspective n point method to solve the 3D to 2D point pair motion, and estimate the position information of the underwater robot by minimizing the reprojection error; Estimate the position information of the underwater robot by minimizing the reprojection error. The solution formula is: ; where represents the coordinate index value of the pixel.

[0012] Compared with the prior art, the present invention has the following advantages: (1) The present invention integrates the bidirectional optical flow consistency and the epipolar geometry constraint strategy, which can effectively solve the positioning problem of underwater robots in weak-texture image scenes.

[0013] (2) The present invention proposes an iterative depth consistency method. By aligning the geometric triangulation depth with the scale-consistent depth, it can effectively solve the depth ambiguity problem in monocular vision positioning in the underwater environment.

[0014] In summary, the present invention effectively avoids the dependence on large-scale underwater datasets with labels through self-supervised learning, uses optical flow and depth networks to effectively interpret the depth information of underwater scenes, and realizes a high-precision monocular vision autonomous positioning method for underwater robots. This is of great significance for improving the autonomy of underwater mobile devices and enhancing the intelligence level of underwater robots in complex marine environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 Trajectory comparison of the monocular vision positioning method for the underwater robot of the present invention Figure 1 .

[0017] Figure 2 Trajectory comparison of the monocular vision positioning method for the underwater robot of the present invention Figure 2 .

[0018] Figure 3 Cumulative distribution function graph of the present invention.

[0019] Figure 4 Schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0022] As Figures 1 - 3 shown, the present invention first corrects the image distortion caused by the influence of the underwater multi-medium environment. Then, the corrected underwater image is sent to the optical flow network and the single-view depth prediction network simultaneously, and the dense correspondence relationship between images is predicted through the optical flow network. Due to the interference of underwater dynamic objects or occlusions, many outliers may appear. To solve this problem, the present invention proposes a reliable optical flow selection strategy based on the bidirectional consistency screening strategy. Then, the epipolar geometry method is used for 2D-2D matching to perform motion estimation. Finally, the iterative depth consistency method proposed by the present invention can solve the depth ambiguity problem by aligning the geometric triangulation depth with the scale-consistent depth.

[0023] Specifically, it includes the following steps: Step 1: Obtain an underwater image through a monocular camera carried by an underwater robot, and correct the distortion of the underwater image to obtain a distortion-corrected underwater image; Different from the ground working environment, underwater light needs to pass through multiple different media, including water, camera waterproof housing, lens, camera sensor, etc., resulting in the change of the light propagation path due to refraction. A calibration method considering the refraction of multiple media is required to correct the radial distortion to ensure the geometric accuracy of the captured image. In addition, due to the fact that the lens cannot be strictly parallel to the photosensitive plane, tangential distortion occurs. Therefore, distortion correction is crucial, otherwise it will seriously affect the accuracy of visual measurement.

[0024] Step 1 includes the following steps: Step 11: Project the 3D points in space into the camera coordinate system, and define the projected coordinates as ; Since there is a refractive index difference between the air medium when light propagates in the underwater scene, the path will change non-linearly. For the radial distortion of the camera lens, by using three parameters , and for correction; for the tangential distortion of the camera lens, use two parameters and perform calibration; Step 12: Define the upper left corner of the image as the origin of the pixel coordinate system. The pixel coordinates are scaled by a scale factor on the axis, and are scaled by a scale factor on the axis and translated by from the origin; There is a difference of a scaling and a translation of the origin between the pixel coordinate system and the imaging plane. The calculation formula for projecting a three-dimensional space point onto the pixel plane is as follows: ; where, represents the pixel coordinates, represents the focal length.

[0025] Furthermore, the coordinates before and after the distortion correction are: ; where, represents the coordinates of the point on the projection plane; represents the coordinates of the distorted point; represents the distance between any point on the plane and the origin of the coordinate system; , and both represent the radial distortion coefficients, and both represent the tangential distortion coefficients.

[0026] Furthermore, Step 2: Define two adjacent frames of the underwater image after distortion correction as ( , ), extract the dense correspondence between the two adjacent frames of the image through an optical flow network, and filter out abnormal optical flow values based on a bidirectional consistency screening strategy; In the optical flow network, first extract pyramid features from different levels of the image through convolution; Then, infer the optical flow field based on the high-level features of image and the high-level features and of image

[0027] The said Step 2 includes the following steps: Step 21: Let the high-level feature representations of image and image be and For any sub-pixel displacement : ; Wherein, represents the spatial resolution factor, represents the source coordinates of the defined sampling points in the high-level feature ; represents four adjacent pixels of; Step 22: Establish the point correspondence between the corrected image and ; During the process of descriptor matching, the point correspondence between the images and is established by calculating the correlation of the high-level feature vectors in the single pyramid feature and ; By aggregating the matching costs of each point in the high-level feature map, the total cost value is constructed; Step 23: Calculate the optical flow field of the current pyramid level : ; Wherein, represents the descriptor matching unit; represents the filtering of the matching cost value; represents the pixel points at the spatial resolution of the upper-layer pyramid feature; represents upsampling; represents the amplitude; Since the cost capacity function in the descriptor matching unit is aggregated by measuring the correlation pixel by pixel, therefore, in order to prevent the wrong optical flow from being amplified by upsampling and passed to the next pyramid level, the pixel-level optical flow field is refined to sub-pixel accuracy, that is, through the optical flow estimation projects to ; The sub-pixel refinement unit calculates the residual flow to minimize the feature space distance between and , thereby obtaining a more accurate optical flow field , and the calculation formula is: ; Step 24: Calculate the photometric error to obtain the difference between the pixels of the image and the image after the projection transformation : ; ; Among them, SSIM represents structural similarity; τ represents a hyperparameter, τ = 0 . 85 to balance the image similarity error and the color intensity error; ω represents a projection function; represents the th image's pixel coordinates reprojected to the th image; represents the camera internal parameters; represents the predicted depth map of the th image; represents the relative pose between two images; Then the re - projection function of pixel x is: For: ; Step 25: Calculate the average value of the loss function for each pixel, and define the forward optical flow as , and the backward optical flow as : Train the optical flow network by minimizing the average value of the loss function for each pixel in the entire image. The loss function is expressed as: ; Among them, , represents the edge - aware depth smoothing error; represents the bidirectional prediction consistency of the forward optical flow and the backward optical flow; represents the correspondence relationship established between the th image and the th image through the optical flow field; Step 26: The monocular vision positioning system performs 2D - 2D matching through the epipolar geometry method to perform the underwater robot motion estimation.

[0028] Furthermore, the said Step 26 includes the following steps: Step 261: For a certain pixel point in Figure , predict its corresponding position in Figure according to the forward optical flow, and then predict its pixel position in Figure of this predicted point according to the backward optical flow; After calculating the consistency of the forward optical flow and the backward optical flow, select the one with the smallest optical flow inconsistency to form the best one match; that is, the projection process at the pixel point is expressed as:

[0029] ; wherein, represents the coordinate index value of the pixel; represents the correspondence between the th image and the th image established through the optical flow field ( ); Step 262: After selecting the 2D-2D correspondence ( ), use the epipolar geometry to solve the fundamental matrix or the essential matrix ; Step 263: Through the decomposition of or obtain the motion estimation of the underwater robot :

[0030] wherein, , , represents an anti-symmetric matrix, represents the camera intrinsic matrix.

[0031] When calculating the reprojection error of multi-source images, the strategy of some self-supervised depth estimation methods is to evenly distribute the reprojection error among each source image. This may lead to the problem that some pixels are visible in the target image but not in some source images. When the network predicts the correct depth of this pixel, the corresponding color in the occluded part of the source image may not match the target, resulting in a relatively high photometric error. The main reason for this phenomenon is the influence of underwater occluding objects or pixels outside the imaging field of view when the underwater robot is moving. To solve the above two problems, instead of calculating the average value of the photometric errors of all source images, taking the minimum value can significantly reduce the artifacts at the image boundary, improve the clarity of the occlusion boundary, and improve the accuracy.

[0032] Furthermore, as a preferred implementation manner, in the present application, Step 3: Input ( , ) into the depth network, use the depth network to predict the scene depth map, and jointly train the optical flow network and the depth network in Step 2 by minimizing the mean value of the objective function of each pixel in the entire image to achieve underwater monocular vision positioning. The steps include: Step 31: Jointly train the network by minimizing the mean value of the objective function for each pixel in the entire image. Then, the regularization form of the edge-aware depth smoothing term is expressed as: ; where and represent the gradients in the horizontal and vertical directions respectively; The depth consistency error is calculated as: ; Take the photometric loss function, smoothness loss function, and depth consistency loss function together as the final loss function , and jointly train the network by minimizing the mean value of the objective function for each pixel in the entire image. The loss function is expressed as: ; where represents the photometric error, represents the depth smoothing error, represents the depth consistency error, and represent the loss weights; Step 32: Use the perspective-n-point method to solve the 3D-to-2D point pair motion, and estimate the position information of the underwater robot by minimizing the reprojection error; Estimate the position information of the underwater robot by minimizing the reprojection error, and the solution formula is: ; where represents the coordinate index value of the pixel.

[0033] Furthermore, the root mean squared error (RMSE) and mean absolute error (MAE) are used to quantitatively analyze the performance of the monocular vision positioning method for the underwater robot proposed in the present invention. After calculation, the RMSE result of this method is 0.44 m, and the MAE result is 0.51 m. In addition, the method proposed in this application is compared with other methods, and RMSE is used as the evaluation index. The detailed comparison results are shown in Table 1.

[0034] Table 1 Comparison table of the accuracy of the monocular vision positioning method for the underwater robot

[0035] Reference document: [1] Zhao C, Dong H, Wang J, et al. Dual-type marker fusion-based underwater visual localization for autonomous docking[J]. IEEE Transactions on Instrumentation and Measurement, 2023, 73: 1-11. [2] Jung J, Lee Y, Kim D, et al. AUV SLAM using forward / downward looking cameras and artificial landmarks[C]. 2017 IEEE Underwater Technology(UT). IEEE, 2017: 1-3. [3] Burguera Burguera A, Bonin-Font F. A trajectory-based approach to multi-session underwater visual SLAM using global image signatures[J]. Journal of Marine Science and Engineering, 2019, 7(8): 278.

[0036] As can be seen from Table 1, the accuracy of monocular vision localization of the underwater robot using the method of the present invention has been significantly improved.

[0037] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0038] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0039] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0040] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0041] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0042] If the above-mentioned integrated units are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0043] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of each embodiment of the present invention.

Claims

1. A monocular vision positioning method for underwater robots in complex environments, characterized in that: The following steps are involved: Step 1: Acquire an underwater image through a monocular camera carried by an underwater robot, and perform distortion correction on the underwater image to obtain a distortion-corrected underwater image; Step 2: The images of two adjacent frames in the distortion-corrected underwater image are defined as ( , ), extracting the dense correspondence between the two adjacent frames of images through an optical flow network, and filtering abnormal optical flow values ​​based on a bidirectional consistency screening strategy; in the optical flow network, firstly extracting pyramid features from different levels of the image through convolution; Then, according to the image and images Advanced features of and Infer the optical flow field and gradually refine the estimated value of the optical flow field through iteration to estimate the camera's pose parameters; Step 3: Combine the images of two adjacent frames ( , ) is input into the deep network, and the deep network is used to predict the scene depth map. By minimizing the mean of the objective function of each pixel in the image, the optical flow network and the deep network are jointly trained to achieve underwater monocular visual positioning.

2. The monocular vision positioning method for underwater robots in complex environments according to claim 1, characterized in that: The step 1 of performing distortion correction on the underwater image comprises the following steps: Step 11: Project the 3D point in space to the camera coordinate system and define the projected coordinates as ; Since there is a difference in refractive index between light propagating in underwater scenes and air media, the path will change nonlinearly, so three parameters are used , and Correct the radial distortion of the camera lens and use the parameters and Correct the tangential distortion of the camera lens to obtain the coordinates after distortion correction; Step 12: Define the upper left corner of the image as the origin of the pixel coordinate system. The pixel coordinates are On axis by scale Zoom in The axis will be scaled Zoom and translate from the origin ; The difference between the pixel coordinate system and the imaging plane is a scaling and a translation of the origin. The calculation formula for projecting a three-dimensional space point onto the pixel plane is: ; in, represents pixel coordinates, Indicates focal length.

3. The monocular vision positioning method for underwater robots in complex environments according to claim 2 is characterized in that: The coordinates after the distortion correction are: ; in, Represents the coordinates of the projection plane point; represents the coordinates of the distortion point; Represents any point on the plane The distance from the origin of the coordinate system; , and Both represent radial distortion coefficients, and Both represent the tangential distortion coefficient.

4. The monocular vision positioning method for underwater robots in complex environments according to claim 1, characterized in that: The step 2 comprises the following steps: Step 21: Make the image and images The high-level features of and , for any sub-pixel displacement : ; in, represents the spatial resolution factor, Indicates that the defined sampling point is in the high-level feature The source coordinates in express The four adjacent pixels of Step 22: Create the Corrected Image and Point correspondence between them; in the process of descriptor matching, by calculating a single pyramid feature and Correlation of mid- and high-level feature vectors to build images and The point correspondence between them; by matching the cost of each point in the high-level feature map Aggregate together to build the total cost value ; Step 23: Calculate the optical flow field of the current pyramid level : ; in, represents the descriptor matching unit; Indicates the filter matching cost; Represents the pixel point at the spatial resolution of the previous pyramid feature; represents upsampling; represents amplitude; Since the cost capacity function in the descriptor matching unit is aggregated by measuring the correlation pixel by pixel, in order to prevent the wrong optical flow from being amplified by upsampling and passing to the next pyramid level, the pixel-level optical flow field Refine to sub-pixel accuracy, i.e. through optical flow estimation Will Projection ; Sub-pixel refinement unit By calculating the residual flow ,make and The feature space distance between them is minimized, thus obtaining a more accurate optical flow field , the calculation formula is: ; Step 24: Calculate the photometric error and obtain the image and the image After projection transformation The difference in pixels: ; ; Among them, SSIM represents structural similarity; τ represents the hyperparameter, τ = 0 . 85 to balance the image similarity error and color intensity error; ω represents the projection function; Indicates that the The pixel coordinates of the first image are reprojected to the images; Represents the camera internal parameters; The predicted The depth map of the image; Represents the relative pose between two images; Then the pixel x The reprojection function for: ; Step 25: Calculate the average value of the loss function for each pixel and define the forward optical flow as , the backward optical flow is : The optical flow network is trained by minimizing the average of the loss function for each pixel in the entire image. The loss function is expressed as: ; in, , represents the edge-aware depth smoothing error; Indicates the bidirectional prediction consistency of forward optical flow and backward optical flow; Represents the first image and The correspondence between the images; Step 26: The monocular vision positioning system performs 2D-2D matching through epipolar geometry method to perform underwater robot motion estimation.

5. The monocular vision positioning method for underwater robots in complex environments according to claim 4, characterized in that: The step 26 comprises the following steps: Step 261: A pixel in According to the forward optical flow prediction, The corresponding position in , that is, a certain pixel The predicted point is then predicted according to the reverse optical flow. The pixel position ; Calculate the consistency of forward and backward optical flow After that, select the one with the smallest optical flow inconsistency To form the best matches; that is, at the pixel The projection process at is expressed as: ; in, Represents the coordinate index value of the pixel; Represents the first image and The correspondence between the images ( ); Step 262: Select 2D-2D correspondence ( ) and then use epipolar geometry to solve the fundamental matrix or the essential matrix ; Step 263: By Decomposition or Get the motion estimate of the underwater robot : in, , , represents an antisymmetric matrix, Represents the camera intrinsic parameter matrix.

6. The monocular vision positioning method for underwater robots in complex environments according to claim 4, characterized in that: By calculating a single pyramid feature and Correlation of mid- and high-level feature vectors to build images and The calculation formula for the point correspondence between is: ; in, express Points in and Points in The matching cost between express The displacement vector, Represents the length of the feature vector.

7. The monocular vision positioning method for underwater robots in complex environments according to claim 1, characterized in that: The step 3 comprises the following steps: Step 31: Jointly train the network by minimizing the mean of the objective function for each pixel in the entire image, then the edge-aware depth smoothing term The regularized form of is expressed as: ; in, and Represent the gradients in the horizontal and vertical directions respectively; Depth consistency error The calculation formula is: ; The photometric loss function, smoothness loss function and depth consistency loss function are used together as the final loss function , and jointly train the network by minimizing the mean of the objective function for each pixel in the entire image. The loss function is expressed as: ; in, represents the photometric error, represents the depth smoothing error, represents the depth consistency error, and represents the loss weight; Step 32: Using perspective n The point method solves the motion of 3D to 2D point pairs and estimates the position information of the underwater robot by minimizing the reprojection error. The position information of the underwater robot is estimated by minimizing the reprojection error. The solution formula is: ; in, Represents the pixel coordinate index value.

Citation Information

Patent Citations

  • Visual simultaneous localization and mapping method based on depth convolution auto-encoder

    CN111325794A

  • Foresight scene depth estimation method based on self-supervised learning

    CN113313732A