An underwater robot nozzle visual positioning method and device

CN122841501APending Publication Date: 2026-09-29UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611072352.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]然而,真实的水下环境与陆地场景存在本质差异

Benefits of technology

[0023]本申请提供了一种水下机器人管口视觉定位方法。在执行所述方法时,先对原始水下图像进行处理,得到增强图像,再利用暗区分割、Hough圆检测和边缘轮廓填洞三条检测线索对增强图像进行检索,得到候选轮廓,接着将候选轮廓进行椭圆拟合和多维综合评分,得到目标二维椭圆参数,并根据目标二维椭圆参数和相机内参构建空间二次锥面矩阵,然后对空间二次锥面矩阵进行特征值分解,得到分解结果;基于分解结果,通过闭式解析方程求解出管口位姿候选解,最后基于管口位姿候选解,通过因果物理消歧与启发式自适应消歧机制确定定位结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841501A_ABST
    Figure CN122841501A_ABST
Patent Text Reader

Abstract

The application provides a kind of underwater robot pipe orifice visual positioning method and device, it is related to image processing technical field.In executing the method, first, the original underwater image is processed to obtain an enhanced image, then the dark area segmentation, Hough circle detection and edge contour hole filling three detection clues are used to search the enhanced image to obtain the candidate contour, then the candidate contour is fitted with ellipse and multi-dimensional comprehensive score to obtain the target two-dimensional ellipse parameter, and the space quadratic conic matrix is constructed according to the target two-dimensional ellipse parameter and camera internal parameter, then the space quadratic conic matrix is eigenvalue decomposed to obtain the decomposition result;Based on the decomposition result, the pipe orifice pose candidate solution is solved through closed-form analytical equation, and finally, based on the pipe orifice pose candidate solution, the positioning result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanism.Furthermore, the accuracy of underwater robot pipe orifice visual positioning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for visual positioning of underwater robot nozzles. Background Technology

[0002] Remotely Operated Vehicles (ROVs or Autonomous Underwater Vehicles, AUVs) are playing an increasingly important role in marine engineering, nuclear power plant operation and maintenance, and subsea pipeline inspection. Taking nuclear power plants as an example, the surface of the reactor pressure vessel (RPV) has multiple nozzles. During refueling and overhauls, underwater robots carrying ultrasonic probes are needed to approach these nozzles for non-destructive testing. This requires the robot to accurately identify the nozzle locations and estimate its six-degree-of-freedom pose relative to the nozzles, thereby performing precise docking or close-range scanning maneuvers.

[0003] Among various sensing methods, visual sensors have become the preferred solution for close-range localization of underwater robots due to their advantages such as large information capacity, small size, low cost, and no need for contact with the object being measured. Under ideal conditions, by detecting the circular outline of the pipe opening in the image and combining it with camera calibration parameters, the spatial pose of the pipe opening can be easily deduced.

[0004] However, the real underwater environment differs fundamentally from terrestrial scenes. Light undergoes severe absorption and scattering as it travels through water: red light, with its longer wavelength, attenuates most rapidly in water, typically being almost completely absorbed within a few meters, resulting in a significant blue-green tint in underwater images; scattering from suspended particles (plankton, silt, bubbles, etc.) drastically reduces image contrast, making distant targets blurry; and artificial lighting at the work site creates strong specular reflections on the metal pipe openings, forming localized overexposed areas. These factors combined result in underwater pipe opening images of significantly lower quality than those on land, posing a significant challenge to vision-based positioning algorithms.

[0005] In conclusion, improving the accuracy of visual positioning of underwater robot nozzles is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, this application provides a method and apparatus for visual positioning of underwater robot nozzles, aiming to improve the accuracy of visual positioning of underwater robot nozzles.

[0007] In a first aspect, this application provides a visual localization method for underwater robot nozzles, including: The original underwater image is processed to obtain an enhanced image; Candidate contours are obtained by retrieving the enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling. The candidate contour is fitted with an ellipse and scored in multiple dimensions to obtain the parameters of the target two-dimensional ellipse. Construct a spatial quadratic cone matrix based on the target two-dimensional ellipse parameters and camera intrinsic parameters; The spatial quadratic cone matrix is ​​subjected to eigenvalue decomposition to obtain the decomposition result; Based on the decomposition results, candidate solutions for the nozzle pose are obtained by solving closed-form analytical equations. Based on the candidate solutions for the nozzle pose, the localization result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanisms.

[0008] Optionally, the process of processing the original underwater image to obtain an enhanced image includes: The color cast intensity is estimated by performing color cast intensity estimation on the original underwater image; If the color cast intensity exceeds the first threshold, then red channel compensation is performed based on the color cast intensity to obtain the compensated image; Dark channel dehazing is performed on the compensated image, and the transmittance map is estimated. The compensated image is color-corrected using the transmittance map to obtain the corrected image; Based on the color cast intensity, the contrast of the corrected image is enhanced to obtain the enhanced image.

[0009] Optionally, the step of using the transmittance map to perform color correction on the compensated image to obtain the corrected image includes: Using the quantiles of the transmittance map, the compensated image is divided into multiple zones; the zones include near zone, middle zone, and far zone. Perform gray-world white balance and Bradford color adaptation transformation within each partition.

[0010] Optionally, the step of performing ellipse fitting and multidimensional comprehensive scoring on the candidate contour to obtain the target two-dimensional ellipse parameters includes: Based on the candidate contour, the two-dimensional ellipse parameters are obtained using the ellipse fitting method. The parameters of the two-dimensional ellipse are scored using a multi-dimensional comprehensive scoring matrix; The two-dimensional ellipse parameter with the highest score is determined as the target two-dimensional ellipse parameter.

[0011] Optionally, before scoring the two-dimensional ellipse parameters using a multi-dimensional comprehensive scoring matrix, the method further includes: Construct the multidimensional comprehensive scoring matrix; the dimensions of the multidimensional comprehensive scoring matrix include darkness score, inner and outer ring contrast, geometric roundness, mask fill degree, major and minor axis ratio, area ratio and centering score.

[0012] Optionally, after scoring the two-dimensional ellipse parameters using a multi-dimensional comprehensive scoring matrix, the method further includes: When the ellipse axis ratio corresponding to the two-dimensional ellipse parameter is less than the second threshold, the score of the two-dimensional ellipse parameter is reduced.

[0013] Optionally, before retrieving candidate contours from the enhanced image using three detection cues—dark area segmentation, Hough circle detection, and edge contour hole filling—the method further includes: The enhanced image is subjected to specular suppression preprocessing to obtain a preprocessed enhanced image; The enhanced image is retrieved using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling, to obtain candidate contours, including: The candidate contours are obtained by retrieving the preprocessed enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling.

[0014] Secondly, this application provides an underwater robot nozzle visual positioning device, comprising: The processing module is used to process the original underwater image to obtain an enhanced image; The retrieval module is used to retrieve candidate contours from the enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling. The scoring module is used to perform ellipse fitting and multi-dimensional comprehensive scoring on the candidate contour to obtain the target two-dimensional ellipse parameters; The first construction module is used to construct a spatial quadratic cone matrix based on the target two-dimensional ellipse parameters and camera intrinsic parameters; The decomposition module is used to perform eigenvalue decomposition on the spatial quadratic cone matrix to obtain the decomposition result. The solution module is used to solve for candidate solutions of the nozzle pose using closed-form analytical equations based on the decomposition results. The determination module is used to determine the positioning result based on the candidate port pose solution through causal physical disambiguation and heuristic adaptive disambiguation mechanism.

[0015] Optionally, the processing module includes: The estimation submodule is used to estimate the color cast intensity of the original underwater image to obtain the color cast intensity; The supplementary submodule is used to perform red channel compensation based on the color cast intensity if the color cast intensity exceeds a first threshold, so as to obtain a compensated image. The processing submodule is used to perform dark channel dehazing on the compensated image and estimate the transmittance map; The correction submodule is used to perform color correction on the compensated image using the transmittance map to obtain the corrected image; An enhancement submodule is used to enhance the contrast of the corrected image based on the color cast intensity to obtain the enhanced image.

[0016] Optionally, the correction submodule includes: A partitioning unit is used to divide the compensated image into multiple partitions using the quantiles of the transmittance map; the partitions include a near zone, a middle zone, and a far zone. The correction unit is used to perform gray-world white balance and Bradford color adaptation transformation within each partition.

[0017] Optionally, the scoring module includes: The acquisition submodule is used to obtain two-dimensional ellipse parameters based on the candidate contour using an ellipse fitting method; The scoring submodule is used to score the two-dimensional ellipse parameters using a multi-dimensional comprehensive scoring matrix; The determination submodule is used to determine the two-dimensional ellipse parameter with the highest score as the target two-dimensional ellipse parameter.

[0018] Optionally, the device further includes: The second construction module is used to construct the multidimensional comprehensive scoring matrix; the dimensions of the multidimensional comprehensive scoring matrix include darkness score, inner and outer ring contrast, geometric roundness, mask fill degree, major and minor axis ratio, area ratio and centering score.

[0019] Optionally, the device further includes: The deduction module is used to deduct the score of the two-dimensional ellipse parameter when the ellipse axis ratio corresponding to the two-dimensional ellipse parameter is less than a second threshold.

[0020] Optionally, the device further includes: The preprocessing module is used to perform specular suppression preprocessing on the enhanced image to obtain the preprocessed enhanced image; The retrieval module is specifically used for: The candidate contours are obtained by retrieving the preprocessed enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling.

[0021] Thirdly, embodiments of this application provide a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the underwater robot port visual positioning method as described in any embodiment of the first aspect of this application.

[0022] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the underwater robot port visual positioning method as described in any of the embodiments of the first aspect of this application.

[0023] This application provides a visual localization method for underwater robot nozzles. When executing the method, the original underwater image is first processed to obtain an enhanced image. Then, three detection cues—dark area segmentation, Hough circle detection, and edge contour hole filling—are used to retrieve candidate contours from the enhanced image. Next, the candidate contours are subjected to ellipse fitting and multi-dimensional comprehensive scoring to obtain the target two-dimensional ellipse parameters. A spatial quadratic cone matrix is ​​constructed based on the target two-dimensional ellipse parameters and camera intrinsic parameters. Then, eigenvalue decomposition is performed on the spatial quadratic cone matrix to obtain the decomposition result. Based on the decomposition result, candidate nozzle pose solutions are solved using closed-form analytical equations. Finally, based on the candidate nozzle pose solutions, the localization result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanisms.

[0024] In this way, adaptive enhancement of the original underwater image provides a high-quality, color-consistent input image for subsequent detection, fundamentally reducing the risk of false detections due to image degradation. Secondly, three complementary detection cues at the physical level—dark area segmentation, Hough circle detection, and edge contour filling—are used in parallel. These cues capture nozzle features from three perspectives: brightness prior, geometric roundness, and boundary gradient. Even in adverse conditions such as specular highlights, partial occlusion, or edge breaks, these cues can compensate, greatly improving the recall and robustness of the nozzle candidate contours. Subsequently, direct ellipse fitting is performed on all candidate contours to select the optimal two-dimensional ellipse parameters that best match the actual nozzle shape. Based on this, the ellipse parameters are inversely constructed into a spatial quadratic conical matrix using camera intrinsic parameters, and strict geometric constraints are obtained through eigenvalue decomposition. Then, a closed-form analytical equation is used to directly solve for the nozzle pose, ensuring the mathematical accuracy and stability of the pose calculation. Finally, through causal physical disambiguation and heuristic adaptive disambiguation mechanisms, virtual solutions located behind the camera are successively eliminated, the normal vector is ensured to face the camera, and the solution closest to the facing posture is selected based on the inner product of the line-of-sight vectors, thus uniquely determining the positioning result that conforms to physical reality. This improves the accuracy of visual positioning of the underwater robot at the nozzle. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart illustrating a method for visual localization of an underwater robot nozzle, as provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of an underwater robot nozzle visual positioning device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. This application provides a visual positioning method and apparatus for underwater robot nozzles, relating to the field of image processing technology. The above are merely examples and do not limit the application areas of the methods and apparatus provided in this application.

[0028] Remotely Operated Vehicles (ROVs or Autonomous Underwater Vehicles, AUVs) are playing an increasingly important role in marine engineering, nuclear power plant operation and maintenance, and subsea pipeline inspection. Taking nuclear power plants as an example, the surface of the reactor pressure vessel (RPV) has multiple nozzles. During refueling and overhauls, underwater robots carrying ultrasonic probes are needed to approach these nozzles for non-destructive testing. This requires the robot to accurately identify the nozzle locations and estimate its six-degree-of-freedom pose relative to the nozzles, thereby performing precise docking or close-range scanning maneuvers.

[0029] Among various sensing methods, visual sensors have become the preferred solution for close-range localization of underwater robots due to their advantages such as large information capacity, small size, low cost, and no need for contact with the object being measured. Under ideal conditions, by detecting the circular outline of the pipe opening in the image and combining it with camera calibration parameters, the spatial pose of the pipe opening can be easily deduced.

[0030] However, the real underwater environment differs fundamentally from terrestrial scenes. Light undergoes severe absorption and scattering as it travels through water: red light, with its longer wavelength, attenuates most rapidly in water, typically being almost completely absorbed within a few meters, resulting in a significant blue-green tint in underwater images; scattering from suspended particles (plankton, silt, bubbles, etc.) drastically reduces image contrast, making distant targets blurry; and artificial lighting at the work site creates strong specular reflections on the metal pipe openings, forming localized overexposed areas. These factors combined result in underwater pipe opening images of significantly lower quality than those on land, posing a significant challenge to vision-based positioning algorithms.

[0031] The inventors, through research, proposed the technical solution of this application. First, the original underwater image is processed to obtain an enhanced image. Then, three detection cues—dark area segmentation, Hough circle detection, and edge contour hole filling—are used to retrieve candidate contours from the enhanced image. Next, the candidate contours are fitted with ellipses and multidimensionally scored to obtain the target two-dimensional ellipse parameters. Based on the target two-dimensional ellipse parameters and camera intrinsic parameters, a spatial quadratic cone matrix is ​​constructed. Then, eigenvalue decomposition is performed on the spatial quadratic cone matrix to obtain the decomposition results. Based on the decomposition results, candidate solutions for the nozzle pose are solved using closed-form analytical equations. Finally, based on the candidate solutions for the nozzle pose, the positioning result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanisms.

[0032] In this way, adaptive enhancement of the original underwater image provides a high-quality, color-consistent input image for subsequent detection, fundamentally reducing the risk of false detections due to image degradation. Secondly, three complementary detection cues at the physical level—dark area segmentation, Hough circle detection, and edge contour filling—are used in parallel. These cues capture nozzle features from three perspectives: brightness prior, geometric roundness, and boundary gradient. Even in adverse conditions such as specular highlights, partial occlusion, or edge breaks, these cues can compensate, greatly improving the recall and robustness of the nozzle candidate contours. Subsequently, direct ellipse fitting is performed on all candidate contours to select the optimal two-dimensional ellipse parameters that best match the actual nozzle shape. Based on this, the ellipse parameters are inversely constructed into a spatial quadratic conical matrix using camera intrinsic parameters, and strict geometric constraints are obtained through eigenvalue decomposition. Then, a closed-form analytical equation is used to directly solve for the nozzle pose, ensuring the mathematical accuracy and stability of the pose calculation. Finally, through causal physical disambiguation and heuristic adaptive disambiguation mechanisms, virtual solutions located behind the camera are successively eliminated, the normal vector is ensured to face the camera, and the solution closest to the facing posture is selected based on the inner product of the line-of-sight vectors, thus uniquely determining the positioning result that conforms to physical reality. This improves the accuracy of visual positioning of the underwater robot at the nozzle.

[0033] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present application. It should be noted that, for ease of description, only the parts related to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of the present application can be combined with each other.

[0034] See Figure 1 , Figure 1 A flowchart of a method for visual localization of an underwater robot nozzle provided in this application embodiment includes: S101: Process the original underwater image to obtain an enhanced image.

[0035] Original underwater images Separate the blue channel Green Channel Red Channel The pixel values ​​for each channel have been normalized to the [0,1] interval. The joint mean of the blue and green channels is calculated. With the red channel mean The ratio of the two values ​​is used to assess the intensity of color cast. The specific formula is as follows: ; ; in, This represents the arithmetic mean of all pixels in the entire channel; the joint mean of the blue and green channels. Used to characterize the overall intensity of the blue-green component in an image; This represents the average value of the red channel; since red light decays most rapidly in water, its value is typically small; ratio This reflects the degree of deviation of blue-green from red; subtracting 1.0 yields the relative color cast. For the truncation function, The value is restricted to the range Inside, that is, when Time to take ,when Time to take Otherwise take ; This is the color cast intensity assessment value, which takes values ​​ranging from... between, The larger the value, the more severe the color cast in the image.

[0036] when If the image color is close to normal, subsequent white balance is skipped, and only mild contrast enhancement (CLAHE) is performed to protect the original colors from being destroyed. CLAHE is a method that divides the image into blocks, performs histogram equalization, and limits the contrast amplification factor, effectively enhancing local contrast while avoiding excessive noise amplification.

[0037] like If the color cast is severe, then the estimated color cast intensity will be used. Gain compensation is applied to the red channel while the blue and green channels are slightly suppressed to mitigate the sharp attenuation of red light as it passes through water. The calculation formula is as follows: ; ; ; in, , as well as The matrix represents the compensated red, green, and blue channel pixel values; coefficients This represents the gain factor of the red channel. The larger the value, the greater the gain, in order to compensate for the attenuation of red light; coefficient This indicates an inhibitory factor for the green and blue channels, whose intensity is slightly reduced to avoid oversaturation; Ensure that the compensated pixel values ​​remain Within the specified range, prevent overflow.

[0038] Dark Channel Prior Descattering and Distance Estimation: The Dark Channel Prior (DCP) is an image dehazing theory that posits that in local regions of a sharp image, at least one color channel's pixel value approaches zero. This prior can be used to estimate the transmittance scattered by the medium and recover the sharp image. Specifically, dark channel dehazing is applied to the compensated image to estimate the transmittance map reflecting the distance relationships within the scene. transmittance The size directly corresponds to the relative distance between the underwater object and the robot camera. The smaller the value, the stronger the scattering and the farther away the object is.

[0039] Using transmittance diagram 33rd percentile With the 66th percentile As a non-linear physical threshold, the image is dynamically divided into three regions: Nearby area: ; Central District: ; Remote areas: .

[0040] Within each partition, the gray world balance gain is calculated independently, based on the color cast intensity. The image is then scaled. Next, a standard D65 white point is constructed as the target source, and the image is projected from RGB space to the Bradford Visual Response (LMS) space to perform color adaptation balance (CAT). Finally, it is inversely projected back to RGB space. This process achieves distance-wise fine-grained color correction. The Bradford Color Adaptive Transform (Bradford CAT) is a white balance method performed in the LMS cone response space, adapting the color temperature of scene light sources to the standard D65 white point, which is more consistent with human visual perception than a simple gray-world assumption or linear stretching.

[0041] The corrected image undergoes adaptive limited contrast histogram equalization (CLAHE) in the L channel of the LAB color space to improve dynamic range. Finally, the enhanced image and the original image are compared according to color cast intensity. Adaptive weighted linear fusion is performed to obtain the final enhanced image. Contrast-limited adaptive histogram equalization (CLAHE) is a method that divides an image into blocks, performs histogram equalization, and limits the contrast amplification factor. It can effectively enhance local contrast while avoiding excessive noise amplification.

[0042] S102: The enhanced image is retrieved using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling, to obtain candidate contours.

[0043] Statistical Enhancement Image The high percentile (e.g., the 96th percentile) of the corresponding grayscale image is used to identify the specular reflection area, and these extremely bright spots are replaced with locally smoothed background generated by large kernel median filtering (e.g., a 31x31 window).

[0044] Candidate contours are obtained by retrieving them from the enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling. Dark area segmentation: Set multiple percentile thresholds (such as 8%, 12%, 18%, 25%, 35%) and use the global Otsu method to extract multi-level dark area masks in the image in batches, and apply morphological operations to remove burrs.

[0045] Hough Circle Detection: The Hough Circles transform is run on the grayscale image after highlight suppression to directly locate the hidden circular regions in the image.

[0046] Edge contour filling: The edges are extracted using the adaptive double-threshold Canny operator, morphological closing operation is applied to connect the broken edges, and then the flood filling method is used to fill the holes inside the closed area.

[0047] S103: Perform ellipse fitting and multidimensional comprehensive scoring on the candidate contours to obtain the target two-dimensional ellipse parameters.

[0048] For all candidate contours generated by the three sub-clues, the more noise-resistant direct ellipse fitting method (fitEllipseDirect) is uniformly applied to obtain the two-dimensional ellipse geometric parameters (including center cx, cy, major and minor axes a, b, and rotation angle θ). Subsequently, a comprehensive scoring matrix integrating seven dimensions is constructed for scoring. As shown in Table 1, Table 1 is a schematic table of the multi-dimensional comprehensive scoring matrix: Table 1

[0049] Penalty Mechanism: When the ellipse axis ratio b / a < 0.45, the contour is determined to be extremely flat, most likely a false feature caused by striped shadows or welding textures on the pipe wall. In this case, the scoring matrix triggers a penalty, deducting 0.8 points and directly disqualifying the candidate. Finally, the algorithm outputs the set of two-dimensional standard ellipse parameters with the highest score. Entering the solution phase.

[0050] S104: S103: Perform ellipse fitting and multidimensional comprehensive scoring on the candidate contours to obtain the target two-dimensional ellipse parameters.

[0051] Based on the optimal two-dimensional ellipse parameters with the highest score Expand it into a quadratic curve matrix in pixel coordinates. Such that the points on the ellipse satisfy ,matrix It is a symmetric matrix, obtained from the general equation of an ellipse. The coefficients are composed of: ; Combined with the camera's known intrinsic parameter matrix The ellipse at the image plane is extended in reverse into three-dimensional space to construct a quadratic cone matrix with the camera's optical center as its vertex. : ; in, The camera intrinsic parameter matrix is ​​a 3×3 upper triangular matrix containing parameters such as focal length and principal point coordinates, in the form of: ; in, and These are the focal lengths in the x and y directions, respectively. The coordinates of the principal point in the image; It is a spatial quadratic cone matrix with the camera's optical center as the vertex and the image ellipse as the cross section, and it is also a 3×3 symmetric matrix.

[0052] S105: Perform eigenvalue decomposition on the spatial quadratic cone matrix to obtain the decomposition results.

[0053] For the spatial quadratic cone matrix Perform eigenvalue decomposition to obtain its eigenvalues. , , and its corresponding unit eigenvector , , According to the principles of perspective geometry, the eigenvalue signature of a valid real conical surface must be... The algorithm then performs sign scaling and ascending sorting on the feature values ​​accordingly (ensuring...). And both are positive. (Negative).

[0054] S106: Based on the decomposition results, candidate solutions for the nozzle pose are obtained by solving closed-form analytical equations.

[0055] Define intermediate auxiliary geometric variables and : , ; and These are intermediate auxiliary geometric variables, calculated from eigenvalues, used to construct the normal vector and center point. The normal vector of the plane containing the nozzle in three-dimensional space. and the three-dimensional coordinates of the pipe center point Only exists in the form of feature vectors and Within the spanned two-dimensional plane (eigenvectors) (The principal section perpendicular to the cone surface). Therefore, two exact dual mathematical solutions can be directly given through closed-form analytical equations: ; ; in, This is a sign control parameter, taking a value of 1 or -1, representing two different combinations of tilt attitudes; The unit normal vector of the plane containing the nozzle is a 3×1 column vector pointing in the direction of the nozzle opening; The three-dimensional coordinates of the center of the nozzle in the camera coordinate system are represented by a 3×1 column vector. This represents the actual physical radius of the pipe opening.

[0056] S107: Based on the candidate solutions for the nozzle pose, the localization result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanisms.

[0057] Hard physical filtration: Excludes the three-dimensional coordinates of the pipe outlet center Solutions with components less than or equal to 0 are excluded, meaning solutions located in the virtual blind zone behind the camera are excluded, ensuring... .

[0058] Line of sight direction constraint: Ensure the normal vector It must be facing the camera side, that is, it meets the requirement. (The nozzle faces the robot.)

[0059] The core heuristic disambiguation mechanism: After the first two hard filtering steps, two geometrically plausible dual pose solutions are still retained. A "maximum orthogonal pose prior" is introduced, which calculates the inner product of the two candidate normal vectors and their respective camera line-of-sight vectors, prioritizing the solution with the larger absolute value of the inner product and closer to "orthogonal to the center of the nozzle" as the final unique physical solution. This mechanism effectively avoids a 180-degree flip in pose calculation due to minor image noise during robot docking and propulsion.

[0060] If the eigenvalue decomposition results in numerical anomalies due to special boundary conditions, the geometric backtracking operator is automatically activated: It directly utilizes the geometric property that the major axis of the ellipse is not affected by tilt and the line of sight is shortened, estimating the distance and tilt angle using the following formula: ; ; in, This is the estimated depth value of the center of the nozzle along the optical axis of the camera; The focal length of the camera; This is the length of the major semi-axis of the optimal ellipse; This is the length of the minor semi-axis of the optimal ellipse; The angle between the normal vector of the nozzle plane and the camera's optical axis is the tilt angle; this back-off mechanism ensures the algorithm remains uninterrupted and doesn't crash. Finally, based on the determined physical properties, the robot's three-dimensional position is output. Euler angles .

[0061] In the embodiments provided in this application, the original underwater image is first processed to obtain an enhanced image. Then, three detection cues—dark area segmentation, Hough circle detection, and edge contour hole filling—are used to retrieve candidate contours from the enhanced image. Next, the candidate contours are fitted with ellipses and scored using multidimensional comprehensive methods to obtain the target two-dimensional ellipse parameters. A spatial quadratic cone matrix is ​​constructed based on the target two-dimensional ellipse parameters and camera intrinsic parameters. Then, eigenvalue decomposition is performed on the spatial quadratic cone matrix to obtain the decomposition results. Based on the decomposition results, candidate solutions for the nozzle pose are solved using closed-form analytical equations. Finally, based on the candidate solutions for the nozzle pose, the positioning result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanisms.

[0062] In this way, adaptive enhancement of the original underwater image provides a high-quality, color-consistent input image for subsequent detection, fundamentally reducing the risk of false detections due to image degradation. Secondly, three complementary detection cues at the physical level—dark area segmentation, Hough circle detection, and edge contour filling—are used in parallel. These cues capture nozzle features from three perspectives: brightness prior, geometric roundness, and boundary gradient. Even in adverse conditions such as specular highlights, partial occlusion, or edge breaks, these cues can compensate, greatly improving the recall and robustness of the nozzle candidate contours. Subsequently, direct ellipse fitting is performed on all candidate contours to select the optimal two-dimensional ellipse parameters that best match the actual nozzle shape. Based on this, the ellipse parameters are inversely constructed into a spatial quadratic conical matrix using camera intrinsic parameters, and strict geometric constraints are obtained through eigenvalue decomposition. Then, a closed-form analytical equation is used to directly solve for the nozzle pose, ensuring the mathematical accuracy and stability of the pose calculation. Finally, through causal physical disambiguation and heuristic adaptive disambiguation mechanisms, virtual solutions located behind the camera are successively eliminated, the normal vector is ensured to face the camera, and the solution closest to the facing posture is selected based on the inner product of the line-of-sight vectors, thus uniquely determining the positioning result that conforms to physical reality. This improves the accuracy of visual positioning of the underwater robot at the nozzle.

[0063] The above are some specific implementations of the underwater robot port visual positioning method provided in the embodiments of this application. Based on this, this application also provides a corresponding device. The device provided in the embodiments of this application will be described below from the perspective of functional modularity.

[0064] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an underwater robot pipe opening visual positioning device 200 provided in an embodiment of this application. The underwater robot pipe opening visual positioning device 200 includes: Processing module 210 is used to process the original underwater image to obtain an enhanced image; The retrieval module 220 is used to retrieve candidate contours from the enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling. The scoring module 230 is used to perform ellipse fitting and multi-dimensional comprehensive scoring on the candidate contour to obtain the target two-dimensional ellipse parameters; The first construction module 240 is used to construct a spatial quadratic cone matrix based on the target two-dimensional ellipse parameters and camera intrinsic parameters. The decomposition module 250 is used to perform eigenvalue decomposition on the spatial quadratic cone matrix to obtain the decomposition result. Solver module 260 is used to solve for candidate solutions of the nozzle pose by means of closed analytical equations based on the decomposition results. The determination module 270 is used to determine the positioning result based on the candidate port pose solution through causal physical disambiguation and heuristic adaptive disambiguation mechanism.

[0065] Optionally, the processing module 210 includes: The estimation submodule is used to estimate the color cast intensity of the original underwater image to obtain the color cast intensity; The supplementary submodule is used to perform red channel compensation based on the color cast intensity if the color cast intensity exceeds a first threshold, so as to obtain a compensated image. The processing submodule is used to perform dark channel dehazing on the compensated image and estimate the transmittance map; The correction submodule is used to perform color correction on the compensated image using the transmittance map to obtain the corrected image; An enhancement submodule is used to enhance the contrast of the corrected image based on the color cast intensity to obtain the enhanced image.

[0066] Optionally, the correction submodule includes: A partitioning unit is used to divide the compensated image into multiple partitions using the quantiles of the transmittance map; the partitions include a near zone, a middle zone, and a far zone. The correction unit is used to perform gray-world white balance and Bradford color adaptation transformation within each partition.

[0067] Optionally, the scoring module 230 includes: The acquisition submodule is used to obtain two-dimensional ellipse parameters based on the candidate contour using an ellipse fitting method; The scoring submodule is used to score the two-dimensional ellipse parameters using a multi-dimensional comprehensive scoring matrix; The determination submodule is used to determine the two-dimensional ellipse parameter with the highest score as the target two-dimensional ellipse parameter.

[0068] Optionally, the device 200 further includes: The second construction module is used to construct the multidimensional comprehensive scoring matrix; the dimensions of the multidimensional comprehensive scoring matrix include darkness score, inner and outer ring contrast, geometric roundness, mask fill degree, major and minor axis ratio, area ratio and centering score.

[0069] Optionally, the device 200 further includes: The deduction module is used to deduct the score of the two-dimensional ellipse parameter when the ellipse axis ratio corresponding to the two-dimensional ellipse parameter is less than a second threshold.

[0070] Optionally, the device 200 further includes: The preprocessing module is used to perform specular suppression preprocessing on the enhanced image to obtain the preprocessed enhanced image; The retrieval module 220 is specifically used for: The candidate contours are obtained by retrieving the preprocessed enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling.

[0071] This application also provides corresponding devices and computer storage media for implementing the solutions provided in this application.

[0072] like Figure 3 As shown, computer device 01 is represented in the form of a general-purpose computing device. Components of computer device 01 may include, but are not limited to: one or more processors or processor units 03, system memory 08, and buses 04 connecting different system components (including system memory 08 and processor units 03).

[0073] Bus 04 represents one or more of several bus architectures, including memory buses or memory controllers, peripheral buses, graphics acceleration ports, processors, or local buses using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0074] Computer device 01 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 01, including volatile and non-volatile media, removable and non-removable media.

[0075] System memory 08 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 09 and / or cache memory 10. Computer device 01 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 11 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 3 Not shown; usually referred to as a "hard drive"). Although Figure 3 As not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 04 via one or more data media interfaces. System memory 08 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0076] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in system memory 08. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 13 typically perform the functions and / or methods described in the embodiments of the present invention.

[0077] Computer device 01 can also communicate with one or more external devices 02 (e.g., keyboard, pointing device, display 07, etc.), and with one or more devices that enable a user to interact with the computer device 01, and / or with any device that enables the computer device 01 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 06. Furthermore, computer device 01 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 05. Figure 3 As shown, network adapter 05 communicates with other modules of computer device 01 via bus 04. It should be understood that, although... Figure 3 As not shown in the diagram, it can be used in conjunction with computer device 01 with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0078] The processor unit 03 executes various functional applications and data processing by running programs stored in the system memory 08, such as implementing a visual positioning method for underwater robot nozzles provided in the embodiments of this application.

[0079] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0080] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus a general-purpose hardware platform. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a router) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0081] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0082] The above description is merely an exemplary implementation of this application and is not intended to limit the scope of protection of this application.

Claims

1. A method for visual positioning of an underwater robot's nozzle, characterized in that, include: The original underwater image is processed to obtain an enhanced image; Candidate contours are obtained by retrieving the enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling. The candidate contour is fitted with an ellipse and scored in multiple dimensions to obtain the parameters of the target two-dimensional ellipse. Construct a spatial quadratic cone matrix based on the target two-dimensional ellipse parameters and camera intrinsic parameters; The spatial quadratic cone matrix is ​​subjected to eigenvalue decomposition to obtain the decomposition result; Based on the decomposition results, candidate solutions for the nozzle pose are obtained by solving closed-form analytical equations. Based on the candidate solutions for the nozzle pose, the localization result is determined through causal physical disambiguation and heuristic adaptive disambiguation mechanisms.

2. The method according to claim 1, characterized in that, The process of processing the original underwater image to obtain the enhanced image includes: The color cast intensity is estimated by performing color cast intensity estimation on the original underwater image; If the color cast intensity exceeds the first threshold, then red channel compensation is performed based on the color cast intensity to obtain the compensated image; Dark channel dehazing is performed on the compensated image, and the transmittance map is estimated. The compensated image is color-corrected using the transmittance map to obtain the corrected image; Based on the color cast intensity, the contrast of the corrected image is enhanced to obtain the enhanced image.

3. The method according to claim 2, characterized in that, The step of using a transmittance map to perform color correction on the compensated image to obtain a corrected image includes: Using the quantiles of the transmittance map, the compensated image is divided into multiple zones; the zones include near zone, middle zone, and far zone. Perform gray-world white balance and Bradford color adaptation transformation within each partition.

4. The method according to claim 1, characterized in that, The step of performing ellipse fitting and multi-dimensional comprehensive scoring on the candidate contour to obtain the target two-dimensional ellipse parameters includes: Based on the candidate contour, the two-dimensional ellipse parameters are obtained using the ellipse fitting method. The parameters of the two-dimensional ellipse are scored using a multi-dimensional comprehensive scoring matrix; The two-dimensional ellipse parameter with the highest score is determined as the target two-dimensional ellipse parameter.

5. The method according to claim 4, characterized in that, Before using the multidimensional comprehensive scoring matrix to score the two-dimensional ellipse parameters, the method further includes: Construct the multidimensional comprehensive scoring matrix; the dimensions of the multidimensional comprehensive scoring matrix include darkness score, inner and outer ring contrast, geometric roundness, mask fill degree, major and minor axis ratio, area ratio and centering score.

6. The method according to claim 4, characterized in that, After scoring the two-dimensional ellipse parameters using a multi-dimensional comprehensive scoring matrix, the method further includes: When the ellipse axis ratio corresponding to the two-dimensional ellipse parameter is less than the second threshold, the score of the two-dimensional ellipse parameter is reduced.

7. The method according to claim 1, characterized in that, Before retrieving candidate contours from the enhanced image using three detection cues—dark area segmentation, Hough circle detection, and edge contour hole filling—the method further includes: The enhanced image is subjected to specular suppression preprocessing to obtain a preprocessed enhanced image; The enhanced image is retrieved using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling, to obtain candidate contours, including: The candidate contours are obtained by retrieving the preprocessed enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling.

8. A visual positioning device for underwater robot nozzles, characterized in that, include: The processing module is used to process the original underwater image to obtain an enhanced image; The retrieval module is used to retrieve candidate contours from the enhanced image using three detection cues: dark area segmentation, Hough circle detection, and edge contour hole filling. The scoring module is used to perform ellipse fitting and multi-dimensional comprehensive scoring on the candidate contour to obtain the target two-dimensional ellipse parameters; The first construction module is used to construct a spatial quadratic cone matrix based on the target two-dimensional ellipse parameters and camera intrinsic parameters; The decomposition module is used to perform eigenvalue decomposition on the spatial quadratic cone matrix to obtain the decomposition result. The solution module is used to solve for candidate solutions of the nozzle pose using closed-form analytical equations based on the decomposition results. The determination module is used to determine the positioning result based on the candidate port pose solution through causal physical disambiguation and heuristic adaptive disambiguation mechanism.

9. A computer device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the underwater robot port visual positioning method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the underwater robot port visual positioning method as described in any one of claims 1-7.