Light vision positioning method, program, equipment and storage medium for AUV (Autonomous Underwater Vehicle) recovery docking

By adaptively adjusting the exposure time and using a geometric threading algorithm to screen the optical beacon outline, combined with the light source feature matching and sliding window optimization of the binocular camera, the optical positioning problem of the AUV in complex ocean environments was solved, and high-precision underwater recovery docking was achieved.

CN120593713APending Publication Date: 2025-09-05HARBIN ENG UNIV

Patent Information

Application Number
CN202510762511.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In complex marine environments, existing technologies make it difficult to achieve efficient and accurate underwater optical positioning and guidance of AUVs. Especially under the influence of background light and seawater scattering, the recognition and positioning accuracy of optical sensors is insufficient, affecting the recovery and docking effect of AUVs.

Method used

Adaptive exposure time adjustment, outermost boundary tracking method and geometric threading algorithm are used to screen the outline of the optical beacon. Combined with the light source feature matching and relative pose estimation of the binocular camera, the trajectory accuracy is improved through sliding window optimization to ensure that the AUV can accurately locate the recovery dock in complex environments.

Benefits of technology

The optical positioning accuracy of AUV in complex marine environments is improved, ensuring that AUV can accurately identify the position of the recovery dock and achieve efficient underwater docking and recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120593713A_ABST
    Figure CN120593713A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of underwater docking and recovery of autonomous underwater vehicles, and particularly relates to a light vision positioning method, program and device for AUV recovery docking and a storage medium. In the image acquisition process, the exposure time is adaptively adjusted to ensure that imaging of the underwater optical beacon in the binocular image is not influenced by ambient light; performing target contour extraction on the binarized binocular image by adopting an outermost layer boundary tracking method, and screening contours of four groups of underwater optical beacons by adopting a geometric thread algorithm; a target tracking method is utilized to realize matching of each light source characteristic between single-frame binocular and global image sequence tracking, an analysis method and an iteration method are combined, a relative pose of a docking base station is solved by utilizing a real spatial position of the underwater optical beacon, and a sliding window is utilized to realize a target tracking function. All AUV poses and light source features in a period of time in history are subjected to nonlinear optimization, and the track precision is improved. According to the method, the AUV can be assisted to accurately identify the pose of the recovery dock station under the interference of the underwater environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater docking and recovery of autonomous underwater vehicles, and in particular relates to an optical vision positioning method, program, equipment and storage medium for AUV recovery and docking. Background Art

[0002] Autonomous underwater vehicles (AUVs) are rapidly developing towards productization and serialization. Compared to remotely operated vehicles (ROVs), AUVs lack optical cables connecting to their mother vessels for energy and data transmission. While AUVs have a wider range, they are also constrained by the limited energy they can carry on a single trip. Efficient energy replenishment and information transmission also present technical challenges for autonomous underwater recovery. Compared to surface recovery, underwater docking offers advantages such as greater concealment and reduced interference from sea conditions. To achieve underwater docking and recovery of AUVs, positioning and guidance technology is a key step. Currently, three types of sensors are commonly used for positioning and guidance: acoustic, optical, and electromagnetic. Acoustic sensors have long ranges, low data transmission rates, and are relatively insensitive to environmental influences, but their calibration process is relatively complex. Electromagnetic sensors are limited by the underwater range of electromagnetic fields and are relatively large in size. Therefore, they are generally not used in positioning and guidance for various docking methods. The close-range detection accuracy of optical sensors during AUV docking or recovery can reach the centimeter level. However, due to the influence of seawater clarity, background light, and seawater scattering, how to ensure efficient and accurate identification and positioning remains a technical difficulty in underwater guided recovery.

[0003] The present invention relates to an underwater optical guidance recovery docking method. The method first uses an adaptive selection algorithm to preprocess the captured light source image, then uses a geometric threading algorithm to identify and extract the semantic features of the light source, matches the light source features between frames, and then solves the relative pose of the recovery dock through a relative pose estimation module. Finally, a sliding window optimization module is used to perform nonlinear optimization on the historical AUV pose and light source features to improve trajectory accuracy.

[0004] The patent application with publication date of July 8, 2015, publication number CN104766312A, and invention name “A method for autonomous docking of an intelligent underwater robot based on binocular visual guidance”, although the optical visual guidance part of this method achieves the same purpose as the present invention, the specific methods used in light source identification, light source matching and light source position calculation are completely different from the invention.

[0005] The patent application, published on September 29, 2020, with publication number CN111721259A and titled "Underwater Robot Recovery and Positioning Method Based on Binocular Vision," implements a single-light-source guidance method without any spurious light source rejection. This invention enables multi-light-source base station positioning in complex interference environments.

[0006] The patent application with a publication date of July 13, 2021, publication number CN113109762A, and invention name "A light vision guidance method for AUV docking and recovery". Although the light vision guidance of this method achieves the same purpose as the present invention, it focuses on binocular matching ranging. The present invention performs semantic recognition between light sources and back-end rolling nonlinear optimization, which can improve the positioning guidance accuracy. Summary of the Invention

[0007] The object of the present invention is to provide an optical vision positioning method for AUV recovery docking, comprising the following steps:

[0008] During the guided recovery process, the AUV uses its onboard underwater binocular camera to take interval shots, acquiring binocular images of the docking base station entrance, which include four sets of underwater optical beacons installed around the docking base station entrance and arranged at the four corners of a square. The binocular images include left and right images of the same frame.

[0009] Grayscale processing is performed on the binocular image to obtain a binarized binocular image; target contour extraction is performed on the binarized binocular image to obtain a candidate contour set; four groups of underwater optical beacon contours are screened from the candidate contour set, and the pixel coordinates of the four groups of underwater optical beacons in the image are obtained;

[0010] Match the binarized binocular image of the current frame to obtain the pixel coordinates of the binarized left image and the pixel coordinates of the binarized right image of each underwater optical beacon in the current frame; match the current frame with the binarized left image of the previous frame to obtain the pixel coordinates of the binarized left image and the pixel coordinates of the binarized left image of each underwater optical beacon in the current frame;

[0011] The real spatial positions of the four underwater optical beacons are combined with their pixel coordinates in the current frame's binarized binocular image to solve the relative pose of the docking base station and obtain the homogeneous extrinsic parameter matrix of the underwater binocular camera carried by the AUV.

[0012] If the sliding window threshold has been reached from the beginning to the current frame, the sliding window optimization is performed to perform nonlinear optimization on all homogeneous external parameter matrices in the current frame and a period of history to improve the trajectory accuracy.

[0013] Furthermore, during the guided recovery process, the AUV uses an adaptive exposure time adjustment method when taking fixed-interval photos with its onboard underwater binocular camera to ensure that the imaging of underwater optical beacons in the binocular image is not affected by ambient light. For each frame of the binocular image, the sizes of the four groups of underwater optical beacons are detected. If the sizes of the four groups of underwater optical beacons in the current frame of the binocular image are smaller than the minimum size threshold, the exposure time is increased; if the sizes of the four groups of underwater optical beacons in the current frame of the binocular image are larger than the maximum size threshold, the exposure time is reduced.

[0014] Furthermore, the target contour is extracted from the binarized binocular image, and the outermost boundary tracking method is used to include contours whose contour edge pixels meet the requirements into the candidate contour set; if the number of candidate contours extracted from the image is less than 4, the grayscale processing threshold is lowered, the original color image of the image is grayscale processed again, and the target contour is extracted again until the number of candidate contours is greater than or equal to 4.

[0015] Furthermore, the four groups of underwater optical beacon profiles are selected from the candidate profile set, specifically:

[0016] Ellipse fitting is performed on the candidate contours, and the center of each candidate contour ellipse fitting is used as a vertex to obtain a vertex set; four different vertices are randomly selected from the vertex set to form a quadrilateral, thereby constructing a candidate quadrilateral set; geometric determination is performed on each quadrilateral in the candidate quadrilateral set to determine the quadrilateral formed by the four groups of underwater optical beacons;

[0017] Step 3.1: If a vertex in the quadrilateral is inside the triangle formed by the other three vertices, the quadrilateral is determined to be a concave quadrilateral and is removed from the candidate quadrilateral set;

[0018] Step 3.2: If any of the interior angles in the quadrilateral is less than 60°, the quadrilateral is determined to be an irregular trapezoid and is removed from the candidate quadrilateral set.

[0019] Step 3.3: If the longest line segment between any two vertices of the quadrilateral is longer than twice the length of the shortest line segment, the quadrilateral is determined to be an irregular trapezoid and is removed from the candidate quadrilateral set;

[0020] Step 3.4: For the remaining quadrilaterals in the candidate quadrilateral set, fit any three vertices of the quadrilateral into a circle. Calculate the absolute value of the difference between the radius and the distance from the center of the circle to the other vertex. Take the largest absolute value of the difference in each quadrilateral as the circle fitting error of the quadrilateral.

[0021] The quadrilateral with the smallest corresponding circle fitting error is taken as the quadrilateral formed by the four groups of underwater optical beacons, and the pixel coordinates of the four vertices in the quadrilateral are obtained as the pixel coordinates of the four groups of underwater optical beacons in the image.

[0022] Furthermore, the matching of the binarized binocular image of the current frame is specifically performed as follows:

[0023] The pixel coordinates of the i-th underwater optical beacon in the binary left image of the current frame are X li (t m )=(u li (t m ),v li (t m )), i=1,2,3,4, t m For the current frame, the method of propagating the light source features to the binary right image is used to predict the pixel coordinates of the i-th underwater light beacon in the binary left image in the binary right image. Expressed as:

[0024]

[0025] Prediction based on four sets of pixel coordinates Construct the predicted bounding box, take the quadrilateral formed by the four groups of underwater light beacons in the current frame binary right image as the actual bounding box, calculate the intersection-over-union (IOU) distance matrix between the actual bounding box and the predicted bounding box, select the Hungarian algorithm as the matcher, and only use the IOU distance between ellipses in the cost matching to solve the problem with the minimum total cost (a u1 ,a v1 ); According to the solution (a u1 ,a v1 ), binarize the current frame to the corresponding right image The smallest j-th underwater optical beacon is used as the matching result of the i-th underwater optical beacon in the binarized left image of the current frame.

[0026] The matching of the current frame with the binarized left image of the previous frame is specifically as follows:

[0027] The pixel coordinates of the i-th underwater optical beacon in the binary left image of the previous frame are predicted by using the method of propagating the light source characteristics. Expressed as:

[0028]

[0029] Wherein, Δt is the time interval between two adjacent frames;

[0030] Prediction based on four sets of pixel coordinates Construct the predicted bounding box, take the quadrilateral formed by the four groups of underwater light beacons in the current frame binary left image as the actual bounding box, calculate the intersection-over-union distance matrix between the actual bounding box and the predicted bounding box, select the Hungarian algorithm as the matcher, and only use the intersection-over-union distance between ellipses in the cost matching to solve the problem with the minimum total cost (au2 ,a v2 ); According to the solution (a u2 ,a v2 ), binarize the current frame to the left image corresponding to The smallest j-th underwater optical beacon is used as the matching result of the i-th underwater optical beacon in the previous frame of the binary left image.

[0031] Furthermore, the real spatial positions of the four groups of underwater optical beacons are combined with their pixel coordinates in the binary binocular image of the current frame to solve the relative position of the docking base station and obtain the homogeneous extrinsic parameter matrix of the underwater binocular camera carried by the AUV, which is specifically:

[0032] Step 6.1: Using the DLT algorithm, construct eight sets of equations using the pixel coordinates and 3D coordinates of the four sets of underwater optical beacons in the binarized left image, and solve them to obtain the homography matrix H;

[0033]

[0034] Among them, X Di =(x Di ,y Di ,z Di ) is the real spatial position of the i-th underwater optical beacon, i.e., the three-dimensional coordinate;

[0035] Step 6.2: Decompose the homography matrix H to obtain the initial values ​​of the rotation matrix and translation matrix of the underwater binocular camera;

[0036]

[0037] Among them, the internal parameter matrix K of the underwater binocular camera is known, Rotation matrix of underwater binocular camera Translation Matrix The scaling factor s satisfies the following conditions:

[0038]

[0039] Let s = 1, solve and get the first two columns of the rotation matrix R. Since the last column of the rotation matrix R must be orthogonal to the other columns, get the last column of elements through the cross product operation and get the initial value of the rotation matrix R. Use the third column vector of the homography matrix H as the initial value of the translation matrix T.

[0040] Step 6.3: Construct the homogeneous extrinsic parameter matrix of the underwater binocular camera The least squares cost function of :

[0041]

[0042] Among them, X ci =(xci ,y ci ,z ci ), b is the binocular baseline length of the underwater binocular camera;

[0043] The initial values ​​of the rotation matrix and the translation matrix are iterated and solved until C converges and output. The three-dimensional coordinates of the centers of the four groups of underwater optical beacons are (x c (t m ),y c (t m ),z c (t m )), x c (t m )=T 00 ,y c (t m )=T 10 , z c (t m )=T 20 .

[0044] Furthermore, if the sliding window threshold n has been reached from the beginning to the current frame, a sliding window optimization is performed to perform nonlinear optimization on all homogeneous external parameter matrices in the current frame and a period of history to improve the trajectory accuracy. Specifically:

[0045] Construct the set β={C(t m -n),...,C(t),...,C(t m )},t m is the current frame, n is the total length of the optimized sliding window; Q={X D1 ,X D2 ,X D3 ,X D4};

[0046] Constructing the optimization problem:

[0047]

[0048] Where, e1(C(t),C(t+1))=log((C(t)) -1 C(t+1)), represents the relative motion constraint between adjacent camera poses; Represents the camera's observation projection error of the light source; ρ(·) represents the robust kernel function;

[0049] The camera pose node C is optimized by the fixed position Q of four sets of underwater optical beacons, focusing on the output results of the current frame. If the distance between the AUV and the docking base station is less than the threshold, it means that the docking and recovery mission has entered the late stage. The distance to the recovery base station is close enough, the pose estimation is accurate enough, and the sliding window threshold n can be reduced.

[0050] A computer device / equipment / system includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-mentioned optical vision positioning method for AUV recovery and docking.

[0051] A computer-readable storage medium stores a computer program / instruction, which, when executed by a processor, implements the steps of the optical vision positioning method for AUV recovery and docking.

[0052] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned optical vision positioning method for AUV recovery and docking.

[0053] The beneficial effects of the present invention are:

[0054] The present invention uses an underwater binocular camera carried by an AUV to perform fixed-interval shooting, and adopts a method of adaptively adjusting the exposure time to ensure that the imaging of the underwater optical beacon in the binocular image is not affected by ambient light; the outermost boundary tracking method is used to extract the target contour from the binary binocular image, and the geometric thread algorithm is used to screen out the contours of four groups of underwater optical beacons; the target tracking method is used to achieve the matching of each light source feature between single-frame binoculars and global image sequence tracking, and the real spatial position of the underwater optical beacon is used to solve the relative posture of the docking base station by combining analytical and iterative methods. The sliding window is used to perform nonlinear optimization on all AUV postures and light source features within a period of time in history to improve the trajectory accuracy. The present invention can assist the AUV in accurately identifying the posture of the recovery dock under the interference of the underwater environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is the overall architecture diagram of the present invention.

[0056] Figure 2 It is a schematic diagram of the installation of the docking base station and the underwater optical beacon in the present invention.

[0057] Figure 3 FIG. 4 is a flowchart of collecting images in an embodiment of the present invention.

[0058] Figure 4 This is a flowchart of light source semantic recognition in an embodiment of the present invention.

[0059] Figure 5This is a flow chart of light source matching and tracking in an embodiment of the present invention.

[0060] Figure 6 4 is a flowchart of relative pose estimation in an embodiment of the present invention.

[0061] Figure 7 4 is a flow chart of sliding window optimization in an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The present invention will be further described below with reference to the accompanying drawings.

[0063] The present invention discloses an optical vision positioning method for AUV recovery docking, which includes image acquisition, light source semantic recognition, light source matching and tracking, relative pose estimation and sliding window optimization.

[0064] like Figure 2 As shown, the docking base station is a truncated cone structure with the largest entrance radius and a gradually decreasing radius from the entrance to the exit. Four sets of underwater optical beacons are installed on the circumference of the entrance, and the four sets of underwater optical beacons are arranged at the four corners of the square.

[0065] Step 1: During the guided recovery process, the AUV uses its onboard underwater binocular camera to take pictures at regular intervals to obtain binocular images of the docking base station entrance containing four sets of underwater optical beacons. The binocular images include left and right images of the same frame.

[0066] During shooting, an adaptive exposure time adjustment method is used to ensure that the imaging of underwater optical beacons in the binocular image is not affected by ambient light. For each binocular image frame, the sizes of the four groups of underwater optical beacons are detected. If the sizes of the four groups of underwater optical beacons in the current binocular image frame are smaller than the minimum size threshold, the exposure time is increased; if the sizes of the four groups of underwater optical beacons in the current binocular image frame are larger than the maximum size threshold, the exposure time is reduced.

[0067] Step 2: Perform grayscale processing on the binocular image to obtain a binarized binocular image; use the outermost boundary tracking method to extract the target contour from the binarized binocular image, and include the contours with the required number of edge pixels into the candidate contours; for the binarized left or right image, if the number of candidate contours extracted from the image is less than 4, reduce the grayscale processing threshold, reconstruct the image, and perform target contour extraction again until the number of candidate contours is greater than or equal to 4;

[0068] Step 3: Using the geometric threading algorithm, the contours of four groups of underwater optical beacons are screened from the candidate contours of the binary left and right images, and the pixel coordinates of the four groups of underwater optical beacons in the images are obtained;

[0069] Perform ellipse fitting on the candidate contours, use the center of each candidate contour ellipse as a vertex, and obtain a vertex set; randomly select four different vertices from the vertex set to form a quadrilateral, and construct a candidate quadrilateral set;

[0070] Performing geometric determination on each quadrilateral in the candidate quadrilateral set to determine the quadrilateral formed by the four groups of underwater optical beacons;

[0071] Step 3.1: If a vertex in the quadrilateral is inside the triangle formed by the other three vertices, the quadrilateral is determined to be a concave quadrilateral and is removed from the candidate quadrilateral set;

[0072] Step 3.2: If any of the interior angles in the quadrilateral is less than 60°, the quadrilateral is determined to be an irregular trapezoid and is removed from the candidate quadrilateral set.

[0073] Step 3.3: If the longest line segment between any two vertices of the quadrilateral is longer than twice the length of the shortest line segment, the quadrilateral is determined to be an irregular trapezoid and is removed from the candidate quadrilateral set;

[0074] Step 3.4: For the remaining quadrilaterals in the candidate quadrilateral set, fit any three vertices of the quadrilateral into a circle. Calculate the absolute value of the difference between the radius and the distance from the center of the circle to the other vertex. Take the largest absolute value of the difference in each quadrilateral as the circle fitting error of the quadrilateral.

[0075] The quadrilateral with the smallest corresponding circle fitting error is selected as the quadrilateral formed by the four groups of underwater optical beacons, and the pixel coordinates of the four vertices in the quadrilateral are obtained as the pixel coordinates of the four groups of underwater optical beacons in the image;

[0076] Step 4: Use the target tracking method to match the binarized binocular image of the current frame and obtain the pixel coordinates (u li (t m ),v li (t m )) and the pixel coordinates of the binary right image (u ri (t m ),v ri (t m ));

[0077] The pixel coordinates of the i-th underwater optical beacon in the binary left image of the current frame are X li (t m )=(u li (t m ),v li (t m )), i=1,2,3,4, t mFor the current frame, the method of propagating the light source features to the binary right image is used to predict the pixel coordinates of the i-th underwater light beacon in the binary left image in the binary right image. Expressed as:

[0078]

[0079] Prediction based on four sets of pixel coordinates Construct the predicted bounding box, take the quadrilateral formed by the four groups of underwater light beacons in the current frame binary right image as the actual bounding box, calculate the intersection-over-union (IOU) distance matrix between the actual bounding box and the predicted bounding box, select the Hungarian algorithm as the matcher, and only use the IOU distance between ellipses in the cost matching to solve the problem with the minimum total cost (a u1 ,a v1 ); According to the solution (a u1 ,a v1 ), binarize the current frame to the corresponding right image The smallest j-th underwater optical beacon is used as the matching result of the i-th underwater optical beacon in the binary left image of the current frame;

[0080] Step 5: Use the target tracking method to match the current frame with the binary left image of the previous frame to obtain the pixel coordinates (u li (t m ),v li (t m )) and the pixel coordinates of the previous frame of the binary left image (u li (t m -1),v li (t m -1);

[0081] The pixel coordinates of the i-th underwater optical beacon in the binary left image of the previous frame are predicted by using the method of propagating the light source characteristics. Expressed as:

[0082]

[0083] Wherein, Δt is the time interval between two adjacent frames;

[0084] Prediction based on four sets of pixel coordinates Construct the predicted bounding box, take the quadrilateral formed by the four groups of underwater light beacons in the current frame binary left image as the actual bounding box, calculate the intersection-over-union distance matrix between the actual bounding box and the predicted bounding box, select the Hungarian algorithm as the matcher, and only use the intersection-over-union distance between ellipses in the cost matching to solve the problem with the minimum total cost (a u2 ,a v2 ); According to the solution (a u2,a v2 ), binarize the current frame to the left image corresponding to The smallest j-th underwater optical beacon is used as the matching result of the i-th underwater optical beacon in the previous frame of the binary left image;

[0085] Step 6: According to the three-dimensional coordinates X of each underwater light beacon Di =(x Di ,y Di ,z Di ), obtain the rotation matrix R and translation matrix T of the underwater binocular camera;

[0086] Step 6.1: Using the DLT algorithm, construct eight sets of equations using the pixel coordinates and 3D coordinates of the four sets of underwater optical beacons in the binarized left image, and solve them to obtain the homography matrix H;

[0087]

[0088] Step 6.2: Decompose the homography matrix H to obtain the initial values ​​of the rotation matrix and translation matrix of the underwater binocular camera;

[0089]

[0090] Among them, the internal parameter matrix K of the underwater binocular camera is known, Rotation matrix of underwater binocular camera Translation Matrix The scaling factor s satisfies the following conditions:

[0091]

[0092] Let s = 1, solve and get the first two columns of the rotation matrix R. Since the last column of the rotation matrix R must be orthogonal to the other columns, get the last column of elements through the cross product operation and get the initial value of the rotation matrix R. Use the third column vector of the homography matrix H as the initial value of the translation matrix T.

[0093] Step 6.3: Construct the external parameter matrix of the underwater binocular camera The least squares cost function is solved by iteratively using the initial values ​​of the rotation matrix and the translation matrix until C converges and is output. The three-dimensional coordinates of the centers of the four groups of underwater optical beacons obtained are (x c (t m ),y c (t m ),z c (t m )), x c (t m )=T 00 ,y c (t m )=T 10 , zc (t m )=T 20 ;

[0094]

[0095] Among them, X ci =(x ci ,y ci ,z ci ), b is the binocular baseline length of the underwater binocular camera;

[0096] Step 7: If the threshold n has been reached from the beginning to the current frame, a sliding window optimization is performed to perform nonlinear optimization on all homogeneous external parameter matrices in the current frame and a period of history to improve trajectory accuracy. If the distance between the AUV and the docking base station is less than the threshold, it means that the docking and recovery mission has entered the late stage and the distance to the recovery base station is close enough. The pose estimation is accurate enough and the threshold n can be reduced.

[0097] Construct the set β={C(t m -n),...,C(t),...,C(t m )},t m is the current frame, n is the total length of the optimized sliding window; Q={X D1 ,X D2 ,X D3 ,X D4};

[0098] Constructing the optimization problem:

[0099]

[0100] Where, e1(C(t),C(t+1))=log((C(t)) -1 C(t+1)), represents the relative motion constraint between adjacent camera poses; Represents the camera's observation projection error of the light source; ρ(·) represents the robust kernel function; the camera pose node C is optimized by the fixed position Q of four groups of underwater optical beacons, focusing on the output result of the current frame.

[0101] Example 1:

[0102] Figure 1 The system flow of the present invention is described, and the specific implementation steps are as follows:

[0103] Step (1): Image acquisition: During the optical guided recovery phase, the underwater binocular camera carried by the AUV is used to take fixed-interval images, the images are pre-processed by the proposed adaptive selection algorithm, and continuously transmitted to the optical guided recovery docking program to enter step (2);

[0104] Step (2): Light source semantic recognition and matching: Using the proposed geometric threshold algorithm, the semantic features of the light source, including position and size information, are identified and extracted, and the target tracking method is used to achieve inter-frame matching of each light source feature, and then proceed to step (3);

[0105] Step (3): Relative pose estimation: Combining analytical and iterative methods, the relative pose of the base station is solved using the spatial position of the light source center and added to the mileage record, and then proceed to step (4);

[0106] Step (4): Sliding window optimization: Perform nonlinear optimization on all AUV poses and light source features within a period of time in the history to improve trajectory accuracy and proceed to step (5);

[0107] Step (5): Output the three-dimensional coordinates of the base station (x o ,y o ,z o ),Finish.

[0108] Figure 2 This is a schematic diagram of the mechanical structure of the docking base station of the present invention. Considering the requirements of mobile docking scenarios and to minimize fluid force interference, the AUV's underwater optically guided docking device primarily consists of an open, tapered cone 1, which gradually narrows toward the center to guide the AUV and correct its trajectory. It is equipped with four underwater optical beacons 2 with a wavelength of 540nm, symmetrically placed on the docking ring. The docking device is secured to the underside of the mother ship or test vehicle by a connecting frame 3 to enable dynamic docking.

[0109] Figure 3 This is the image acquisition flow chart of the present invention, and the specific implementation steps are as follows:

[0110] Step (1): Set the initial exposure time, capture the image, and proceed to step (2);

[0111] Step (2): Determine the light source size based on the light source detection result of the previous frame. If the light source size is smaller than size, min , go to step (3), otherwise go to step (4), if it is the first frame, shoot according to the initial exposure value;

[0112] Step (3): Increase the exposure time by Exposure = 1.25 × Exposure to obtain more light source contours, and proceed to step (6);

[0113] Step (4): Determine the contour point of the light source detected in the previous frame. If the light source size is greater than size max , go to step (5), otherwise go to step (6);

[0114] Step (5): Reduce the exposure time Exposure = 0.8 × Exposure. This process usually occurs when approaching the recycling dock, and then proceed to step (6).

[0115] Step (6): grayscale the input color image, with the grayscale value range being [0, 255], and proceed to step (7);

[0116] Step (7): Set the initial grayscale processing threshold Gray, set the pixel value to 0 if the pixel value is less than the threshold value, and set it to 1 otherwise, to obtain a binary image, and proceed to step (8);

[0117] Step (8): Use the outermost boundary tracking method to extract the target contour from the binary image, include the contours whose edge pixels meet the requirements into the candidate contours, and determine the number of candidate contours N c Is it greater than 4? If not, go to step (9); otherwise, go to step (10);

[0118] Step (9): Lower the grayscale processing threshold Gray = Gray-T Gray , return to step (7);

[0119] Step (10): Finally, a binary image and candidate contours are obtained, and the image acquisition process ends.

[0120] Figure 4 This is the light source semantic recognition flow chart of the present invention, and the specific implementation steps are as follows:

[0121] Step (1): First, perform ellipse fitting on the candidate contours to obtain the center coordinates of each contour and use them as the vertex P = {p i ,i=1,2,...,N}, and perform Arrange the obtained quadrilateral candidate array Perform geometric judgment and proceed to step (2);

[0122] Step (2): Concave quadrilateral detection: By checking whether the fourth corner point of each candidate quadrilateral is within the triangle formed by the line connecting the other three corner points, concave quadrilaterals with an internal angle greater than 180° are eliminated. If it is determined to be a concave quadrilateral, proceed to step (6); otherwise, proceed to step (3);

[0123] Step (3): Irregular quadrilateral detection: This step includes two parts of detection. The first is whether the internal angle deviates too much from 90°. If the internal angle is less than 60°, it can be considered as an irregular quadrilateral. The second is the problem of excessive side length. Compare the six line segments in the quadrilateral (including the two diagonals) to check whether the longest line segment is greater than twice the shortest line segment:

[0124]

[0125] If it is determined to be an irregular trapezoid, go to step (6), otherwise go to step (4);

[0126] Step (4): Cocircularity detection: In order to eliminate the influence of glare on light source identification, the minimum circle fitting error is checked by the cocircularity test principle. Considering that the candidate quadrilateral containing glare is similar in size to the actual base station light source quadrilateral, the circle fitting error is redefined as:

[0127] Δr=max||r i=1,2,3,4 -Op j=1,2,3,4 ||

[0128] Where r is the radius of the circle fitted to the three vertices, and Op is the length from the center of the circle to the fourth vertex. The largest distance is selected as the circle fitting error to reduce the false negative rate. If the cocircularity test is passed, proceed to step (5), otherwise proceed to step (6);

[0129] Step (5): Determine whether the quadrilateral candidate array Q is traversed. If each candidate passes the geometric thread algorithm, proceed to step (7); otherwise, proceed to step (6).

[0130] Step (6): Select the next candidate array and proceed to step (3);

[0131] Step (7): Output the light source semantic set and end.

[0132] Figure 5 This is a flow chart of light source matching and tracking of the present invention, and the specific implementation steps are as follows:

[0133] Step (1): Binocular target matching: Using the geometric threading algorithm described above, a unique base station light source array is identified, and a tracking-based method is used to find a semantic pairing for the right image. This invention uses a representation method and motion model that propagates light source features to the right image. Considering the camera motion, it is assumed that the speed of the light source moving from the left image to the right image is constant over a short period of time:

[0134]

[0135] Among them, u, v are the pixel coordinates of the light source, and the values ​​w and h represent the pixel width and height of the light source object, respectively, which come from the envelope bounding box given by the light source ellipse semantic feature. Four light source objects are initialized in each frame image, and the speeds are and A Kalman filter algorithm is used to track these light sources.

[0136] After estimating the position of each tracked object bounding box by predicting its new position in the current right image, the intersection-over-union distance matrix between each bounding box from the light source array and all predicted bounding boxes is calculated. At the same time, the Hungarian algorithm is selected as the matcher, and the minimum cost allocation is solved from the loss intersection-over-union distance:

[0137] LIoU = 1-IoU

[0138] In the cost matching, only the intersection-over-union distance between ellipses is used, and the matching combination with the minimum total cost is selected to complete the feature matching of the left and right binocular images.

[0139] Step (2): Global light source tracking, also using the above steps and methods. Maintain a continuous tracking on the global sequence of the left image, starting from the first valid frame. Subsequent frames are updated and predicted through Kalman filtering to ensure the temporal consistency of the light source in the left image. Global feature matching and tracking are completed.

[0140] Figure 6 This is the relative pose estimation flow chart of the present invention, and the specific implementation steps are as follows:

[0141] Step (1): Calculate the preliminary pose estimate using the homography matrix: The homography transformation that projects a 2D point from the docking station light array plane to the image plane can be defined as:

[0142]

[0143] Among them, (u, v) represents the pixel coordinates, X′ D (x D ,y D ) is the docking station coordinate system plane O D X D Y D The coordinates on , go to step (2);

[0144] Step (2): Decompose the homography matrix to obtain the initial external parameters: The 3×3 homography matrix H has 8 degrees of freedom, corresponding to 8 equations, and requires 4 point pairs. Therefore, the DLT algorithm is used to calculate H using the four 2D point pairs corresponding to the optical beacon. Then, considering the pinhole model, the camera's projection transformation is also expressed as:

[0145]

[0146] Among them, X D (x D ,y D ,z D ) Base station coordinate system O D X D Y D Z DThe coordinates of the camera are as follows, K is the camera intrinsic parameter, including the focal length and the pixel coordinates of the optical center, s is the scale factor, and P = [R|T] is the camera extrinsic matrix, including translation and rotation. The pose from DS to the camera can be obtained from the extrinsic matrix. Among them, all four light sources are coplanar, that is, z D =0, so we can simplify:

[0147]

[0148] Where P′ represents the simplified extrinsic parameter matrix without the third column element. The relationship between the homography matrix and the extrinsic and intrinsic parameters can be obtained:

[0149]

[0150] As mentioned above, the homography matrix is ​​known and is given by the DLT algorithm. Therefore, a set of elements R can be obtained from the above formula ij and T ij , while the scale factor s cannot be calculated directly. Since the columns of the rotation matrix must be normalized and orthogonal, the scale factor s is the geometric mean of the first two columns of the rotation matrix:

[0151]

[0152] where R i ' j It is assumed that the size is s = 1. Since the light source is always in front of the AUV, T 20 <0. Then, the real elements of the P′ part can be recovered by dividing by the scale factor s. The last column of the rotation matrix must be orthogonal to the other columns, so R can be obtained by cross product operation. 02 , R 12 and R 22 Singular value decomposition is used to minimize the error of the Frobenius matrix norm. The translation vector T is directly extracted through the third column of H and divided by s:

[0153]

[0154] Get the initial external parameter matrix P of the camera and go to step (3);

[0155] Step (3): Improve the estimate by constructing an optimization problem to minimize the measurement error: The initial estimate of the extrinsic parameter matrix calculated by purely analytical methods is very sensitive to measurement noise. This measurement noise comes from camera correction and the estimation of the light source center point, so it is unavoidable. To improve this problem, the theory of graph optimization is used to further refine the extrinsic parameter matrix. First, a least squares cost function is constructed:

[0156]

[0157] Among them, X ci =(x ci ,y ci ,z ci ), b is the binocular baseline length of the underwater binocular camera;

[0158] Among them, all variables are known and can be solved iteratively through the optimization library g2o, using the initial values ​​of the rotation matrix and the translation matrix to iteratively solve until C converges and outputs. The three-dimensional coordinates of the centers of the four groups of underwater optical beacons obtained are (x c (t m ),y c (t m ),z c (t m )), x c (t m )=T 00 ,y c (t m )=T 10 , z c (t m )=T 20 .

[0159] Figure 7 This is the flow chart of the sliding window optimization posture of the present invention, and the specific implementation steps are as follows:

[0160] Step (1): Input the pose estimation result of each frame into the sliding window optimization module and proceed to step (2);

[0161] Step (2): Determine whether the current accumulated number of historical poses is greater than 10. If so, proceed to step (3); otherwise, return to step (1).

[0162] Step (3): Determine whether the pose estimation distance to the base station is less than D threshold , if the current distance is less than D threshold , go to step (4), otherwise go to step (5) according to the total amount of sliding window n = 50;

[0163] Step (4): Reduce the total sliding window amount n. This means that the docking recovery task has entered the late stage, the distance to the recovery base station is close enough, the pose estimation is accurate enough, and the sliding window amount required for optimization is reduced.

[0164] Step (5): Perform sliding window pose optimization. First, define the optimization problem:

[0165] Construct the set β={C(t m -n),...,C(t),...,C(t m )},t mis the current frame, n is the total length of the optimized sliding window; Q={X D1 ,X D2 ,X D3 ,X D4};

[0166] Constructing the optimization problem:

[0167]

[0168] Where, e1(C(t),C(t+1))=log((C(t)) -1 C(t+1)), represents the relative motion constraint between adjacent camera poses; Represents the camera's observation projection error of the light source; ρ(·) represents the robust kernel function, which is used to suppress the influence of outliers.

[0169] The camera pose node C is optimized by the fixed position Q of four sets of underwater optical beacons, focusing on the output results of the current frame. If the distance between the AUV and the docking base station is less than the threshold, it means that the docking and recovery mission has entered the late stage. The distance to the recovery base station is close enough, the pose estimation is accurate enough, and the sliding window threshold n can be reduced.

[0170] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A light vision positioning method for AUV recovery docking, characterized by: During the guided recovery process, the AUV uses its onboard underwater binocular camera to take interval shots, acquiring binocular images of the docking base station entrance, which include four sets of underwater optical beacons installed around the docking base station entrance and arranged at the four corners of a square. The binocular images include left and right images of the same frame. Grayscale processing is performed on the binocular image to obtain a binarized binocular image; target contour extraction is performed on the binarized binocular image to obtain a candidate contour set; four groups of underwater optical beacon contours are screened from the candidate contour set, and the pixel coordinates of the four groups of underwater optical beacons in the image are obtained; Match the binarized binocular image of the current frame to obtain the pixel coordinates of the binarized left image and the pixel coordinates of the binarized right image of each underwater optical beacon in the current frame; match the current frame with the binarized left image of the previous frame to obtain the pixel coordinates of the binarized left image and the pixel coordinates of the binarized left image of each underwater optical beacon in the current frame; The real spatial positions of the four underwater optical beacons are combined with their pixel coordinates in the current frame's binarized binocular image to solve the relative pose of the docking base station and obtain the homogeneous extrinsic parameter matrix of the underwater binocular camera carried by the AUV. If the sliding window threshold has been reached from the beginning to the current frame, the sliding window optimization is performed to perform nonlinear optimization on all homogeneous external parameter matrices in the current frame and a period of history to improve the trajectory accuracy.

2. The optical vision positioning method for AUV recovery docking according to claim 1, characterized in that: During the guided recovery process, the AUV uses an adaptive exposure time adjustment method when taking fixed-interval photos with its onboard underwater binocular camera to ensure that the imaging of underwater optical beacons in the binocular image is not affected by ambient light. For each frame of the binocular image, the sizes of the four groups of underwater optical beacons are detected. If the sizes of the four groups of underwater optical beacons in the current frame of the binocular image are smaller than a minimum size threshold, the exposure time is increased; if the sizes of the four groups of underwater optical beacons in the current frame of the binocular image are larger than a maximum size threshold, the exposure time is reduced.

3. The optical vision positioning method for AUV recovery docking according to claim 1, characterized in that: The target contour is extracted from the binarized binocular image, and the outermost boundary tracking method is adopted to include the contours whose contour edge pixel numbers meet the requirements into the candidate contour set; If the number of candidate contours extracted from the image is less than 4, the grayscale processing threshold is lowered, the original color image is grayscale processed again, and the target contour is extracted again until the number of candidate contours is greater than or equal to 4.

4. The optical vision positioning method for AUV recovery docking according to claim 1, characterized in that: The four groups of underwater optical beacon profiles are selected from the candidate profile set, specifically: Ellipse fitting is performed on the candidate contours, and the center of each candidate contour ellipse fitting is used as a vertex to obtain a vertex set; four different vertices are randomly selected from the vertex set to form a quadrilateral, thereby constructing a candidate quadrilateral set; geometric determination is performed on each quadrilateral in the candidate quadrilateral set to determine the quadrilateral formed by the four groups of underwater optical beacons; Step 3.1: If a vertex in the quadrilateral is inside the triangle formed by the other three vertices, the quadrilateral is determined to be a concave quadrilateral and is removed from the candidate quadrilateral set; Step 3.2: If any of the interior angles in the quadrilateral is less than 60°, the quadrilateral is determined to be an irregular trapezoid and is removed from the candidate quadrilateral set. Step 3.3: If the longest line segment between any two vertices of the quadrilateral is longer than twice the length of the shortest line segment, the quadrilateral is determined to be an irregular trapezoid and is removed from the candidate quadrilateral set; Step 3.4: For the remaining quadrilaterals in the candidate quadrilateral set, fit any three vertices of the quadrilateral into a circle. Calculate the absolute value of the difference between the radius and the distance from the center of the circle to the other vertex. Take the largest absolute value of the difference in each quadrilateral as the circle fitting error of the quadrilateral. The quadrilateral with the smallest corresponding circle fitting error is taken as the quadrilateral formed by the four groups of underwater optical beacons, and the pixel coordinates of the four vertices in the quadrilateral are obtained as the pixel coordinates of the four groups of underwater optical beacons in the image.

5. The optical vision positioning method for AUV recovery docking according to claim 1, characterized in that: The matching of the binarized binocular image of the current frame is specifically as follows: The pixel coordinates of the i-th underwater optical beacon in the binary left image of the current frame are X li (t m )=(u li (t m ),v li (t m )), i=1,2,3,4, t m For the current frame, the method of propagating the light source features to the binary right image is used to predict the pixel coordinates of the i-th underwater light beacon in the binary left image in the binary right image. Expressed as: Prediction based on four sets of pixel coordinates Construct the predicted bounding box, take the quadrilateral formed by the four groups of underwater light beacons in the current frame binary right image as the actual bounding box, calculate the intersection-over-union (IOU) distance matrix between the actual bounding box and the predicted bounding box, select the Hungarian algorithm as the matcher, and only use the IOU distance between ellipses in the cost matching to solve the problem with the minimum total cost (a u1 ,a v1 ); According to the solution (a u1 ,a v1 ), binarize the current frame to the corresponding right image The smallest j-th underwater optical beacon is used as the matching result of the i-th underwater optical beacon in the binarized left image of the current frame. The matching of the current frame with the binarized left image of the previous frame is specifically as follows: The pixel coordinates of the i-th underwater optical beacon in the binary left image of the previous frame are predicted by the method of propagating the light source characteristics. Expressed as: Among them, Δt is the time interval between two adjacent frames; Prediction based on four sets of pixel coordinates Construct the predicted bounding box, take the quadrilateral formed by the four groups of underwater light beacons in the current frame binary left image as the actual bounding box, calculate the intersection-over-union distance matrix between the actual bounding box and the predicted bounding box, select the Hungarian algorithm as the matcher, and only use the intersection-over-union distance between ellipses in the cost matching to solve the problem with the minimum total cost (a u2 ,a v2 ); According to the solution (a u2 ,a v2 ), binarize the current frame to the left image corresponding to The smallest j-th underwater optical beacon is used as the matching result of the i-th underwater optical beacon in the previous frame of the binary left image.

6. The optical vision positioning method for AUV recovery docking according to claim 5, characterized in that: The real spatial positions of the four groups of underwater optical beacons are combined with their pixel coordinates in the binary binocular image of the current frame to solve the relative position of the docking base station and obtain the homogeneous extrinsic parameter matrix of the underwater binocular camera carried by the AUV, specifically: Step 6.1: Using the DLT algorithm, construct eight sets of equations using the pixel coordinates and 3D coordinates of the four sets of underwater optical beacons in the binarized left image, and solve them to obtain the homography matrix H; Among them, X Di =(x Di ,y Di ,z Di ) is the real spatial position of the i-th underwater optical beacon, i.e., the three-dimensional coordinate; Step 6.2: Decompose the homography matrix H to obtain the initial values ​​of the rotation matrix and translation matrix of the underwater binocular camera; Among them, the internal parameter matrix K of the underwater binocular camera is known, Rotation matrix of underwater binocular camera Translation Matrix The scaling factor s satisfies the following conditions: Let s = 1, solve and get the first two columns of the rotation matrix R. Since the last column of the rotation matrix R must be orthogonal to the other columns, get the last column of elements through the cross product operation and get the initial value of the rotation matrix R. Use the third column vector of the homography matrix H as the initial value of the translation matrix T. Step 6.3: Construct the homogeneous extrinsic parameter matrix of the underwater binocular camera The least squares cost function of : Among them, X ci =(x ci ,y ci ,z ci ), b is the binocular baseline length of the underwater binocular camera; The initial values ​​of the rotation matrix and the translation matrix are iterated and solved until C converges and output. The three-dimensional coordinates of the centers of the four groups of underwater optical beacons are (x c (t m ),y c (t m ),z c (t m )), x c (t m )=T 00 ,y c (t m )=T 10 , z c (t m )=T 20 .

7. The optical vision positioning method for AUV recovery docking according to claim 6, characterized in that: If the sliding window threshold n has been reached from the beginning to the current frame, the sliding window optimization is performed to perform nonlinear optimization on all homogeneous external parameter matrices in the current frame and a period of history to improve the trajectory accuracy. Specifically: Construct the set β={C(t m -n),...,C(t),...,C(t m )},t m is the current frame, n is the total length of the optimized sliding window; Q = {X D1 ,X D2 ,X D3 ,X D4 }; Constructing the optimization problem: Where, e1(C(t),C(t+1))=log((C(t)) -1 C(t+1)), represents the relative motion constraint between adjacent camera poses; Represents the camera's observation projection error of the light source; ρ(·) represents the robust kernel function; The camera pose node C is optimized by the fixed position Q of four sets of underwater optical beacons, focusing on the output results of the current frame. If the distance between the AUV and the docking base station is less than the threshold, it means that the docking and recovery mission has entered the late stage. The distance to the recovery base station is close enough, the pose estimation is accurate enough, and the sliding window threshold n can be reduced.

8. A computer device / apparatus / system comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Intelligent underwater robot autonomous butting method based on bi-sight-vision guiding

    CN104766312A

  • Underwater robot recycling and positioning method based on binocular vision

    CN111721259A

  • Light vision guiding method for docking and recycling of AUV

    CN113109762A

Cited By

  • Optical guiding method for recovery of deep-sea unmanned underwater vehicle

    CN120991881A