An underwater target detection and pose estimation method
By designing an underwater target detection method that combines nested Apriltags and Siamese networks, the problem of inaccurate target pose estimation in ROV autonomous recovery was solved, achieving real-time and accurate pose estimation and autonomous guidance for underwater targets.
Patent Information
- Application Number
- CN202310485191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-04-28
AI Technical Summary
Existing technologies struggle to achieve autonomous and safe recovery of ROVs in underwater environments, especially due to the insufficient accuracy and robustness of optical image-guided tracking methods in underwater environments, leading to inaccurate target pose estimation.
Nested Apriltags are designed as underwater cooperative targets. Combining a monocular vision system and a Siamese network, the target region is captured by the ROI extraction network, the pose parameters are solved using PnP, and an identifier code fusion strategy is adopted to improve the positioning accuracy. The attenuation characteristics of light underwater are considered to enhance the recognition distance.
It achieves real-time and accurate pose estimation of underwater targets, reduces camera dead zones, improves positioning accuracy and robustness, and enhances the effective recognition range of the camera.
Smart Images

Figure CN116486247B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection and positioning technology, and in particular to a method for underwater target detection and pose estimation. Background Technology
[0002] The ocean is not only rich in resources but also shrouded in mystery. With the scarcity of land resources and the rapid advancement of science and technology, countries worldwide are accelerating their exploration and research of the ocean. Currently, deploying underwater observation equipment and utilizing its onboard sensors to collect vast amounts of marine information is a crucial method for human ocean research. However, the complexity and unpredictability of the marine environment, along with the mobility of the observation equipment, present numerous challenges to the autonomous and safe recovery of underwater observation devices.
[0003] ROVs, due to their advantages of high safety, maneuverability, small size, long operating time, and large diving depth, are frequently used for autonomous underwater recovery missions. Previously, ROVs were mostly equipped with sonar equipment to perceive ocean information and detect, monitor, and track underwater targets. However, the autonomous recovery process of underwater observation equipment requires precise target pose information, which sonar alone cannot meet. Visual images, with their non-invasive, passive, and high information content, can provide detailed and rich target feature information. Image processing technology can obtain precise target pose information, providing accurate data for ROVs to autonomously track and retrieve underwater observation equipment. Therefore, a recovery strategy combining acoustic and optical guidance can be used to achieve the autonomous and safe recovery of underwater observation equipment by ROVs. Currently, long-range acoustic positioning technology is relatively mature; however, the complexity and unknown nature of the underwater environment, as well as the attenuation and scattering characteristics of light during underwater propagation, present many challenges to overcome in terms of accuracy and robustness of short-range optical image-based guidance and tracking methods.
[0004] Although research on ROV image processing and target tracking methods has received high attention from experts and scholars both domestically and internationally in recent years, and related technologies have developed rapidly, many technical challenges still remain to be solved. Therefore, to meet the mission requirements under the current circumstances, it is urgent to conduct research on the real-time detection and positioning accuracy of underwater targets, laying a theoretical and technical foundation for ROVs to perform real-time autonomous underwater target docking and precise operations. Summary of the Invention
[0005] To address the aforementioned problems and technical requirements, the inventors propose an underwater target detection and pose estimation method. This method utilizes a set of monocular vision systems on an ROV to solve the target's pose parameters relative to the ROV in real time, accurately, and effectively, providing crucial technical support for ROVs to perform vision-based near-field tasks. The technical solution of this invention is as follows:
[0006] A method for underwater target detection and pose estimation includes the following steps:
[0007] The nested Apriltag is designed as an underwater cooperative target, and the nested Apriltag contains a second identifier nested in the first identifier;
[0008] The acquired video sequence images are input into the ROI extraction network to capture the target region in the video sequence images;
[0009] Target detection is performed on the captured area to obtain the pixel coordinates of the four corner points of the first and / or second identifiers. Combined with the actual physical size of the first and / or second identifiers, the pose parameters of the target relative to the ROV are solved using PnP.
[0010] If only the first or second identifier is detected, the obtained pose parameters are output directly; if both the first and second identifiers are detected at the same time, the result of fusing the pose parameters of the target relative to the ROV obtained by using the two identifiers is used as the pose parameter output.
[0011] The beneficial technical effects of this invention are:
[0012] (1) Nested Apriltags were designed as underwater cooperative targets. The advantages are: ① It can reduce the effective field-of-view dead zone of the camera. That is, when the large-sized first tag exceeds the camera's field-of-view boundary, the small-sized second tag can still be used to locate the underwater target; ② The fusion strategy can be used to improve the positioning accuracy. That is, when the first and second tags are detected at the same time, the fusion strategy based on the distance function is used to calculate the result after fusing the pose parameters of the two tags, which is used as the final pose parameter output; ③ Considering the attenuation characteristics of light propagation underwater, setting the background of the two tags to blue can effectively increase the effective recognition distance of the camera for the tags.
[0013] (2) An underwater target region of interest extraction network based on Siamese network was designed to capture the region of the target in the video sequence image, providing clear target features for the subsequent detection layer, thereby improving the detection speed and suppressing background noise.
[0014] (3) A template image update strategy is proposed for Siamese networks to extract template images of the target region of interest in the next frame, which can effectively enhance the generalization ability of the proposed method. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the underwater target detection and pose estimation method provided in this application.
[0016] Figure 2 This is a schematic diagram of the nested Apriltag provided in this application.
[0017] Figure 3 These are actual images of the underwater cameras used in this application.
[0018] Figure 4 This is a flowchart of the proposed method's visual algorithm provided in this application.
[0019] Figure 5 This is a comparative diagram showing whether or not a template update strategy is used, provided in this application.
[0020] Figure 6 These are some of the actual water tank test visualization results provided in this application.
[0021] Figure 7 This is a schematic diagram illustrating the positioning accuracy of the method proposed in this application. Detailed Implementation
[0022] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0023] like Figure 1 As shown, this embodiment provides an underwater target detection and pose estimation method, including the following steps:
[0024] Step 1: Design nested Apriltags as underwater cooperative targets.
[0025] During the identification process, the more pixels the identifier occupies under the camera's imaging plane, the more details are preserved, and the more accurate the identification will be. Therefore, when identifying at a distance, it is necessary to ensure that the identifier occupies a sufficient number of pixels. As the camera gradually approaches, it is necessary to ensure that a cooperative target can still be found within the field of view to complete the final accurate docking. Therefore, two identifiers of different sizes are nested together to form an underwater cooperative target, which includes the following steps:
[0026] Step 1.1: As Figure 2 As shown, the first identifier 1 and the second identifier 2 are constructed using alternating dark and light colors, and the first and second identifiers have different IDs. In this embodiment, considering the attenuation characteristics of light propagation underwater, blue is used for the dark color and white is used for the light color to improve the effective recognition distance of the camera underwater and improve positioning accuracy.
[0027] Step 1.2: Design the size of the identification code: Based on the farthest distance the camera can recognize and the impact of the underwater environment on the camera, determine the optimal physical size of the first identification code. The second identification code should still be found within the camera's field of view when the camera and the target are at close range, making it the optimal physical size for this purpose. In this embodiment, the actual physical size of the first identification code is designed to be 165mm × 165mm, and the actual physical size of the second identification code is designed to be 17mm × 17mm.
[0028] Step 1.3: Nesting of Identifier Codes: To ensure that the four corners of the second identifier code can be identified at close range, the second identifier code needs to be placed in the light-colored area in the middle of the first identifier code to form a nested Apriltag.
[0029] Step 2: Deploy and calibrate the cameras.
[0030] Step 2.1: A monocular camera is fixed to the bottom of the ROV, and ideally positioned below the ROV's center of gravity to minimize the impact of camera rotation and translation matrix deviations relative to the ROV on positioning accuracy. In this embodiment, the underwater monocular camera used is as follows... Figure 3 As shown, the model used is the Shengyou HD network underwater camera SW01.
[0031] Step 2.2: Before acquiring images, calibrate the monocular camera to obtain the camera's intrinsic parameters and distortion parameters, and calibrate the camera's rotation and translation matrices relative to the ROV's body coordinate system.
[0032] ;
[0033] ;
[0034] In the formula: and for x and y Focal length in direction, and This represents the translation dimension of the origin in the camera coordinate system. and The radial distortion coefficient is... and denoted as the tangential distortion coefficient.
[0035] Step 3: Input the acquired video sequence images into the Region of Interest (ROI) extraction network to capture the region of the target in the video sequence images.
[0036] During the operation, the ROV needs to first autonomously navigate to an image-recognizable area, and then use a monocular camera to acquire a video sequence of images as the raw search images, such as... Figure 4 As shown in (a) above. Before being input into the ROI extraction network, the acquired image needs to undergo distortion correction and filtering preprocessing to lay the foundation for accurate target detection and localization. In this embodiment, the ROI extraction network is implemented based on a Siamese network. The video sequence image and the template image are used as inputs to the Siamese network. The Siamese network performs a full-image search on each frame of the video sequence image based on the template image to capture the target region in the video sequence image, providing clear target features for the subsequent detection layer, thereby improving detection speed and suppressing background noise.
[0037] The network structure of the twin network and the processing procedure of each layer are referenced. Figure 1 As shown, the input template image size is required to be 127×127×3, the search image size is 255×255×3, and the output result is the region of interest (ROI) in the image. Figure 4 As shown in (b) of the diagram.
[0038] Step 4: Perform target detection on the captured area to obtain the pixel coordinates of the four corner points of the first and / or second identifier. This includes the following sub-steps:
[0039] Step 4.1: Line segment detection: Calculate the gradient intensity and direction of each pixel on the captured region image, then use a clustering algorithm to cluster the gradient intensity and direction, and use weighted least squares to fit the straight line equation. The weight of each point is the gradient intensity of that point. The coordinates of the two endpoints of the directed line segment are obtained through straight line fitting.
[0040] Step 4.2: Quadrilateral Detection: After detecting all line segments, group the segments according to the following rules: the distance between the end point of the previous edge and the start point of the next edge should be less than a threshold, and the connecting line segments should form a counterclockwise / clockwise direction. Apply a depth-first search algorithm to traverse all directed line segments. When the tree depth is 4, if the last edge forms a closed loop with the first edge, it means that four edges of the suspected identifier have been found, thus obtaining the pixel coordinates of its four corner points.
[0041] Step 4.3: Identifier Code Decoding and Matching: Since some quadrilaterals obtained may not contain identifier codes, decoding is required. Furthermore, the same identifier code may have four different values in the camera's field of view due to angle variations. Therefore, identification from different angles is necessary. Each quadrilateral is rotated three times in the same direction (clockwise / counterclockwise), and the four angles of the resulting quadrilateral are decoded separately. The decoded values are then matched against the encoding library. If a match is found, the quadrilateral is identified as an identifier code, and the pixel coordinates of its four corner points in the camera coordinate system are obtained. Otherwise, the quadrilateral is identified as a non-identifier code. The encoding library stores the binary code values and IDs of all identifier codes.
[0042] The method for decoding a quadrilateral at any angle includes: dividing the quadrilateral into n×n squares; for each square, if its pixel value is greater than a certain threshold, assigning the coordinates of its interior points to 1, otherwise assigning them to 0; extracting the coordinates of the interior points of the squares located in the valid area and stringing them together into binary code as the decoded value of the quadrilateral. In this embodiment, as... Figure 2 As shown, the nested Apriltag design is divided into 8×8 squares. The dark-colored squares around the edge are removed, and the inner 6×6 squares are retained as the effective area. The final identifier obtained at the detection layer is as follows: Figure 4 As shown in (c) in the figure.
[0043] Step 5: Based on the results of Step 4, and combined with the actual physical dimensions of the first and / or second identifier codes, use PnP to solve for the pose parameters of the target relative to the ROV.
[0044] Since the solution methods for these three cases are the same, this embodiment takes the method of solving the pose parameters of the target relative to the ROV using PnP based on the pixel coordinates of the four corner points of any identification code and the actual physical size of the identification code as an example. The specific steps include the following:
[0045] Step 5.1: Based on the pixel coordinates of the four corner points of the first or second identifier in the camera coordinate system and their coordinates in the target coordinate system, solve for the rotation matrix of the camera coordinate system relative to the target coordinate system. Translation matrix In this step, the target coordinate system is used as the world coordinate system, which is a coordinate system established with the center point of the identifier code as the origin. Geometric relationships are established using the pixel coordinates of three detected corner points and the world coordinates of those three corner points. Let the camera origin be denoted as 3D point . A , B , C 2D points are a ,b , c ,in a , b , c for A , B , C The projection onto the imaging plane, with the pixel coordinates and world coordinates of the remaining corner point serving as verification points. D It is used to select the most suitable solution from the possible solutions.
[0046] Based on geometric relationships and using the Law of Cosines, we get:
[0047] ;
[0048] For the equation above, divide both sides by... ,make , Then we have:
[0049] ;
[0050] Again , get about x and y The two quadratic equations in two variables are as follows:
[0051] ;
[0052] By solving the above quadratic equation, we obtain... , , The four solutions in the camera coordinate system pass the fourth verification point. D The coordinates were used to verify the final unique solution. , , Find the length of the line segment, and then find the point. A , B , C Coordinates in the camera coordinate system; then combined with A , B , C Calculate the rotation matrix between the camera coordinate system and the target coordinate system using the coordinates in the target coordinate system. Translation matrix .
[0053] Step 5.2: Similarly, using the ROV body coordinate system as the world coordinate system, and marking four known points on different lines within the ROV body coordinate system, and combining the pixel coordinates of the four known points captured by the camera in the camera coordinate system, solve for the rotation matrix of the camera coordinate system relative to the ROV body coordinate system. Translation matrix .
[0054] Step 5.3: Calculate the pose parameters of the target relative to the ROV based on the following formula:
[0055] ;
[0056] in, This indicates the spatial position of the target in the target coordinate system (the target coordinate system takes the target center as the origin, so it is used in this chapter). ), This represents the spatial position parameter of the target in the ROV body coordinate system. This represents the target's attitude parameters in the ROV body coordinate system.
[0057] Step 6: Based on the results of Step 5, there are two cases: ① If only the first or second identifier is detected, the obtained pose parameters are directly output. ② If both the first and second identifiers are detected simultaneously, the result of fusing the pose parameters of the target relative to the ROV obtained using the two identifiers is used as the pose parameter output. A schematic diagram of the final pose parameter output is shown below. Figure 4 As shown in (d) in the diagram. For the second case, the pose parameters of the target relative to the ROV obtained from the two identification codes are substituted into the designed distance function for fusion. The expression of the distance function is:
[0058] ;
[0059] in, This indicates the pose parameters obtained based on the second identifier. This represents the pose parameters obtained based on the first identifier. This indicates the target obtained based on the second identifier in the ROV body coordinate system. z To distance parameter, This indicates the target obtained based on the first identifier in the ROV body coordinate system. z The distance parameter is defined by the formula. When the distance between the target and the ROV is 1 meter, the pose information solved by the first or second identifier is considered to have the same reliability. As the distance between the two decreases, the pose solved by the second identifier is considered to become more reliable. Conversely, as the distance between the two increases, the pose solved by the first identifier is considered to become more reliable.
[0060] Step 7: Based on the pixel coordinates of the four corner points of the first and / or second identifier obtained in Step 4, update the template image used by the Siamese network to extract the target region of interest in the next frame. This specifically includes the following sub-steps:
[0061] If the first identifier is detected, the pixel coordinates of the four corner points of the first identifier are substituted into the update function to obtain the updated information of the template image; if only the second identifier is detected, the second identifier is adjusted according to a certain ratio coefficient. k The image is enlarged, and the pixel coordinates of the four corner points of the enlarged second identifier are used as the pixel coordinates of the four corner points of the first identifier in the update function to obtain the updated information of the template image. If the four corner points of the enlarged second identifier exceed the image boundary, the pixel coordinates of the four corner points of the image boundary are used as the pixel coordinates of the four corner points of the first identifier. The expression of the update function is as follows:
[0062] ;
[0063] in, This indicates the coordinates of the center point of the template image in the search image. and These represent the width and height of the template image, respectively. , , and The pixel coordinates of the top left corner, top right corner, bottom left corner, and bottom right corner of the first identifier are represented sequentially. Indicates the correction factor, with a default value of 0. b =20.
[0064] Please refer to Figure 5 ,in Figure 5 (a) in the image is the original image of a single frame of the video sequence. Figure 5 (b) in the image shows the detection result for the identifier code without an update strategy. Figure 5 (c) in the image shows the identification code detection result with the update strategy set. It can be seen that the ROI region outlined by the updated template image fits the identification code better.
[0065] To verify the feasibility of the method of the present invention, a water tank test was conducted based on the above method. The results of the water tank test are as follows: Figure 6 and Figure 7 As shown, where Figure 7 (a), (b), and (c) in the figure correspond to the target in the ROV body coordinate system, respectively. x , y , z Position parameters. Experimental results show that the present invention can effectively identify and accurately locate underwater targets.
[0066] The above descriptions are merely preferred embodiments of this application, and the present invention is not limited to the above embodiments. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the spirit and concept of the present invention should be considered to be included within the protection scope of the present invention.
Claims
1. A method for underwater target detection and pose estimation, characterized in that, The method includes: The nested Apriltag is designed as an underwater cooperative target, wherein the nested Apriltag contains a second identifier nested within a first identifier; The acquired video sequence images are input into the ROI extraction network to capture the target region in the video sequence images; Target detection is performed on the captured area to obtain the pixel coordinates of the four corner points of the first and / or second identifiers. Combined with the actual physical size of the first and / or second identifiers, the pose parameters of the target relative to the ROV are solved using PnP. If only the first or second identifier is detected, the obtained pose parameters are output directly; if both the first and second identifiers are detected at the same time, the result of fusing the pose parameters of the target relative to the ROV obtained by using the two identifiers is output as the pose parameters. The method for solving the pose parameters of the target relative to the ROV using PnP, based on the pixel coordinates of the four corner points of any identifier and the actual physical size of the identifier, includes: Based on the pixel coordinates of the four corner points of the first or second identifier in the camera coordinate system and their coordinates in the target coordinate system, solve for the rotation matrix of the camera coordinate system relative to the target coordinate system. Translation matrix Mark four known points on different lines in the ROV body coordinate system. Combine this with the pixel coordinates of the four known points captured by the camera in the camera coordinate system, and solve for the rotation matrix of the camera coordinate system relative to the ROV body coordinate system. Translation matrix The target's pose parameters relative to the ROV are calculated based on the following formula: ; in, This indicates the spatial position of the target in the target coordinate system. This represents the spatial position parameter of the target in the ROV body coordinate system. This represents the target's attitude parameters in the ROV body coordinate system; the target coordinate system is a coordinate system established with the center point of the identification code as the origin; The method for fusing the pose parameters of the target relative to the ROV obtained by using two different identification codes includes: The target's pose parameters relative to the ROV, obtained from the two identification codes respectively, are substituted into the designed distance function for fusion. The expression of the distance function is as follows: ; in, This indicates the pose parameters obtained based on the second identifier. This represents the pose parameters obtained based on the first identifier. This indicates the target obtained based on the second identifier in the ROV body coordinate system. z To distance parameter, This indicates the target obtained based on the first identifier in the ROV body coordinate system. z Distance parameter.
2. The underwater target detection and pose estimation method according to claim 1, characterized in that, Methods for designing nested Apriltags include: The first and second identifiers are constructed by alternating dark and light colors, and the first and second identifiers have different IDs; Based on the farthest distance the camera can identify and the influence of the underwater environment on the camera, the optimal physical size of the first identifier is determined. When the camera and the target are close to each other, the second identifier can still be found within the camera's field of view as its optimal physical size. The second identifier is placed in the light-colored area in the middle of the first identifier to form a nested Apriltag.
3. The underwater target detection and pose estimation method according to claim 2, characterized in that, The dark color is blue, and the light color is white, to improve the camera's effective recognition distance underwater.
4. The underwater target detection and pose estimation method according to claim 1, characterized in that, The method for performing target detection on the captured area and obtaining the pixel coordinates of the four corner points of the first and / or second identifier includes: Line segment detection and quadrilateral detection are performed sequentially on the captured area. Each quadrilateral is rotated three times in the same direction. The quadrilaterals at the four angles are decoded, and the decoded values are matched with the encoding library. If a match is successful, the quadrilateral is identified as an identifier, and the pixel coordinates of the four corner points of the identifier in the camera coordinate system are obtained. Otherwise, the quadrilateral is identified as a non-identifier. The encoding library stores the code values and IDs of all identifiers.
5. The underwater target detection and pose estimation method according to claim 4, characterized in that, Methods for decoding a quadrilateral at any angle include: The quadrilateral is divided into n×n squares. For each square, if its pixel value is greater than a certain threshold, the coordinates of the interior point of the square are assigned a value of 1, otherwise they are assigned a value of 0. The coordinates of the interior points of the squares located in the effective area are extracted and concatenated into a binary code as the decoding value of the quadrilateral. The code value of the identifier stored in the encoding library is also in binary code form.
6. The underwater target detection and pose estimation method according to any one of claims 1-5, characterized in that, The ROI extraction network is implemented based on a Siamese network. The video sequence images and the template image serve as inputs to the Siamese network, which uses the template image to capture the region of the target in the video sequence images.
7. The underwater target detection and pose estimation method according to claim 6, characterized in that, The method further includes updating the template image used by the Siamese network to extract the target region of interest in the next frame based on the pixel coordinates of the four corner points of the first and / or second identifier.
8. The underwater target detection and pose estimation method according to claim 7, characterized in that, The method for updating the template image used by the Siamese network to extract the target region of interest in the next frame based on the pixel coordinates of the four corner points of the first and / or second identifier includes: If the first identifier is detected, the pixel coordinates of the four corner points of the first identifier are substituted into the update function to obtain the update information of the template image; if only the second identifier is detected, the second identifier is adjusted according to a certain ratio coefficient. k The image is magnified, and the pixel coordinates of the four corner points of the magnified second identifier are used as the pixel coordinates of the four corner points of the first identifier and substituted into the update function to obtain the update information of the template image; if the four corner points of the magnified second identifier exceed the image boundary, then the pixel coordinates of the four corner points of the image boundary are used as the pixel coordinates of the four corner points of the first identifier; wherein, the expression of the update function is: ; in, This represents the coordinates of the center point of the template image within the video sequence. and These represent the width and height of the template image, respectively. , , and The pixel coordinates of the top left corner, top right corner, bottom left corner, and bottom right corner of the first identifier are represented sequentially. This represents the correction factor.
Citation Information
Patent Citations
Underwater robot pose control method based on OpenMV
CN112214028A
Detection and identification method and device of underwater robot and computer storage medium
CN112862865A