Method for Estimating Relative Distance and Azimuth of Vision-Based Underwater Bionic Manta Ray Robot Fish

Through deep convolutional neural network and binocular distance measurement technology, the problem of underwater bionic manta ray robotic fish being difficult to accurately estimate relative distance and orientation during cluster formation is solved, and real-time and accurate relative distance and orientation estimation is achieved.

CN115170654BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210595435.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-29
Publication Date
2025-06-27
Estimated Expiration
2042-05-29

AI Technical Summary

Technical Problem

When underwater bionic manta ray robots are formed in clusters, due to their flexible maneuver, complex movement, large posture changes and few fixed points, it is difficult to accurately estimate the relative distance and orientation through traditional methods.

Method used

Deep convolutional neural network is used for object detection, combined with binocular distance measurement, and the relative distance and orientation information of underwater bionic manta ray robot fish is obtained in real time. Specific steps include image dataset annotation and training, binocular camera calibration and correction, parallax map acquisition and processing, prediction box coordinate calculation and relative distance estimation.

Benefits of technology

The accurate estimation of the relative distance and orientation between underwater bionic manta ray robots has been achieved, and the problem of difficulty in ensuring the accuracy of traditional methods under complex movements and posture changes is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170654B_ABST
    Figure CN115170654B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for estimating the relative distance and azimuth of an underwater biomimetic manta ray robot fish based on vision. The target detection network is trained by obtaining a data set. After the training is completed, the network can be used to obtain the azimuth of the underwater biomimetic manta ray robot fish and the position of the prediction box in the pixel coordinate system. Then, the semi-global binocular matching method is used to obtain the binocular disparity map, and the relative distance estimation of the biomimetic manta ray robot fish can be obtained by combining the camera calibration information and the disparity pixel search method. It can effectively overcome the problems that the single-body motion state of the underwater biomimetic manta ray robot fish is complex, the attitude change range is large, the fixed points are few, and it is difficult to obtain the relative distance and azimuth estimation by using the fixed-point indicator method similar to traditional vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and navigation positioning, and in particular, to a method for estimating the relative distance and azimuth of an underwater biomimetic manta ray robot fish based on vision. Background Technique

[0002] With the development of related fields and technologies such as bioengineering, fluid mechanics, materials science, and computer science, the pectoral fin propulsion method, which is efficient, low-noise, and flexible under hundreds of millions of years of natural evolution of fish, has emerged in the research field of underwater vehicles. As a kind of fish with pectoral fin propulsion, the research and development of underwater biomimetic manta ray robot fish by drawing on the pectoral fin propulsion characteristics of manta rays has also made certain progress.

[0003] When traditional underwater vehicles achieve cluster formation, indicator lights can be fixed on the vehicle carrier to calculate the relative distance and azimuth. However, underwater biomimetic manta ray robot fish are flexible, have complex movements, large attitude changes, and few fixed points. When achieving cluster formation, if the traditional vehicle fixed-point indicator lights are still used, it will be difficult to obtain accurate relative distance and azimuth information, thus affecting the mutual positioning between formation members and the task execution of the formation. Summary of the Invention

[0004] Technical Problems to be Solved

[0005] In order to avoid the deficiencies of the prior art and solve the problem of difficult estimation of the relative distance and azimuth between current underwater biomimetic manta ray robot fish. The present invention provides a method for estimating the relative distance and azimuth of an underwater biomimetic manta ray robot fish based on vision, which directly detects the underwater biomimetic manta ray robot fish by using a deep convolutional neural network method to obtain the azimuth information relative to its neighboring robot fish, and then combines binocular distance measurement to obtain accurate and stable distance and azimuth information of the underwater biomimetic manta ray robot fish in real time.

[0006] Technical Solution

[0007] A method for estimating the relative distance and azimuth of an underwater biomimetic manta ray robot fish based on vision, characterized by the following steps:

[0008] S1: Obtain the motion image dataset of the underwater biomimetic manta ray robot fish, and perform category annotation and real box calibration;

[0009] S2: Send the image and the label file containing category annotation and real box calibration into the target detection neural network for training and verification at the same time; the target detection neural network outputs the azimuth estimation, the confidence of the azimuth prediction, and the upper left coordinate and the lower right coordinate of the prediction box in the pixel coordinate system;

[0010] S3: Use the binocular calibration method to obtain the internal parameter matrix, distortion parameters of the left camera and the right camera, the rotation matrix and translation vector of the right camera relative to the left camera;

[0011] S4: Correct the radial distortion and tangential distortion of the binocular camera;

[0012] S5: Obtain the disparity map through the binocular stereo matching method;

[0013] S6: Input the left camera image in the binocular camera after correction in S4 into the object detection neural network trained in S2, and obtain the upper left coordinate, lower right coordinate of the predicted box of the underwater biomimetic manta ray robot fish in the pixel coordinate system, the azimuth estimation of the camera located at the underwater biomimetic manta ray robot fish, and the confidence of the azimuth prediction;

[0014] S7: Calculate the center point coordinate and pixel area of the predicted box from the upper left coordinate and lower right coordinate of the predicted box obtained in S6;

[0015] S8: Calculate the proportion of the pixel area of the predicted box in the total image area, and record it as the search range parameter;

[0016] S9: Calculate the disparity search pixel range from the search range parameter, and extract all the disparity values corresponding to each pixel point in the disparity map obtained in S5 according to the disparity search pixel range;

[0017] S10: Use the camera calibration data obtained in S3 and the disparity values obtained in S9 to calculate the relative distance values corresponding to each pixel point, and take the minimum value of all the relative distances within the search range, which is the relative distance of the underwater biomimetic manta ray robot fish. Finally, combine the azimuth estimation obtained in S6 to estimate the relative distance and azimuth of the underwater biomimetic manta ray robot fish.

[0018] Further technical solution of the present invention: The category annotation and true box calibration in S1: Divide by the azimuth of the camera located at the robot fish, and label eight azimuth categories C on the horizontal plane, where: C0 is the front left side, C1 is the rear left side, C2 is the front right side, C3 is the rear right side, C4 is the front side, C5 is the rear side, C6 is the left side, and C7 is the right side; The true box calibration content is the upper left coordinate (x left , y left ) and the lower right coordinate (x right , y right ) of the biomimetic manta ray robot fish in the pixel coordinate system.

[0019] Further technical solution of the present invention: The target detection neural network described in S2 includes a backbone feature extraction part, a neck feature aggregation part, and a detection head part; the backbone feature extraction part is composed of Module 1 and Module 2. Module 1 is composed of three convolutional layers with a convolutional kernel size of 3x3, and is used to adjust the dimension and size of the input original image; Module 2 is composed of 22 residual connection blocks connected in segments in a ratio of 1:2:7:8:4, and is used to extract features; each residual connection block is composed of a convolutional layer with a convolutional kernel size of 3x3 and a 1x1 convolutional layer used as a residual connection to ensure that the input dimension and output dimension are consistent; convolutional layers with a convolutional kernel size of 3x3 for secondary adjustment of dimension and image size are interspersed at the segment connection points. Denote the five segments as bloCk1~block5. Input the image processed by Module 1, and the corresponding output results of each segment are out1~out5; the neck feature aggregation part is composed of three upsampling units and three aggregation units, and performs upsampling and aggregation operations on the outputs out2~out5 of the backbone feature extraction part, that is, four feature maps; the detection head part is composed of three classification detection heads and three regression detection heads. Each classification detection head and each regression detection head are each composed of three convolutional layers with a convolutional kernel size of 3x3.

[0020] Further technical solution of the present invention: The internal parameter matrix K in S3:

[0021]

[0022] where, f is the focal length; d x is the width in the x direction of the pixel; d y is the width in the y direction of the pixel; f / d x uses pixels to describe the length of the focal length in the x-axis direction; f / d y uses pixels to describe the length of the focal length in the y-axis direction; u0, v0 are the actual positions of the principal points;

[0023] Distortion parameter D:

[0024] D=(k1, k2, k3, p1, p2)

[0025] where, (k1, k2, k3) are radial distortion parameters, and (p1, p2) are tangential distortion parameters;

[0026] The rotation matrix R and translation vector T of the right camera relative to the left camera:

[0027]

[0028] where, R l , T l are the rotation matrix and translation vector of Camera 1 relative to the calibration object obtained through single-object calibration, R r, T r is the rotation matrix and translation vector of camera 2 relative to the calibration object obtained through single calibration.

[0029] A further technical solution of the present invention: The calculation formula of S7 is as follows:

[0030]

[0031] where (x pl , y pl ) is the upper left corner coordinate of the prediction box, (x pr , y pr ) is the lower right corner coordinate, (x c , y c ) is the center point coordinate of the prediction box, and S is the pixel area of the prediction box.

[0032] A further technical solution of the present invention: The calculation formula of S9 is as follows:

[0033]

[0034] where λ is the search range parameter, (x c , y c ) is the center point of the prediction box, and (x s , y s ) is the disparity search pixel range.

[0035] A further technical solution of the present invention: The calculation formula of S10 is as follows:

[0036]

[0037]

[0038] target = min(Distance)

[0039] where Distance is the relative distance value corresponding to each pixel point, target is the minimum value of all relative distances within the search range, u and v represent the coordinates of a point in the pixel coordinate system, that is, the x s and y s obtained in step 8, u0 and v0 are the coordinate values of the origin of the left camera image plane in the pixel coordinate system, and T x is the center distance between the optical centers of the two cameras, that is, the modulus of the translation vector T.

[0040] Beneficial effects

[0041] A method for estimating the relative distance and orientation of an underwater biomimetic manta ray robot fish based on vision. The method trains a target detection network by obtaining a data set. After the training is completed, the network can be used to obtain the orientation of the underwater biomimetic manta ray robot fish and the position of the prediction box in the pixel coordinate system. Then, a semi-global binocular matching method is used to obtain a binocular disparity map, and the relative distance estimation of the biomimetic manta ray robot fish can be obtained by combining the camera calibration information and the disparity pixel search method.

[0042] Experimental verification shows that when the relative distance and orientation of underwater biomimetic manta ray robot fish are estimated according to the method provided by the present invention, it can effectively overcome the problems that the single motion state of underwater biomimetic manta ray robot fish is complex, the attitude change range is large, the fixed points are few, and it is difficult to obtain the relative distance and orientation estimation by using the fixed-point indicator method similar to traditional vehicles. Description of the Drawings

[0043] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0044] Figure 1 The working flowchart of the present invention;

[0045] Figure 2 Prototype diagram of the underwater biomimetic manta ray robot fish;

[0046] Figure 3 Schematic diagram of category annotation and ground truth calibration;

[0047] Figure 4 Schematic diagram of the backbone feature extraction part;

[0048] Figure 5 Schematic diagram of the neck feature aggregation part;

[0049] Figure 6 Schematic diagram of the detection head part;

[0050] Figure 7 PR curve;

[0051] Figure 8 Partial visualization results of orientation estimation;

[0052] Figure 9 Original image and visualization results of orientation estimation when the true relative distance is 2m;

[0053] Figure 10 Disparity map and relative distance estimation results when the true relative distance is 2m. Detailed Description of the Invention

[0054] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0055] The style of the underwater biomimetic manta ray robot fish is as shown in Figure 2 , and the specific working flow chart of the present invention is as shown in Figure 1 , including the following steps:

[0056] Step 1: Use a waterproof camera to collect a dataset of underwater biomimetic manta ray robot fish motion images with 8000 high-quality images, complete angles and reasonable category arrangements, and perform category annotation and true box calibration. Divided by the orientation of the camera relative to the robot fish, eight orientation categories C are marked on the horizontal plane, where: C0 is the front left side, C1 is the rear left side, C2 is the front right side, C3 is the rear right side, C4 is the front side, C5 is the rear side, C6 is the left side, and C7 is the right side. The content of the true box calibration is the upper left corner coordinates (x left , y left ) and the lower right corner coordinates (x right , y right ) of the biomimetic manta ray robot fish in the pixel coordinate system. The category annotation and true box calibration are detailed in Figure 3 . Then, the dataset is randomly shuffled and divided, with 80% used for network model training, 10% used for model verification, and 10% used for the final model testing.

[0057] Step 2: Send the images and the label files containing category annotations and true box calibrations into the target detection neural network for training and verification at the same time. The structure of the target detection neural network can be divided into three parts, namely the backbone feature extraction part, the neck feature aggregation part, and the detection head part.

[0058] Among them, the backbone feature extraction part can be refined into two sub-modules. Module one consists of three convolutional layers with a convolutional kernel size of 3x3, which is used to adjust the dimension and size of the input original image; Module two consists of 22 residual connection blocks connected in segments in a ratio of 1:2:7:8:4, which is used to extract features. Among them, each residual connection block consists of a convolutional layer with a convolutional kernel size of 3x3 and a 1x1 convolutional layer used as a residual connection to ensure that the input dimension and output dimension are consistent. Convolutional layers with a convolutional kernel size of 3x3 for secondary adjustment of dimension and image size are interspersed at the segment connection points. Denote the five segments as block1~block5, input the image processed by module one, and the corresponding output results of each segment are out1~out5. The specific process is as shown in Figure 4 .

[0059] The neck feature aggregation part consists of three upsampling units and three aggregation units, which perform upsampling and aggregation operations on the outputs out2 to out5 of the backbone feature extraction part, namely four feature maps, as Figure 5 shown. Finally, three feature maps of sizes final1 to final3 are output, which are respectively used to detect targets with different scale sizes.

[0060] The detection head part consists of three classification detection heads and three regression detection heads. Each classification detection head and each regression detection head are each composed of three convolutional layers with a convolutional kernel size of 3x3. The three feature maps of sizes final1 to final3 are copied to obtain (final 11 , final 12 ) to (final 31 , final 32 ) for a total of six feature maps. Then, according to the rule that two identical feature maps of each size respectively correspond to a classification detection head and a regression detection head, all feature maps are classified and regressed, corresponding to obtaining classification category outputs, bounding box regression errors, and predicted center point offset error outputs. After the output results are solved, the azimuth estimation C i , the confidence p of the azimuth prediction, and the upper left coordinates (x pl , y pl ) and lower right coordinates (x pr , y pr ) of the prediction box in the pixel coordinate system can be obtained. The specific process is as Figure 6 shown.

[0061] After setting the detection network structure, the training parameters are set. The training process is set to 300 rounds of loops. During the process, mini-batch stochastic gradient descent is used as the optimizer; the peak learning rate is set to 0.01, the momentum is set to 0.9, the weight decay coefficient is 5e-4, the initial learning rate is 1e-4, and it rises to the peak after 10 loops and then enters the cosine decay stage, and finally drops to 5e-4. Data augmentation is not used in the last 15 loops to improve the model accuracy. In terms of data augmentation, the method of random flipping is not adopted, but a scheme of random color perturbation is adopted, and random adjustments are made in the HSV color space. After the target detection network is trained, the parameters of each neuron in the network are fixed and tested on the test set, and the various indicators of the model are tested respectively. After determining convergence, the next step is carried out.

[0062] Step 3: Use the binocular calibration method to obtain the internal parameter matrix K of camera 1 (left camera) and camera 2 (camera):

[0063]

[0064] where f is the focal length in millimeters; d x is the width of a pixel in the x direction in millimeters; d y is the width of a pixel in the y direction in millimeters; f / d x uses pixels to describe the length of the focal length in the x-axis direction; f / d y uses pixels to describe the length of the focal length in the y-axis direction; u0, v0 are the actual positions of the principal points, also in pixels.

[0065] Distortion parameter D:

[0066] D = (k1, k2, k3, p1, p2)

[0067] where (k1, k2, k3) are the radial distortion parameters and (p1, p2) are the tangential distortion parameters;

[0068] The rotation matrix R and translation vector T of camera 2 relative to camera 1:

[0069]

[0070] where R l , T l are the rotation matrix and translation vector of camera 1 relative to the calibration object obtained through single-object calibration, R r , T r are the rotation matrix and translation vector of camera 2 relative to the calibration object obtained through single-object calibration. By performing single-object calibration on camera 1 and camera 2 respectively, we can obtain R l , T l , R r , T r . Substituting into the above formula, we can calculate the rotation matrix R and translation matrix T between camera 1 and camera 2.

[0071] Step 4: Perform radial and tangential distortion correction on the binocular camera. The radial distortion correction formula is:

[0072]

[0073] The tangential distortion correction formula is:

[0074]

[0075] where (x r , y r ) are the pixel coordinates after radial distortion, (x t , y t ) are the pixel coordinates after tangential distortion, and (x, y) are the coordinates after undistortion.

[0076] Considering the simultaneous existence of radial distortion and tangential distortion, the undistorted coordinates (x, y) can be obtained by rectification according to the following formula:

[0077]

[0078] Among them, (x0, y0) is the pixel coordinate after superimposing radial distortion and tangential distortion, that is, the pixel coordinate before distortion correction.

[0079] Step 5: Obtain the disparity map using the binocular stereo matching method. Further, the specific operations are as follows:

[0080] Step 5.1: Use the rotation matrix R and translation vector T obtained in Step 3 to perform binocular stereo rectification on the image processed in Step 4.

[0081] Step 5.2: Preprocess the binocular rectified image, and use the Sobel operator to filter to obtain the gradient information of the image.

[0082] Step 5.3: Calculate the binocular matching cost for each pixel point within a pixel window of a specified size for the gradient information obtained from the preprocessing and the original image.

[0083] Step 5.4: Establish a global Markov energy equation in the image according to the idea of dynamic programming, and add the matching costs in each direction of a single pixel to obtain the total corresponding pixel matching cost.

[0084] Step 5.5: Select the disparity with the smallest total matching cost within the disparity search range as the disparity value of the corresponding pixel point, and visualize the disparity value of each pixel to obtain the disparity map before post-processing.

[0085] Step 5.6: Perform post-processing, including uniqueness detection, sub-pixel interpolation, left-right consistency detection, and connected region detection operations, to obtain the final output binocular disparity map.

[0086] Step 6: Input the image of Camera 1 (left camera) in the rectified binocular camera (i.e., the image obtained in Step 4) into the detection network trained in Step 2 to obtain the upper-left corner coordinates (x pl , y pl ) and lower-right corner coordinates (x pr , y pr ) of the predicted box of the underwater biomimetic manta ray robot fish in the pixel coordinate system, the azimuth estimation C i of the camera relative to the underwater biomimetic manta ray robot fish, and the confidence p of the azimuth prediction.

[0087] Step 7: The upper-left corner coordinates (x pl , y pl ) and lower-right corner coordinates (x pr , ypr ) The center coordinates (x c , y c ) of the predicted bounding box and the pixel area S of the predicted bounding box are obtained after calculation by the following formula.

[0088]

[0089] Step 8: Calculate the ratio of the area S of the predicted bounding box to the total area S max of the image, denoted as the search range parameter λ ∈ (0, 1).

[0090] Step 9: Obtain the disparity search pixel range (x c , y c ) centered on the center point (x s , y s ) of the predicted bounding box, where After that, according to this range, all the disparity values d of each pixel point corresponding in the disparity map obtained in Step 5 are taken out.

[0091] Step 10: Use the camera calibration data obtained in Step 3 and the disparity value d obtained in Step 9 to calculate the relative distance value corresponding to each pixel point according to the following formula, denoted as Distance. Take the minimum value of all the relative distances within the search range and denote it as target, which is the relative distance of the underwater biomimetic manta ray robot fish. Finally, combine the azimuth prediction C i (i = 0, 1,..., 7), and the relative distance target and azimuth C i of the underwater biomimetic manta ray robot fish can be estimated.

[0092]

[0093]

[0094] target = min(Distance)

[0095] where u and v represent the coordinates of a point in the pixel coordinate system, which are the x s and y s obtained in Step 8, u0 and υ0 are the coordinate values of the origin of the left camera image plane in the pixel coordinate system, and T x is the center distance between the optical centers of the two cameras, that is, the modulus of the translation vector T.

[0096] Experimental verification

[0097] Here, the method proposed in the present invention will be experimentally verified from two aspects of azimuth estimation and relative position estimation respectively.

[0098] I. Azimuth estimation verification

[0099] Since the azimuth estimation is directly output after the training of the target detection network model, the PR curve characteristics and mAP in the model testing stage are used for evaluation. The specific meanings of the indicators are shown in Table 1.

[0100] Table 1 Azimuth Estimation Evaluation Index

[0101]

[0102] According to the above index settings, after the network training is completed, it is tested on the previously set test set. The average estimated precision rate for each azimuth category reaches 0.714, and the PR curve is as Figure 7 ., indicating that the method of the present invention has better advantages. Specifically, by visualizing the estimation results, it can be seen that when performing azimuth estimation, in the case of relatively dim underwater environment and relatively far target distance, the azimuth estimation confidence of the underwater biomimetic manta ray robot fish can reach above 0.5, and the azimuth estimation result is consistent with the actual relative azimuth result of the camera relative to the underwater biomimetic manta ray robot fish, and it can output the relative azimuth result in the dim underwater environment; when the target is relatively close and the target is relatively clear, the estimation confidence reaches 0.9, and the estimation effect is excellent. Therefore, the present invention can accurately estimate the azimuth result of the underwater biomimetic manta ray robot fish. Part of the azimuth estimation visualization is as Figure 8 shown.

[0103] II. Relative Distance Estimation Verification

[0104] Taking two manta ray prototypes as examples, a relative distance estimation experiment is carried out. The position of the rear manta ray prototype is set as the starting point, that is, at 0 m. The front manta ray prototype is measured at positions with relative actual distances of 1 m, 1.5 m, 2 m, 2.5 m, 3 m, 3.5 m, 4 m, 4.5 m, and 5 m from the rear prototype respectively. At each position, 10 relative distance measurements are carried out respectively, and the single relative distance estimation error and the average relative distance estimation error are calculated. Since the relative distance estimation method of the present invention depends on the azimuth estimation output, two new indicators are set here, namely the correct detection rate and the global correct detection rate, to refine the relative distance estimation effect. The specific meanings of the above four indicators are shown in Table 2.

[0105] Table 2 Relative Distance Estimation Index

[0106]

[0107]

[0108] Table 3 Relative Distance Estimation Verification Results

[0109]

[0110]

[0111] As can be seen from Table 3, when the actual relative distance between the underwater biomimetic manta ray robot fish is relatively close, at 1-3 m, the relative distance estimation accuracy is relatively high, and the maximum error occurs at 3 m. The others are all below 0.41 m and increase as the distance shortens. At the same time, the azimuth estimation result is relatively good, not less than 90%. When the relative distance is far, both the relative distance estimation accuracy and the azimuth estimation result decrease. When the relative distance is within 5 m, effective estimation can be carried out, further verifying the effectiveness of the present invention. When the actual relative distance is 2 m, the visualization result of the azimuth estimation and the relative distance estimation effect are respectively as Figure 9 and as Figure 10 shown.

[0112] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A method for estimating the relative distance and orientation of a vision-based underwater biomimetic manta ray robot fish, characterized in that The steps are as follows: S1: Obtain the motion image dataset of the underwater biomimetic manta ray robot fish, and perform class annotation and ground truth box calibration; S2: Send the images and the label files containing class annotation and ground truth box calibration into the target detection neural network for training and validation at the same time; the target detection neural network outputs the azimuth estimation, the confidence of azimuth prediction, and the upper left corner coordinates and lower right corner coordinates of the prediction box in the pixel coordinate system; S3: Use the binocular calibration method to obtain the internal parameter matrix, distortion parameters of the left camera and the right camera, the rotation matrix and translation vector of the right camera relative to the left camera; S4: Correct the radial distortion and tangential distortion of the binocular cameras; S5: Obtain the disparity map through the binocular stereo matching method; S6: Input the left camera image in the binocular cameras after correction in S4 into the trained target detection neural network in S2 to obtain the upper left corner coordinates, lower right corner coordinates of the prediction box of the underwater biomimetic manta ray robot fish in the pixel coordinate system, the azimuth estimation of the camera located at the underwater biomimetic manta ray robot fish, and the confidence of azimuth prediction; S7: Calculate the center point coordinates of the prediction box and the pixel area of the prediction box from the upper left corner coordinates and lower right corner coordinates of the prediction box obtained in S6; S8: Calculate the proportion of the pixel area of the prediction box in the total image area, and record it as the search range parameter; S9: Calculate the disparity search pixel range from the search range parameter, and extract all the disparity values corresponding to each pixel point in the disparity map obtained in S5 according to the disparity search pixel range; S10: Use the camera calibration data obtained in S3 and the disparity values obtained in S9 to calculate the relative distance values corresponding to each pixel point, and take the minimum value of all the relative distances within the search range, which is the relative distance of the underwater biomimetic manta ray robot fish. Finally, by integrating the azimuth estimation obtained in S6, the relative distance and azimuth of the underwater biomimetic manta ray robot fish can be estimated.

2. The method for estimating the relative distance and azimuth of the vision-based underwater biomimetic manta ray robot fish according to claim 1, wherein: Category annotation and ground truth bounding box calibration described in S1: Divided by the orientation of the camera relative to the robotic fish, eight orientation categories C are annotated on the horizontal plane, where: C0 is the front left side, C1 is the back left side, C2 is the front right side, C3 is the back right side, C4 is the front side, C5 is the back side, C6 is the left side, and C7 is the right side; the ground truth bounding box calibration content is the upper left corner coordinates (x left , y left ) and the lower right corner coordinates (x right , y right ) of the biomimetic manta ray robotic fish in the pixel coordinate system.

3. The method for estimating the relative distance and azimuth of the vision-based underwater biomimetic manta ray robot fish according to claim 2, wherein: The object detection neural network described in S2 includes a backbone feature extraction part, a neck feature aggregation part, and a detection head part; the backbone feature extraction part consists of Module 1 and Module 2. Module 1 consists of three convolutional layers with a convolutional kernel size of 3x3 and is used to adjust the dimension and size of the input original image; Module 2 consists of 22 residual connection blocks connected in segments at a ratio of 1:2:7:8:4 and is used to extract features; each residual connection block consists of a convolutional layer with a convolutional kernel size of 3x3 and a 1x1 convolutional layer serving as a residual connection and used to ensure that the input dimension and output dimension are consistent; convolutional layers with a convolutional kernel size of 3x3 for secondary adjustment of dimension and image size are interspersed at the segment connection points. Denote the five segments as bloCk1 to block5. Input the image processed by Module 1, and the corresponding output results of each segment are out1 to out5; the neck feature aggregation part consists of three upsampling units and three pooling units, and performs upsampling and pooling operations on the outputs out2 to out5 of the backbone feature extraction part, that is, four feature maps; the detection head part consists of three classification detection heads and three regression detection heads, where each classification detection head and each regression detection head both consist of three convolutional layers with a convolutional kernel size of 3x3.

4. The method for estimating the relative distance and azimuth of the vision-based underwater biomimetic manta ray robot fish according to claim 3, wherein: The internal parameter matrix K in S3: Among them, f is the focal length; d x Width of the pixel in the x direction; d y Width of the pixel in the y direction; f / d x Using pixels to describe the length of the focal length in the x-axis direction; f / d y Using pixels to describe the length of the focal length in the y-axis direction; u0, v0 are the actual positions of the principal points; distortion parameter D: D=(k1, k2, k3, p1, p2) where (k1, k2, k3) are radial distortion parameters and (p1, p2) are tangential distortion parameters; The rotation matrix R and translation vector T of the right camera relative to the left camera: where, R l , T l are the rotation matrix and translation vector of Camera 1 relative to the calibration object obtained through single-object calibration, and R r , T r are the rotation matrix and translation vector of Camera 2 relative to the calibration object obtained through single-object calibration.

5. The method for estimating the relative distance and azimuth of the vision-based underwater biomimetic manta ray robot fish according to claim 4, wherein: The calculation formula of S7 is as follows: Among them, (x pl , y pl ) is the upper-left coordinate of the prediction box, (x pr , y pr ) is the lower-right coordinate, (x c , y c ) is the center point coordinate of the prediction box, and S is the pixel area of the prediction box.

6. The method for estimating the relative distance and azimuth of the vision-based underwater biomimetic manta ray robot fish according to claim 5, wherein: The calculation formula of S9 is as follows: Among them, λ is the search range parameter, (x c , y c ) is the center point of the prediction box, and (x s , y s ) is the pixel range for disparity search.

7. The method for estimating the relative distance and azimuth of the vision-based underwater biomimetic manta ray robot fish according to claim 6, characterized in that: The calculation formula of S10 is as follows: target = min(Distance) Among them, Distance is the relative distance value corresponding to each pixel point, target is the minimum value of all relative distances within the search range, u and v represent the coordinates of a point in the pixel coordinate system, which are the x obtained in step 8 s and y s , u0 and υ0 are the coordinate values of the origin of the left camera image plane in the pixel coordinate system, and T x is the center distance between the optical centers of the two cameras, that is, the modulus of the translation vector T.

Citation Information

Patent Citations

  • Underwater biomimetic robot fish binocular stereoscopic ranging method

    CN109798877A

  • Hovering control method, system and device of bionic underwater robot of visual servo

    CN110488847A