Underwater binocular positioning method and device based on target assistance and storage medium
By using binocular cameras and advanced image processing algorithms in underwater environments, the problem of insufficient accuracy of existing underwater positioning methods in complex environments is solved, and high-precision target positioning is achieved.
Patent Information
- Application Number
- CN202510106998.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The existing underwater positioning methods have shortcomings in accuracy and stability, especially in complex underwater environments, which are difficult to achieve high-precision positioning.
The target-assisted underwater binocular positioning method is adopted to obtain internal and external parameters through the binocular camera, establish a refractive model of the underwater binocular camera, and optimize the refractive parameters using genetic algorithms, combine the YOLOv8 model and the ORB algorithm to detect and match the target's three-dimensional coordinates to calculate the target.
High-precision target positioning in complex underwater environments is achieved, the shortcomings of monocular vision systems in depth calculations and complex environments are overcome, and the accuracy and efficiency of underwater positioning are significantly improved.
Smart Images

Figure CN119992038A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an underwater binocular positioning method, a device and a storage medium based on target assistance, and belongs to the technical field of underwater environment perception. Background Art
[0002] Traditional underwater positioning methods, such as sonar positioning and GPS positioning, have problems such as low accuracy, high cost or limited applicable conditions. With the increasing intelligence and high-precision requirements of underwater operations, there is a need for an automatic and accurate positioning method and system for underwater targets. In recent years, vision-based positioning methods have attracted widespread attention due to their advantages such as high precision and low cost. In the prior art, monocular vision systems face many problems in underwater environments, especially in feature extraction and depth calculation, and it is difficult to meet the needs of high-precision positioning. In addition, the existing underwater visual positioning methods still face defects such as insufficient positioning accuracy and poor stability when dealing with complex underwater environments, such as changes in light and refractive index. Summary of the invention
[0003] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide an underwater binocular positioning method, device and storage medium based on target assistance, which can realize accurate positioning of underwater targets through binocular vision method and meet the working requirements in complex underwater environments.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0005] In a first aspect, the present invention provides an underwater binocular positioning method based on target assistance, comprising the following steps:
[0006] Obtaining the internal and external parameters of the binocular camera, which are obtained by calibrating the binocular camera in air;
[0007] Obtain a calibration evaluation function, which is constructed based on the underwater binocular camera refraction model, and then use a genetic algorithm to optimize the calibration evaluation function to obtain the refraction parameters of the binocular camera;
[0008] Controlling the binocular camera to shoot and collect the target image to obtain an initial underwater cooperative positioning target image, and using the internal and external parameters and refraction parameters of the binocular camera to correct the initial underwater cooperative positioning target image to obtain a corrected underwater cooperative positioning target image;
[0009] The YOLOv8 model is used to detect the target in the corrected underwater cooperative positioning target image to obtain the target area in the underwater cooperative positioning target image;
[0010] The ORB algorithm is used to extract and match feature points in the target area to obtain matching feature points;
[0011] The three-dimensional coordinates of the target in the camera coordinate system are calculated using the matching feature points.
[0012] The method for calibrating the internal and external parameters of the binocular camera includes:
[0013] Using Zhang's calibration method to solve the camera intrinsic parameter matrix With the external parameter matrix The product of , and then solve the internal parameter matrix and the external parameter matrix ,
[0014] ;
[0015] in, is the extrinsic rotation matrix, is the extrinsic translation matrix, is the scaling factor, and is the focal length The normalized coordinates of is the principal point coordinate of the camera, that is, the intersection of the camera optical axis and the imaging plane, is the coordinate in the pixel coordinate system, is the coordinate in the world coordinate system;
[0016] Based on Zhang's calibration method, the tangential distortion is described as:
[0017] ;
[0018] in, is the actual observed image coordinate, is the image coordinate without tangential distortion, is the tangential distortion coefficient of the camera, is the square distance from the point to the center of the image;
[0019] The three radial distortion coefficients and two tangential distortion coefficients are solved by the least squares method, and combined with the internal parameter matrix , extrinsic rotation matrix and the extrinsic translation matrix The reprojection error function is constructed, and then the maximum likelihood estimation method is used for nonlinear optimization to obtain the optimized internal and external parameters of the binocular camera. The reprojection error function is shown in the following formula:
[0020] ;
[0021] Among them, F represents the error size, n represents the number of calibration plate images obtained in different directions in the air, and m represents the total number of corner points on each calibration plate image. Indicates On the image The coordinates of the corner points, Representing 3D points In the The reprojection result on the image is is the camera distortion coefficient, is the radial distortion coefficient of the camera, is the tangential distortion coefficient of the camera, For the The extrinsic rotation matrix and extrinsic translation matrix of the image, is the three-dimensional coordinate in the world coordinate system, indicating the The true 3D position of the corner points is achieved by minimizing the reprojection error function F. Optimization of six parameters.
[0022] The calibration evaluation function construction method includes:
[0023] The first refraction model and the second refraction model that occur when light passes through the glass surface are established. The expression of the first refraction model is as follows:
[0024]
[0025] The expression of the double refraction model is as follows:
[0026] ;
[0027] in, , , Represent the relative refractive index of air, glass, and water, respectively. is the thickness of the glass cover, Indicates the distance from the optical center to the glass cover; Indicates point To the camera coordinate system The vertical distance on the axis, Indicates point The distance to the outer surface of the glass cover, Indicates point Projecting point on the outer surface of the glass cover To the camera coordinate system The vertical distance on the axis, Indicates point Refraction point on the inner surface of the glass cover To the camera coordinate system The vertical distance on the axis;
[0028] Combining the first refraction model and the second refraction model, we get The only solution is to obtain the refraction surface of the glass cover point, and then use the air camera model to find the projection point on the imaging surface , and finally establish the space point Projection point to the camera imaging surface the relationship between;
[0029] According to the underwater binocular camera refraction model, the spatial object point Projection point to the camera imaging surface The relationship between the two is transformed to obtain the binocular model and the projection point of the camera imaging surface is obtained and the actual image point on the camera imaging surface The relationship between them is established by minimizing the difference between them:
[0030] ;
[0031] Among them, E represents the value of the calibration evaluation function. The smaller the E value is, the closer the actual detected corner point coordinates are to the corner point coordinates calculated using the imaging model. is the total number of corner points extracted from each calibration plate image, , are the actual coordinates of the i-th corner point detected on the left and right camera imaging surfaces, respectively. , are the coordinates of the i-th corner point obtained by projecting the underwater 3D point onto the left and right camera imaging surfaces through model calculation, as well as They respectively represent the normal vector of the glass cover refraction surface, the distance from the optical center to the glass cover, and the thickness of the glass cover in the left and right camera refraction models in the underwater binocular camera refraction model.
[0032] The method of optimizing the calibration evaluation function by using a genetic algorithm to obtain the refractive parameters of the binocular camera includes:
[0033] Set the initial population size to 100 individuals and randomly generate the initial population. Each individual in the population is represented by a ten-dimensional vector express, , where the parameter to be optimized is the vector Each parameter is generated within a defined threshold range, where and They are the thickness of the left and right glass waterproof covers respectively. and are the distances from the optical centers of the left and right cameras to the glass refraction surface, and The normal vectors of the glass cover refraction surface in front of the left and right cameras are The component in the axial direction, The components in the axial direction and in Component in the axial direction;
[0034] Define the fitness function , which accepts each individual in the population As input and return the above calibration evaluation function The negative value of , the optimization goal is to maximize the fitness function, that is, to minimize the calibration evaluation function;
[0035] Using the fitness function, calculate each individual in the population The fitness of the individual is 30%, and 30% of the individuals are selected from the population as the parent individuals according to the roulette wheel selection method. , the remaining 70% of individuals are used as offspring individuals ;
[0036] For the selected parent individuals Pairing and single-point crossover to generate new offspring individuals , that is, each pair of parents exchanges parameters through a randomly selected crossover point, and each crossover generates two offspring individuals If the number of parent individuals is even, a perfect pairing is formed; if it is odd, the last parent individual can be directly copied as an offspring, and the generated offspring individual Directly replace the parent individual ;
[0037] From the generated offspring individuals Select 5% of the individuals for mutation operation, that is, select A certain parameter is randomly adjusted slightly, and the mutated individual replaces the original individual;
[0038] A new population is generated through selection, crossover and mutation operations, that is, all individuals in the original population , now a new individual ,
[0039] Set the termination condition, the maximum number of iterations is 100 or the average fitness value of the individuals in the population is > -0.05. Before any termination condition is met, repeat the above selection, crossover and mutation operations to optimize the population;
[0040] After reaching the termination condition, the individual with the highest fitness is selected as the optimal solution, and the parameter value corresponding to the individual is output.
[0041] The correction processing of the initial underwater cooperative positioning target image using the internal and external parameters and refraction parameters of the binocular camera includes:
[0042] By obtaining the distortion coefficients of the binocular camera, a distortion model is constructed for the cooperative target image. For each pixel (x, y), the corresponding point in the original image is calculated by the distortion model using the reverse mapping method. :
[0043] ;
[0044] in, , is the radial distortion coefficient, is the tangential distortion coefficient;
[0045] Cooperative target image After distortion correction, the cooperative target image Perform epipolar correction so that the corresponding pixels of the left and right images are on the same horizontal line, and finally obtain the corrected underwater target image .
[0046] The YOLOv8 model uses a target detection model , the YOLOv8 model is used to detect the target in the corrected underwater cooperative positioning target image, and the target area in the underwater cooperative positioning target image includes:
[0047] If the target detection model Correcting the underwater target images of the left and right eyes If the target is successfully detected in all, the target detection model is used In the left and right images, the target area is selected by the detection box, where the detection box contains the position information of the box selection, that is, the pixel coordinates of the upper left corner and the lower right corner of the target area. These coordinates will be used in the subsequent image cropping operation to extract the target part;
[0048] If the target detection model Failed to calibrate the left and right underwater target images If the target is successfully detected in both the left and right images, it means that the underwater binocular camera is too far away from the target or the image is not clear enough. Adjust the position of the underwater binocular camera until the detection model can successfully detect the target in both the corrected left and right images.
[0049] Underwater left and right eye images based on successful detection of the target part , from the target detection model Extract the pixel coordinates of the upper left corner of the target area from the output and the pixel coordinates of the lower right corner , and use the cropping function of the image processing tool to crop within a certain coordinate range to obtain an underwater binocular image that only contains the target part ;
[0050] Wherein, the target detection model The construction includes the following steps:
[0051] Use binocular cameras to capture several sets of underwater target images at different angles, scenes, and distances to form an underwater cooperative target image dataset. ;
[0052] Underwater cooperative target image dataset using underwater camera calibration parameters All images in the dataset are corrected to obtain the underwater cooperative target image dataset. ;
[0053] Use the labeling tool LabelImg to annotate the underwater cooperative target image dataset Label the objects in all images in the image, specifically draw a bounding box for each target object in the image and assign a category label to it. Each image will generate a corresponding .txt format label text to record the corresponding annotation information;
[0054] The underwater target image dataset is constructed in a 7:3 ratio. Divide into training set and validation set;
[0055] Choose to use the C2f RFCA module to replace the original C2f module in the YOLOv8-s version model, and then use the above-divided training set to train the improved model;
[0056] Use the trained model to evaluate on the validation set divided above to determine the performance of the model and adjust the model parameters as needed;
[0057] Set the number of training times and repeat the model training and validation process;
[0058] Compare the evaluation indicators of the models obtained in each round of training on the validation set to select the best target detection model .
[0059] The ORB algorithm is used to extract and match feature points in the target area, and the feature point extraction part of the matching feature points includes two steps: extracting feature points using the improved oFAST method and converting image feature points into binary descriptors using the rBRIEF method:
[0060] The method of extracting feature points using the improved oFAST method includes:
[0061] Adopting the idea of adaptive threshold segmentation, a threshold t is defined for each pixel p: ,in, is the adjustment coefficient, is the brightness of the pixel with the highest brightness on the circumference; is the lowest pixel brightness on the circumference; is the average brightness of 16 pixels. is the brightness of the i-th pixel on the circumference;
[0062] Harris corner detection method is used to sort feature points and screen potential important feature points;
[0063] Apply FAST feature extraction at different scales of the image;
[0064] Define the Rosin domain moment as , where m pq for The two-dimensional moment in the window area, p and q represent the powers of the x and y coordinates when calculating the moment, I(x,y) represents the gray value of the pixel point, and the direction of the offset vector between the corner point and the centroid is defined as the direction of the feature point using the gray centroid method, including the following steps:
[0065] Calculate a feature point around Grayscale centroid of the window area;
[0066] ;
[0067] in, for The 0th-order moment in the window area represents the sum of the brightness values of all pixels in the window area. for In the window area The first-order moment in the direction represents the weighted average of the brightness values of all pixels in the window area in the horizontal direction. for In the window area The first-order moment in the direction represents the weighted average of the brightness values of all pixels in the window area in the vertical direction;
[0068] Calculate the offset vector between the feature point and the centroid, and define its direction as the direction of the feature point;
[0069] ;
[0070] in, for The image coordinates within the window area, For coordinates The brightness value at ;
[0071] The method of converting image feature points into binary descriptors using the rBRIEF method includes:
[0072] Define a fixed size around the feature point. Window, extract the pixel value of this window area to obtain the grayscale information around the feature point;
[0073] Select a set of predefined sampling points for generating descriptors;
[0074] For each pair of sampling points , use the following formula to compare their gray values;
[0075] ;
[0076] in, Pixel The gray value of Pixel In this way, all comparison results are combined into a binary string to form the final BRIEF descriptor;
[0077] The ORB algorithm is used to extract and match feature points in the target area, and the matching part of the feature points in the matching feature points includes:
[0078] Calculate the shortest and second shortest Hamming distances between the feature descriptors of the left image and the right image. If the shortest Hamming distance is less than 50 and the ratio of the shortest Hamming distance to the second shortest Hamming distance is less than 0.8, the two feature points are considered to be matched and the matching feature points are obtained.
[0079] Use the RANSAC algorithm to remove the left and right target images For the mismatched points in the image, we first randomly select a small number of points from the feature matching points in the left and right images for model fitting to obtain a model that can represent the true matching relationship in the data. Then, we use this model to calculate the reprojection errors of all feature point pairs.
[0080] According to the set reprojection error threshold, the point pairs that meet the model are marked as inliers. These inliers are considered to be accurate matching relationships, while the points that do not meet the threshold conditions are marked as outliers, indicating that they are the result of mismatching. Finally, the left and right eye images are obtained. All matching feature point pairs.
[0081] The method of calculating the three-dimensional coordinates of the target in the camera coordinate system by matching feature points includes:
[0082] Based on obtaining left and right eye images All matching feature point pairs are obtained by using the parallax principle and triangulation method to obtain the depth information Z of the target from the camera in the image. Through the depth information Z, the three-dimensional coordinates of the target in the camera coordinate system are further calculated to achieve the positioning of the target. The specific steps are as follows:
[0083] The disparity value between the same pair of matching feature points in the left and right images is calculated by the difference of the pixel horizontal coordinates, which is the absolute value of the two horizontal coordinates. , the calculation formula is as follows:
[0084] ;
[0085] in, is the pixel horizontal coordinate of the i-th feature point in the left image, is the pixel horizontal coordinate of the i-th feature point in the right image, is the disparity value between the i-th pair of matching feature points of the left and right images;
[0086] Then, the triangulation method is used to calculate the depth distance from each pair of feature points on the left and right images to the binocular camera. , the calculation formula is as follows:
[0087] ;
[0088] in, is the focal length of the camera in the comprehensive calibration of underwater cameras, is the baseline distance between the binocular cameras;
[0089] Take the left camera coordinate system as the base coordinate system and normalize the left camera coordinate system. Pixel coordinates of feature points in the image Convert to image coordinates , using the focal length and image coordinates in the underwater binocular camera comprehensive calibration parameters, the cropped left target image The pixel coordinates of the feature points in the image are normalized to the image coordinates. Specifically, the cropped left target image is first The pixel coordinates of the feature points in the left image are converted into the feature points in the left image before cropping The pixel coordinates of , the normalized calculation formula is as follows:
[0090] ;
[0091] in are the principal point coordinates of the image before cropping, , The normalized focal length of the internal parameters in the comprehensive calibration of the underwater binocular camera. is the cropped left target image The pixel coordinates of the feature points in Left eye image before cropping The pixel coordinates of the upper left corner of the target area, is the image coordinate of the feature point;
[0092] Convert the image coordinates of the feature points to the left camera coordinate system to obtain the left camera coordinate system. The three-dimensional coordinates of each feature point in the image in the left camera coordinate system , the calculation formula is as follows:
[0093] ;
[0094] The 3D coordinates of each feature point are averaged to obtain a 3D coordinate representing the actual position of the cooperative target in the left camera coordinate system. :
[0095] ;
[0096] in, is the deviation of the lateral distance of the cooperative target relative to the optical center of the camera, is the deviation of the longitudinal distance of the cooperative target relative to the optical center of the camera, is the depth information of the target from the camera.
[0097] Second aspect. The present invention provides an underwater binocular positioning device based on target assistance, comprising:
[0098] An internal and external parameter acquisition module, used to acquire the internal and external parameters of the binocular camera, wherein the internal and external parameters of the binocular camera are obtained by calibrating the binocular camera in air;
[0099] The refraction parameter acquisition module is used to obtain the calibration evaluation function. The calibration evaluation function is constructed based on the refraction model of the underwater binocular camera, and then the genetic algorithm is used to optimize the calibration evaluation function to obtain the refraction parameters of the binocular camera;
[0100] The target image acquisition module is used to control the binocular camera to shoot and collect the target image, obtain the initial underwater cooperative positioning target image, and use the internal and external parameters and refraction parameters of the binocular camera to correct the initial underwater cooperative positioning target image to obtain the corrected underwater cooperative positioning target image;
[0101] The target area selection module is used to detect the target in the corrected underwater cooperative positioning target image using the YOLOv8 model, and select the target area in the underwater cooperative positioning target image;
[0102] The feature point extraction and matching module is used to extract and match the feature points of the target area using the ORB algorithm to obtain matching feature points;
[0103] The three-dimensional coordinate calculation module is used to calculate the three-dimensional coordinates of the target in the camera coordinate system by matching feature points.
[0104] A third aspect: The present invention provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the target-assisted underwater binocular positioning method.
[0105] The beneficial effects of the present invention are as follows: the present invention provides a target-assisted underwater binocular positioning method, device and storage medium, which adopts underwater binocular positioning to overcome the limitations of monocular positioning in depth calculation, calibrates the binocular camera in the air, and obtains the camera's internal parameters and distortion coefficients; establishes an underwater binocular camera refraction model, optimizes the refraction parameters and camera parameters using a genetic algorithm, and realizes accurate correction of underwater binocular images; uses an improved YOLOv8 model to perform target detection and select a target area; performs feature extraction and matching on the target area, removes mismatched points, and calculates the parallax of key points of the target; finally, calculates the three-dimensional coordinates of the target based on the parallax principle and triangulation, and accurately locates the underwater target. In summary, the present invention can maintain efficient image processing and positioning capabilities in a complex underwater environment, significantly improves the accuracy and efficiency of underwater positioning, and provides more reliable technical support for underwater operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Figure 1 It is a flow chart of a target-assisted underwater binocular positioning method of the present invention;
[0107] Figure 2 is a schematic diagram of an imaging device in the present invention;
[0108] Figure 3 It is an underwater secondary refraction projection cross-sectional diagram in the present invention;
[0109] Figure 4 It is the flow chart of the genetic algorithm in the present invention;
[0110] Figure 5 This is a schematic diagram of the cooperative target proposed in the present invention;
[0111] Figure 6 Schematic diagram of the C2f RFCA module in the present invention. DETAILED DESCRIPTION
[0112] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the protection scope of the present invention.
[0113] Example 1
[0114] like Figure 1As shown, the present invention discloses an underwater binocular positioning method based on target assistance, which captures the image data of the underwater target through binocular vision, and then uses advanced algorithms in the image processing module to analyze and process the captured image to accurately extract the position information of the target, thereby realizing the rapid and accurate positioning of the underwater target. The method of the present invention is realized based on the image processing module and the underwater binocular vision positioning software on the computer. The image processing module is mainly composed of a binocular camera and a waterproof shell, and the two waterproof shells are strictly encapsulated on the outside to meet the function of underwater image acquisition. The waterproof shell adopts a cylindrical design, and the top is made of glass material so that the binocular camera can shoot images underwater; the barrel and the bottom are made of waterproof plastic material to prevent water from entering. A round opening of the size of the camera industrial data cable is left at the bottom to facilitate the data cable to pass through. An industrial binocular camera with a resolution not less than is used for image acquisition, and the image acquisition frame rate is set to 30~40FPS. The binocular camera is connected to the computer via a USB3.0 industrial camera data cable, and the data cable is equipped with an M2 metric screw to ensure that the plug is firmly fixed on the camera body. Through this data cable, not only can the camera be powered, but also the acquired image can be transmitted to the computer.
[0115] The underwater binocular positioning method of the present invention specifically comprises the following steps:
[0116] Step 1: Obtain the internal and external parameters of the binocular camera. The internal and external parameters of the binocular camera are obtained by calibrating the binocular camera in air.
[0117] Binocular cameras are used in underwater scenes. The light emitted by underwater objects needs to be refracted through the three media of water, glass and air before reaching the imaging surface inside the camera. Therefore, the camera model in the air is no longer valid, and the traditional camera calibration method cannot accurately calibrate it. Therefore, an underwater camera refraction model is established, and the underwater camera imaging model is comprehensively calibrated.
[0118] Since the basic internal parameters of the binocular camera are determined by the factory and will not change with the camera posture and external environment, the calibration of the basic parameters of the binocular camera is completed in the air.
[0119] like Figure 2 As shown, first, the binocular camera is placed in parallel on the binocular board with a baseline distance of 8 cm to 15 cm in the air, and the calibration plate is placed at a distance of 1.0 m to 2.0 m from the camera to ensure that the calibration plate appears in the field of view of the binocular camera.
[0120] When the calibration plate occupies 1 / 3 to 2 / 3 of the binocular camera's field of view and the image is clear, the distance from the calibration plate to the binocular camera is fixed, and the underwater binocular vision positioning server software system implemented by the present invention controls the binocular camera to synchronously shoot calibration plate images in different postures. In order to reduce the error caused by the small number of calibration plate images, 15 groups or more of calibration plate images in different postures need to be shot.
[0121] Using Zhang's calibration method to solve the camera intrinsic parameter matrix With the external parameter matrix The product of , and then solve the internal parameter matrix and the external parameter matrix ,
[0122] ;
[0123] in, is the extrinsic rotation matrix, is the extrinsic translation matrix, is the scaling factor, and is the focal length The normalized coordinates of is the principal point coordinate of the camera, that is, the intersection of the camera optical axis and the imaging plane, is the coordinate in the pixel coordinate system, is the coordinate in the world coordinate system.
[0124] The intrinsic parameter matrix contains information such as the camera's focal length, principal point coordinates, and distortion coefficients, while the extrinsic parameter matrix describes the camera's position and orientation in the world coordinate system.
[0125] Based on Zhang's calibration method, the description of camera tangential distortion is further improved, and the tangential distortion is described as:
[0126] ;
[0127] in, is the actual observed image coordinate, is the image coordinate without tangential distortion, is the tangential distortion coefficient of the camera, is the square distance from the point to the center of the image;
[0128] The three radial distortion coefficients and two tangential distortion coefficients are solved by the least squares method, and combined with the internal parameter matrix , extrinsic rotation matrix and the extrinsic translation matrix The reprojection error function is constructed, and then the maximum likelihood estimation method is used for nonlinear optimization to obtain the optimized internal and external parameters of the binocular camera. The reprojection error function is shown in the following formula:
[0129] ;
[0130] Among them, F represents the error size, n represents the number of calibration plate images obtained in different directions in the air, and m represents the total number of corner points on each calibration plate image. Indicates On the image The coordinates of the corner points, Representing 3D points In the The reprojection result on the image is is the camera distortion coefficient, is the radial distortion coefficient of the camera, is the tangential distortion coefficient of the camera, For the The extrinsic rotation matrix and extrinsic translation matrix of the image, is the three-dimensional coordinate in the world coordinate system, indicating the The true 3D position of the corner points is achieved by minimizing the reprojection error function F. Optimization of six parameters.
[0131] In each iteration, the reprojection error is calculated based on the current parameters, and the parameters are updated until they converge to the optimal solution to obtain the basic parameters of the camera in air.
[0132] Step 2: Obtain a calibration evaluation function. The calibration evaluation function is constructed based on the refraction model of the underwater binocular camera, and then the genetic algorithm is used to optimize the calibration evaluation function to obtain the refraction parameters of the binocular camera.
[0133] Based on the above method of shooting the calibration plate image in the air, the binocular camera is placed in a waterproof housing device, and 15 groups or more of underwater calibration plate images are shot in the same way.
[0134] like Figure 3 As shown, an object point in underwater three-dimensional space The emitted light propagates to The process of imaging the camera surface with the optical center is divided into two stages. The first stage: the underwater object emits The light propagates to a point on the outer surface of the glass Refraction occurs. Stage 2: The light passes through the first refraction point. Continue to propagate to the point It undergoes secondary refraction and is then projected onto the camera imaging surface. Image point. The first refraction model and the second refraction model that occur when the light passes through the glass surface in these two stages are established through formulas. Among them, , are the incident angle and the outgoing angle of the double refraction, is the first refraction angle, , , represent the relative refractive index of air, glass, and water respectively. is the thickness of the glass surface, Indicates the distance from the optical center to the glass cover.
[0135] The first refraction model and the second refraction model that occur when light passes through the glass surface are established. The expression of the first refraction model is as follows:
[0136]
[0137] The expression of the double refraction model is as follows:
[0138] ;
[0139] in, , are the incident angle and the outgoing angle of the double refraction, is the first refraction angle, , , represent the relative refractive index of air, glass, and water respectively. is the thickness of the glass cover, Indicates the distance from the optical center to the glass cover; Indicates point To the camera coordinate system The vertical distance on the axis, Indicates point The distance to the outer surface of the glass cover, Indicates point Projecting point on the outer surface of the glass cover To the camera coordinate system The vertical distance on the axis, Indicates point Refraction point on the inner surface of the glass cover To the camera coordinate system The vertical distance on the axis.
[0140] Combine the first refraction model and the second refraction model, and The restrictions are available The only solution is to obtain the refraction surface of the glass cover point, because To camera optical center The projection process is carried out in the air, and then the projection point on the imaging surface can be obtained using the air camera model. Finally, the space point Projection point to the camera imaging surface The relationship between Point transformation solution The process requires the use of the refractive structure parameters of the above-mentioned binocular system model, such as: glass refractive surface , glass thickness , distance from optical center to glass surface , so the final projection point is The expression of also contains this information, so that the spatial object point can be analyzed from the binocular model The projection point and the actual image point on the camera imaging surface The relationship between them is established by minimizing the difference between them:
[0141] ;
[0142] Among them, E represents the value of the calibration evaluation function. The smaller the E value, the closer the actual detected corner coordinates are to the corner coordinates calculated using the imaging model. m is the number of calibrated plate images obtained at different underwater orientations. is the total number of corner points extracted from each calibration plate image, , are the actual coordinates of the i-th corner point detected on the left and right camera imaging surfaces, respectively. , are the coordinates of the i-th corner point obtained by projecting the underwater 3D point onto the left and right camera imaging surfaces through model calculation, as well as They respectively represent the normal vector of the glass cover refraction surface, the distance from the optical center to the glass cover, and the thickness of the glass cover in the left and right camera refraction models in the underwater binocular camera refraction model.
[0143] The genetic algorithm is used to optimize the calibration evaluation function to obtain the refractive parameters of the binocular camera. First, the threshold range of the unknown parameters to be optimized is set, and the initial population is randomly generated. Each individual represents a set of initial parameter values. The performance of each individual is evaluated by the fitness function, and individuals with high fitness are selected as parents for crossover and mutation operations to generate new offspring individuals. After multiple generations of selection, crossover, and mutation, the fitness in the population will gradually increase, and the optimal solution will eventually be found. Through iterative updates, the genetic algorithm can effectively explore the parameter space, thereby finding the global optimal solution of the calibration evaluation function, such as Figure 4 As shown, it mainly includes the following steps:
[0144] Step a, initializing the population: Set the initial population size to 100 individuals and randomly generate the initial population. Each individual in the population is represented by a ten-dimensional vector express, , where the parameter to be optimized is the vector Each parameter is generated within a defined threshold range, where and They are the thickness of the left and right glass waterproof covers respectively. and are the distances from the optical centers of the left and right cameras to the glass refraction surface, and The normal vectors of the glass cover refraction surface in front of the left and right cameras are The component in the axial direction, The components in the axial direction and in Component in the direction of the axis.
[0145] Step b, set the fitness function: define the fitness function , which accepts each individual in the population As input and return the above calibration evaluation function The negative value of , the optimization goal is to maximize the fitness function, that is, to minimize the calibration evaluation function.
[0146] Step c, select operation: use the fitness function to calculate each individual in the population The fitness of the individual is 30%, and 30% of the individuals are selected from the population as the parent individuals according to the roulette wheel selection method. , the remaining 70% of individuals are used as offspring individuals .
[0147] Step d, crossover operation: for the selected parent individuals Pairing and single-point crossover to generate new offspring individuals , that is, each pair of parents exchanges parameters through a randomly selected crossover point, and each crossover generates two offspring individuals If the number of parent individuals is even, a perfect pairing is formed; if it is odd, the last parent individual can be directly copied as an offspring, and the generated offspring individual Directly replace the parent individual .
[0148] Step e, mutation operation: from the generated offspring individuals Select 5% of the individuals for mutation operation, that is, select A certain parameter is randomly adjusted slightly, and the mutated individual replaces the original individual.
[0149] Step f, new population generation: Generate a new population through selection, crossover and mutation operations, that is, all individuals in the original population , now a new individual ,
[0150] Step g, iterative process: set the termination condition, the maximum number of iterations is 100 times or the average fitness value of the individuals in the population is >-0.05, and before any termination condition is met, repeat the above selection, crossover and mutation operations to optimize the population.
[0151] Step h, output the optimal solution: after reaching the termination condition, select the individual with the highest fitness as the optimal solution, and output the parameter value corresponding to the individual.
[0152] At this point, the comprehensive calibration of the underwater binocular camera based on the refraction model is completed, and the basic parameters of the left and right cameras are obtained. and And the normal vector of the waterproof cover glass surface , , thickness of left and right glass waterproof cover and , the distance from the optical center of the left and right cameras to the glass refractive surface , , these parameters will be used for subsequent image correction processing and target positioning.
[0153] Step three, control the binocular camera to shoot and collect the target image to obtain the initial underwater cooperative positioning target image, use the internal and external parameters and refraction parameters of the binocular camera to correct the initial underwater cooperative positioning target image to obtain the corrected underwater cooperative positioning target image.
[0154] Figure 5 A schematic diagram of a cooperative target provided by the present invention, the cooperative target is in the form of an active target, which is square in shape, 30 cm in side length, 2 cm in thickness, made of aluminum alloy, and coated with a waterproof coating on the surface. Eight green high-brightness LED lights are set on the front of the target as active light sources, evenly distributed to form a ring.
[0155] Based on the above positioning cooperative target, the target image is photographed and collected using a binocular camera to obtain the underwater positioning cooperative target image. However, due to the particularity of the underwater environment, the captured target image will produce various distortions. Therefore, the underwater camera calibration parameters obtained in step 1 are used to correct the target image to obtain the corrected underwater target image. .
[0156] First, the distortion model is constructed by obtaining the distortion coefficient of the binocular camera. For each pixel (x, y), the corresponding point in the original image is calculated by the distortion model using the reverse mapping method. .
[0157]
[0158] in, , is the radial distortion coefficient, is the tangential distortion coefficient.
[0159] During the mapping process, since non-integer coordinates may appear, the bilinear interpolation method is used to obtain the pixel value of the new coordinate position.
[0160] Step 4: Use the YOLOv8 model to detect the target in the corrected underwater cooperative positioning target image, and select the area where the target is located in the underwater cooperative positioning target image.
[0161] right After distortion correction, Epipolar correction is performed to make the corresponding pixels of the left and right images on the same horizontal line, and finally the corrected underwater target image is obtained. .
[0162]
[0163] Based on the improved YOLOv8 model, the underwater cooperative target image is For target detection, due to the particularity of the underwater environment and the disturbance of the water body, the underwater target image obtained above may have motion blur, making it difficult to extract the features of the target part. Therefore, the original YOLOv8 model is difficult to effectively capture the features of blurred objects in this environment.
[0164] Combination Figure 6 To solve this problem, this paper proposes a C2f RFCA module as an enhanced version of the original C2f module in YOLOv8. This module improves feature extraction by utilizing multi-scale receptive fields and complex feature fusion strategies, enabling the model to better capture local and global information. Even under low-contrast and motion blur conditions, this module can more accurately detect cooperative targets.
[0165] Based on the above improved YOLOv8 model, target detection is performed on underwater cooperative target images. The specific steps are as follows:
[0166] Step a: Use a binocular camera to capture more than 2000 sets of underwater target images at different angles, scenes, and distances to form an underwater cooperative target image dataset. .
[0167] The underwater cooperative target image dataset is calibrated using underwater camera calibration parameters. All images in the dataset are corrected to obtain the underwater cooperative target image dataset. .
[0168] Step b: Use the labeling tool LabelImg to label the underwater cooperative target image dataset The specific operation is to draw a bounding box for each target object in the image and assign a category label to it. For example, the target object is framed by a bounding box and labeled as category 0. Each image will generate a corresponding .txt format label text to record the corresponding annotation information.
[0169] Step c: the underwater target image dataset is processed according to the ratio of 7:3 Divide it into training set and validation set.
[0170] Step d, select the YOLOv8-s version model, and improve and train it based on the model. The specific operation is to replace the original C2f module in the YOLOv8-s version model with the C2f RFCA module proposed by the present invention, and then train the improved model using the above-divided training set.
[0171] Step e: Use the trained model to evaluate on the validation set divided above to determine the performance of the model and adjust the model parameters as needed.
[0172] Step f, set the number of training rounds to 200, and repeat steps e and f.
[0173] Step g: Compare the evaluation indicators of the models obtained in each round of training on the validation set, such as mAP (Mean Average Precision) and P (Precision), to select the best training model. .
[0174] Based on the above steps, a target detection model trained with the underwater cooperative target image dataset can be obtained, which is used to detect the target part in the image taken by the underwater binocular camera.
[0175] Then use the binocular camera to shoot underwater images to get the original left and right images , and based on the above image correction processing, the corrected left and right eye images are obtained .
[0176] Based on the best target detection model obtained above Correct the left and right images To perform target detection, the specific steps are as follows:
[0177] If the target detection model In correcting the left and right images If the target is successfully detected in both images, the detection model is used to select the target area in the two images. The detection box will contain the position information of the box, that is, the pixel coordinates of the upper left corner and the lower right corner of the target area. These coordinates will be used in the subsequent image cropping operation to extract the target part.
[0178] If the target detection model Failed to calibrate left and right images If both the left and right images successfully detect the target, it means that the underwater binocular camera is too far away from the target or the image is not clear enough. In this case, the position of the underwater binocular camera needs to be adjusted until the detection model can successfully detect the target in both the corrected left and right images.
[0179] Underwater left and right eye images based on successful detection of the target part , from the target detection model Extract the pixel coordinates of the upper left corner of the target area from the output and the pixel coordinates of the lower right corner , and use the cropping function of the image processing tool to crop within a certain coordinate range to obtain an underwater binocular image that only contains the target part , reduce the feature matching range of underwater cooperative targets and improve the efficiency and accuracy of subsequent feature matching.
[0180] Step 5: Use the ORB algorithm to extract and match feature points in the target area to obtain matching feature points.
[0181] The improved ORB algorithm is used for feature matching, which mainly includes two steps: feature point extraction and feature point matching. The improved ORB algorithm is used to extract and match feature points of an image, which specifically includes the following steps:
[0182] First of all, the feature extraction of the ORB algorithm is improved by the FAST algorithm, which can also be called oFAST. When using the oFAST method to detect feature points in an image, you first need to set a parameter, which is a fixed threshold between the center of the pixel and the center of the ring. Take the candidate feature point A as the center of the circle and compare its grayscale value with all the points on the circumference of the circle with A as the center. If there are enough points on the circumference whose grayscale value difference with point A exceeds the set threshold, the candidate point A is considered a feature point.
[0183] In an underwater environment, the brightness of the image captured by the binocular camera will change, so there may be large differences in the grayscale values of different areas in the image. In this case, it is not appropriate to continue to use a fixed threshold to determine whether each pixel is a feature point, which may lead to the misextraction and misexclusion of feature points. Therefore, for underwater environments, the present invention adopts the idea of adaptive threshold segmentation to improve the fixed threshold strategy in the oFAST method. Specifically, for each pixel in the image, a different dynamic local threshold is set to achieve more flexible and accurate feature point extraction.
[0184] The improved oFAST method is used to extract feature points. The specific steps are as follows:
[0185] Step a, using the idea of adaptive threshold segmentation, defines a threshold t for each pixel p, that is:
[0186]
[0187] in, is the adjustment coefficient, is the brightness of the pixel with the highest brightness on the circumference; is the lowest pixel brightness on the circumference; is the average brightness of 16 pixels. is the brightness of the i-th pixel on the circumference. , , Neither is a fixed value, so it is a dynamic local threshold.
[0188] Step b: Harris corner detection method is used to sort the feature points to help screen potential important feature points.
[0189] In step c, FAST feature extraction is applied at different scales of the image so as to extract corresponding features in each scale layer.
[0190] Step d, define the Rosin domain moment as , where m pq for The two-dimensional moment in the window area, p and q represent the powers of the x and y coordinates when calculating the moment, and I(x,y) represents the gray value of the pixel. The direction of the offset vector between the corner point and the centroid is defined as the direction of the feature point using the grayscale centroid method, including the following steps:
[0191] Calculate a feature point around Grayscale centroid of the window area.
[0192]
[0193] in, for The 0th-order moment in the window area represents the sum of the brightness values of all pixels in the window area. for In the window area The first-order moment in the direction represents the weighted average of the brightness values of all pixels in the window area in the horizontal direction. for In the window area The first-order moment in the vertical direction represents the weighted average of the brightness values of all pixels in the window area in the vertical direction.
[0194] Calculate the offset vector between the feature point and the centroid, and define its direction as the direction of the feature point.
[0195]
[0196] in, for The image coordinates within the window area, For coordinates The brightness value at .
[0197] In this way, the improved oFAST method is used to successfully extract the feature points of the cooperative target in the image. Next, the rBRIEF method is used to convert the detected image feature points into binary descriptors to facilitate subsequent feature point matching.
[0198] Based on the rBRIEF method, the image feature points are converted into binary descriptors. The main steps are as follows:
[0199] Step a: define a fixed size around the feature point. The pixel values of this window area are extracted to obtain the grayscale information around the feature point.
[0200] Step b: Select a set of predefined sampling points for generating descriptors.
[0201] Step c, for each pair of sampling points , use the following formula to compare their grayscale values.
[0202]
[0203] in, Pixel The gray value of Pixel Gray value
[0204] In this way, the results of all comparisons are combined into a binary string to form the final BRIEF descriptor.
[0205] Based on the above, the left and right images can be In the extraction of feature points of the left image and the generation of its descriptor, feature points in the right image can also be extracted and descriptors can be generated.
[0206] Next, All feature points in the left and right images are matched. The specific steps are as follows:
[0207] Calculate the shortest and second shortest Hamming distances between the feature descriptors of the left image and the right image. If the shortest Hamming distance is less than 50 and the ratio of the shortest Hamming distance to the second shortest Hamming distance is less than 0.8, the two feature points are considered to match.
[0208] Based on the above method, the improved ORB algorithm is used to realize the feature point extraction and matching of the left and right images. However, due to the objective noise interference in the underwater environment, the feature points may be mismatched. Therefore, the mismatch point elimination algorithm is applied to remove these mismatch point pairs.
[0209] Use the RANSAC algorithm to remove the left and right target images To identify the mismatched points in the image, we first randomly select a small number of points from the feature matching points in the left and right images for model fitting. The goal is to find a model that can represent the true matching relationship in the data. Next, based on this model, we calculate the reprojection error of all feature point pairs.
[0210] The reprojection error is a key metric for evaluating the quality of model fit, which represents the difference between the actual observed points and the model's predicted points. By comparing these reprojection errors, we can determine which feature point pairs meet the model's expectations.
[0211] According to the set reprojection error threshold, the point pairs that meet the model are marked as inliers, which are considered to be accurate matches. Points that do not meet the threshold conditions are marked as outliers, indicating that they may be the result of mismatches.
[0212] After multiple iterations and random selections, the RANSAC algorithm will find the best model with the largest number of inliers, thereby eliminating mismatched points.
[0213] So far, based on the above, the improved ORB algorithm is used to realize the left and right eye images that only retain the target part Feature points are extracted and matched, and the RANSAC algorithm is used to eliminate mismatched points.
[0214] Step 6: Use the matching feature points to calculate the three-dimensional coordinates of the target in the camera coordinate system.
[0215] According to the above steps, only the left and right images of the target part are retained. All matching feature point pairs of Since the image has been calibrated, the matched feature point pairs on the image are all located on the horizontal line, that is, the pixel vertical coordinates of the same feature point pair are the same.
[0216] Based on the obtained left and right images All matching feature point pairs are obtained by using the parallax principle and triangulation method to obtain the depth information Z of the target from the camera in the image. Through the depth information Z, the three-dimensional coordinates of the target in the camera coordinate system are further calculated to achieve the positioning of the target. The specific steps are as follows:
[0217] First, the left image The pixel coordinates of the feature points are , then the right eye image The pixel coordinates of the feature points are .
[0218] Since the same pair of matching feature points of the left and right images are located on the same epipolar line, the disparity value between them can be calculated by the difference of the horizontal coordinates of the pixels, specifically the absolute values of the two horizontal coordinates. The calculation formula is as follows:
[0219]
[0220] in, is the pixel horizontal coordinate of the feature point of the left image, is the pixel horizontal coordinate of the feature point in the right image.
[0221] Based on the above, the disparity value between each pair of matching feature points of the left and right images is obtained Then, the depth distance from each pair of feature points on the left and right images to the binocular camera is calculated using triangulation. , the calculation formula is as follows:
[0222]
[0223] in, is the focal length of the camera in the underwater camera comprehensive calibration in step 1, is the baseline distance between the binocular cameras.
[0224] Get the image The distance information of each pair of feature points to the imaging plane of the binocular camera , taking the left camera coordinate system as the base coordinate system, normalize the left camera Pixel coordinates of feature points in the image Convert to image coordinates , since the target area in the underwater binocular image is cropped after target detection, the cropped target image The pixel coordinates of the feature points in the image before cropping The pixel coordinates of the feature points in the image are inconsistent. The focal length and image coordinates in the comprehensive calibration parameters of the underwater binocular camera are used to crop the left target image. To normalize the pixel coordinates of the feature points in the image to the image coordinates, you need to first crop the left target image The pixel coordinates of the feature points in the left image are converted into the feature points in the left image before cropping The pixel coordinates of , the normalized calculation formula is as follows:
[0225]
[0226] in are the principal point coordinates of the image before cropping, , The normalized focal length of the internal parameters in the comprehensive calibration of the underwater binocular camera. is the cropped left target image The pixel coordinates of the feature points in Left eye image before cropping The pixel coordinates of the upper left corner of the target area, is the image coordinate of the feature point.
[0227] Convert the image coordinates of the feature points to the left camera coordinate system to obtain the left camera coordinate system. The three-dimensional coordinates of each feature point in the image in the left camera coordinate system , the calculation formula is as follows:
[0228]
[0229] Due to the existence of objective errors, the 3D coordinates of each feature point are averaged to reduce the influence of inaccurate positioning caused by the large 3D coordinate error of a single feature point, and a 3D coordinate representing the actual position of the cooperative target in the left camera coordinate system is obtained. .
[0230]
[0231] in, It is the deviation of the lateral distance of the cooperative target relative to the optical center of the camera. It is the deviation of the longitudinal distance of the cooperative target relative to the optical center of the camera. That is the depth information of the target from the camera.
[0232] Based on the above, by taking a group of images of the underwater cooperative target with a binocular camera, the three-dimensional coordinates of the target at this moment in the left camera coordinate system can be obtained.
[0233] After the target is accurately positioned, it can be used to accurately locate the target object. Simply place the target on the target object, and by positioning the target, the position of the target object can be accurately determined. This method is not only simple and efficient, but can also be widely used in the tracking and positioning of various objects.
[0234] Example 2
[0235] This embodiment discloses an underwater binocular positioning device based on target assistance, comprising:
[0236] An internal and external parameter acquisition module is used to obtain the internal and external parameters of the binocular camera. The internal and external parameters of the binocular camera are obtained by calibrating the binocular camera in the air.
[0237] The refraction parameter acquisition module is used to obtain the calibration evaluation function. The calibration evaluation function is constructed based on the refraction model of the underwater binocular camera, and then the genetic algorithm is used to optimize the calibration evaluation function to obtain the refraction parameters of the binocular camera;
[0238] The target image acquisition module is used to control the binocular camera to shoot and collect the target image, obtain the initial underwater cooperative positioning target image, and use the internal and external parameters and refraction parameters of the binocular camera to correct the initial underwater cooperative positioning target image to obtain the corrected underwater cooperative positioning target image;
[0239] The target area selection module is used to detect the target in the corrected underwater cooperative positioning target image using the YOLOv8 model, and select the target area in the underwater cooperative positioning target image;
[0240] The feature point extraction and matching module is used to extract and match the feature points of the target area using the ORB algorithm to obtain matching feature points;
[0241] The three-dimensional coordinate calculation module is used to calculate the three-dimensional coordinates of the target in the camera coordinate system by matching feature points.
[0242] Example 3
[0243] This embodiment discloses a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, an underwater binocular positioning method based on target assistance is implemented.
[0244] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0245] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0246] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0247] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An underwater binocular positioning method based on target assistance, characterized in that: The following steps are involved: Obtaining the internal and external parameters of the binocular camera, which are obtained by calibrating the binocular camera in air; Obtain a calibration evaluation function, which is constructed based on the underwater binocular camera refraction model, and then use a genetic algorithm to optimize the calibration evaluation function to obtain the refraction parameters of the binocular camera; Controlling the binocular camera to shoot and collect the target image to obtain an initial underwater cooperative positioning target image, and using the internal and external parameters and refraction parameters of the binocular camera to correct the initial underwater cooperative positioning target image to obtain a corrected underwater cooperative positioning target image; The YOLOv8 model is used to detect the target in the corrected underwater cooperative positioning target image to obtain the target area in the underwater cooperative positioning target image; The ORB algorithm is used to extract and match feature points in the target area to obtain matching feature points; The three-dimensional coordinates of the target in the camera coordinate system are calculated using the matching feature points.
2. The target-assisted underwater binocular positioning method according to claim 1, characterized in that: The method for calibrating the internal and external parameters of the binocular camera includes: Using Zhang's calibration method to solve the camera intrinsic parameter matrix With the external parameter matrix The product of , and then solve the internal parameter matrix and the external parameter matrix , ; in, is the extrinsic rotation matrix, is the extrinsic translation matrix, is the scaling factor, and is the focal length The normalized coordinates of is the principal point coordinate of the camera, that is, the intersection of the camera optical axis and the imaging plane, is the coordinate in the pixel coordinate system, is the coordinate in the world coordinate system; Based on Zhang's calibration method, the tangential distortion is described as: ; in, is the actual observed image coordinate, is the image coordinate without tangential distortion, is the tangential distortion coefficient of the camera, is the square distance from the point to the center of the image; The three radial distortion coefficients and two tangential distortion coefficients are solved by the least squares method, and combined with the internal parameter matrix , extrinsic rotation matrix and the extrinsic translation matrix The reprojection error function is constructed, and then the maximum likelihood estimation method is used for nonlinear optimization to obtain the optimized internal and external parameters of the binocular camera. The reprojection error function is shown in the following formula: ; Among them, F represents the error size, n represents the number of calibration plate images obtained in different directions in the air, and m represents the total number of corner points on each calibration plate image. Indicates On the image The coordinates of the corner points, Representing 3D points In the The reprojection result on the image is is the camera distortion coefficient, is the radial distortion coefficient of the camera, is the tangential distortion coefficient of the camera, For the The extrinsic rotation matrix and extrinsic translation matrix of the image, is the three-dimensional coordinate in the world coordinate system, indicating the The true 3D position of the corner points is achieved by minimizing the reprojection error function F. Optimization of six parameters.
3. The target-assisted underwater binocular positioning method according to claim 2, characterized in that: The calibration evaluation function construction method includes: The first refraction model and the second refraction model that occur when light passes through the glass surface are established. The expression of the first refraction model is as follows: ; The expression of the double refraction model is as follows: ; in, , , Represent the relative refractive index of air, glass, and water, respectively. is the thickness of the glass cover, Indicates the distance from the optical center to the glass cover; Indicates point To the camera coordinate system The vertical distance on the axis, Indicates point The distance to the outer surface of the glass cover, Indicates point Projecting point on the outer surface of the glass cover To the camera coordinate system The vertical distance on the axis, Indicates point Refraction point on the inner surface of the glass cover To the camera coordinate system The vertical distance on the axis; Combining the first refraction model and the second refraction model, we get The only solution is to obtain the refraction surface of the glass cover point, and then use the air camera model to find the projection point on the imaging surface , and finally establish the space point Projection point to the camera imaging surface the relationship between; According to the underwater binocular camera refraction model, the spatial object point Projection point to the camera imaging surface The relationship between the two is transformed to obtain the binocular model and the projection point of the camera imaging surface is obtained and the actual image point on the camera imaging surface The relationship between them is established by minimizing the difference between them: ; Among them, E represents the value of the calibration evaluation function. The smaller the E value is, the closer the actual detected corner point coordinates are to the corner point coordinates calculated using the imaging model. is the total number of corner points extracted from each calibration plate image, , are the actual coordinates of the i-th corner point detected on the left and right camera imaging surfaces, respectively. , are the coordinates of the i-th corner point obtained by projecting the underwater 3D point onto the left and right camera imaging surfaces through model calculation, as well as They respectively represent the normal vector of the glass cover refraction surface, the distance from the optical center to the glass cover, and the thickness of the glass cover in the left and right camera refraction models in the underwater binocular camera refraction model.
4. The target-assisted underwater binocular positioning method according to claim 3, characterized in that: The method of optimizing the calibration evaluation function by using a genetic algorithm to obtain the refractive parameters of the binocular camera includes: Set the initial population size to 100 individuals and randomly generate the initial population. Each individual in the population is represented by a ten-dimensional vector express, , where the parameter to be optimized is the vector Each parameter is generated within a defined threshold range, where and They are the thickness of the left and right glass waterproof covers respectively. and are the distances from the optical centers of the left and right cameras to the glass refraction surface, and The normal vectors of the glass cover refraction surface in front of the left and right cameras are The component in the axial direction, The components in the axial direction and in Component in the axial direction; Define the fitness function , which accepts each individual in the population As input and return the above calibration evaluation function The negative value of , the optimization goal is to maximize the fitness function, that is, to minimize the calibration evaluation function; Using the fitness function, calculate each individual in the population The fitness of the individual is 30%, and 30% of the individuals are selected from the population as the parent individuals according to the roulette wheel selection method. , the remaining 70% of individuals are used as offspring individuals ; For the selected parent individuals Pairing and single-point crossover to generate new offspring individuals , that is, each pair of parents exchanges parameters through a randomly selected crossover point, and each crossover generates two offspring individuals If the number of parent individuals is even, a perfect pairing is formed; if it is odd, the last parent individual can be directly copied as an offspring, and the generated offspring individual Directly replace the parent individual ; From the generated offspring individuals Select 5% of the individuals for mutation operation, that is, select A certain parameter is randomly adjusted slightly, and the mutated individual replaces the original individual; A new population is generated through selection, crossover and mutation operations, that is, all individuals in the original population , now a new individual , Set the termination condition, the maximum number of iterations is 100 or the average fitness value of the individuals in the population is > -0.
05. Before any termination condition is met, repeat the above selection, crossover and mutation operations to optimize the population; After reaching the termination condition, the individual with the highest fitness is selected as the optimal solution, and the parameter value corresponding to the individual is output.
5. The target-assisted underwater binocular positioning method according to claim 4, characterized in that: The correction processing of the initial underwater cooperative positioning target image using the internal and external parameters and refraction parameters of the binocular camera includes: By obtaining the distortion coefficients of the binocular camera, a distortion model is constructed for the cooperative target image. For each pixel (x, y), the corresponding point in the original image is calculated by the distortion model using the reverse mapping method. : ; in, , is the radial distortion coefficient, is the tangential distortion coefficient; Cooperative target image After distortion correction, the cooperative target image Perform epipolar correction so that the corresponding pixels of the left and right images are on the same horizontal line, and finally obtain the corrected underwater target image .
6. The target-assisted underwater binocular positioning method according to claim 5, characterized in that: The YOLOv8 model uses a target detection model , the YOLOv8 model is used to detect the target in the corrected underwater cooperative positioning target image, and the target area in the underwater cooperative positioning target image includes: If the target detection model Correcting the underwater target images of the left and right eyes If the target is successfully detected in all, the target detection model is used In the left and right images, the target area is selected by the detection box, where the detection box contains the position information of the box selection, that is, the pixel coordinates of the upper left corner and the lower right corner of the target area. These coordinates will be used in the subsequent image cropping operation to extract the target part; If the target detection model Failed to calibrate the left and right underwater target images If the target is successfully detected in both the left and right images, it means that the underwater binocular camera is too far away from the target or the image is not clear enough. Adjust the position of the underwater binocular camera until the detection model can successfully detect the target in both the corrected left and right images. Underwater left and right eye images based on successful detection of the target part , from the target detection model Extract the pixel coordinates of the upper left corner of the target area from the output and the pixel coordinates of the lower right corner , and use the cropping function of the image processing tool to crop within a certain coordinate range to obtain an underwater binocular image that only contains the target part ; Wherein, the target detection model The construction includes the following steps: Use binocular cameras to capture several sets of underwater target images at different angles, scenes, and distances to form an underwater cooperative target image dataset. ; Underwater cooperative target image dataset using underwater camera calibration parameters All images in the dataset are corrected to obtain the underwater cooperative target image dataset. ; Use the labeling tool LabelImg to annotate the underwater cooperative target image dataset Label the objects in all images in the image, specifically draw a bounding box for each target object in the image and assign a category label to it. Each image will generate a corresponding .txt format label text to record the corresponding annotation information; The underwater target image dataset is constructed in a 7:3 ratio. Divide into training set and validation set; Choose to use the C2f RFCA module to replace the original C2f module in the YOLOv8-s version model, and then use the above-divided training set to train the improved model; Use the trained model to evaluate on the validation set divided above to determine the performance of the model and adjust the model parameters as needed; Set the number of training times and repeat the model training and validation process; Compare the evaluation indicators of the models obtained in each round of training on the validation set to select the best target detection model .
7. The target-assisted underwater binocular positioning method according to claim 6, characterized in that: The ORB algorithm is used to extract and match feature points in the target area, and the feature point extraction part of the matching feature points includes two steps: extracting feature points using the improved oFAST method and converting image feature points into binary descriptors using the rBRIEF method: Wherein, the feature point extraction using the improved oFAST method includes: Adopting the idea of adaptive threshold segmentation, a threshold t is defined for each pixel p: ,in, is the adjustment coefficient, is the brightness of the pixel with the highest brightness on the circumference; is the lowest pixel brightness on the circumference; is the average brightness of 16 pixels. is the brightness of the i-th pixel on the circumference; Harris corner detection method is used to sort feature points and screen potential important feature points; Apply FAST feature extraction at different scales of the image; Define the Rosin domain moment as , where m pq for The two-dimensional moment in the window area, p and q represent the powers of the x and y coordinates when calculating the moment, I(x,y) represents the gray value of the pixel point, and the direction of the offset vector between the corner point and the centroid is defined as the direction of the feature point using the gray centroid method, including the following steps: Calculate a feature point around Grayscale centroid of the window area; ; in, for The 0th-order moment in the window area represents the sum of the brightness values of all pixels in the window area. for In the window area The first-order moment in the direction represents the weighted average of the brightness values of all pixels in the window area in the horizontal direction. for In the window area The first-order moment in the direction represents the weighted average of the brightness values of all pixels in the window area in the vertical direction; Calculate the offset vector between the feature point and the centroid, and define its direction as the direction of the feature point; ; in, for The image coordinates within the window area, For coordinates The brightness value at ; The method of converting image feature points into binary descriptors using the rBRIEF method includes: Define a fixed size around the feature point. Window, extract the pixel value of this window area to obtain the grayscale information around the feature point; Select a set of predefined sampling points for generating descriptors; For each pair of sampling points , use the following formula to compare their gray values; ; in, Pixel The gray value of Pixel In this way, all comparison results are combined into a binary string to form the final BRIEF descriptor; The ORB algorithm is used to extract and match feature points in the target area, and the matching part of the feature points in the matching feature points includes: Calculate the shortest and second shortest Hamming distances between the feature descriptors of the left image and the right image. If the shortest Hamming distance is less than 50 and the ratio of the shortest Hamming distance to the second shortest Hamming distance is less than 0.8, the two feature points are considered to be matched and the matching feature points are obtained. Use the RANSAC algorithm to remove the left and right target images For the mismatched points in the image, we first randomly select a small number of points from the feature matching points in the left and right images for model fitting to obtain a model that can represent the true matching relationship in the data. Then, we use this model to calculate the reprojection errors of all feature point pairs. According to the set reprojection error threshold, the point pairs that meet the model are marked as inliers. These inliers are considered to be accurate matching relationships, while the points that do not meet the threshold conditions are marked as outliers, indicating that they are the result of mismatching. Finally, the left and right eye images are obtained. All matching feature point pairs.
8. The target-assisted underwater binocular positioning method according to claim 7, characterized in that: The method of calculating the three-dimensional coordinates of the target in the camera coordinate system by matching feature points includes: Based on obtaining left and right eye images All matching feature point pairs are obtained by using the parallax principle and triangulation method to obtain the depth information Z of the target from the camera in the image. Through the depth information Z, the three-dimensional coordinates of the target in the camera coordinate system are further calculated to achieve the positioning of the target. The specific steps are as follows: The disparity value between the same pair of matching feature points in the left and right images is calculated by the difference of the pixel horizontal coordinates, which is the absolute value of the two horizontal coordinates. , the calculation formula is as follows: ; in, is the pixel horizontal coordinate of the i-th feature point in the left image, is the pixel horizontal coordinate of the i-th feature point in the right image, is the disparity value between the i-th pair of matching feature points of the left and right images; Then, the triangulation method is used to calculate the depth distance from each pair of feature points on the left and right images to the binocular camera. , the calculation formula is as follows: ; in, is the focal length of the camera in the comprehensive calibration of underwater cameras, is the baseline distance between the binocular cameras; Take the left camera coordinate system as the base coordinate system and normalize the left camera coordinate system. Pixel coordinates of feature points in the image Convert to image coordinates , using the focal length and image coordinates in the underwater binocular camera comprehensive calibration parameters, the cropped left target image The pixel coordinates of the feature points in the image are normalized to the image coordinates. Specifically, the cropped left target image is first The pixel coordinates of the feature points in the left image are converted into the feature points in the left image before cropping The pixel coordinates of , the normalized calculation formula is as follows: ; in are the principal point coordinates of the image before cropping, , The normalized focal length of the internal parameters in the comprehensive calibration of the underwater binocular camera. is the cropped left target image The pixel coordinates of the feature points in Left eye image before cropping The pixel coordinates of the upper left corner of the target area, is the image coordinate of the feature point; Convert the image coordinates of the feature points to the left camera coordinate system to obtain the left camera coordinate system. The three-dimensional coordinates of each feature point in the image in the left camera coordinate system , the calculation formula is as follows: ; The three-dimensional coordinates of each feature point are averaged to obtain a three-dimensional coordinate representing the actual position of the cooperative target in the left camera coordinate system. : ; in, is the deviation of the lateral distance of the cooperative target relative to the optical center of the camera, is the deviation of the longitudinal distance of the cooperative target relative to the optical center of the camera, is the depth information of the target from the camera.
9. An underwater binocular positioning device based on target assistance, characterized in that: include: An internal and external parameter acquisition module, used to acquire the internal and external parameters of the binocular camera, wherein the internal and external parameters of the binocular camera are obtained by calibrating the binocular camera in air; The refraction parameter acquisition module is used to obtain the calibration evaluation function. The calibration evaluation function is constructed based on the refraction model of the underwater binocular camera, and then the genetic algorithm is used to optimize the calibration evaluation function to obtain the refraction parameters of the binocular camera; The target image acquisition module is used to control the binocular camera to shoot and collect the target image, obtain the initial underwater cooperative positioning target image, and use the internal and external parameters and refraction parameters of the binocular camera to correct the initial underwater cooperative positioning target image to obtain the corrected underwater cooperative positioning target image; The target area selection module is used to detect the target in the corrected underwater cooperative positioning target image using the YOLOv8 model, and select the target area in the underwater cooperative positioning target image; The feature point extraction and matching module is used to extract and match the feature points of the target area using the ORB algorithm to obtain matching feature points; The three-dimensional coordinate calculation module is used to calculate the three-dimensional coordinates of the target in the camera coordinate system by matching feature points.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the target-assisted underwater binocular positioning method described in any one of claims 1-8 is implemented.
Citation Information
Cited By
Transparent cell culture dish three-dimensional positioning method based on pure binocular vision
CN120374737A
A three-dimensional positioning method for transparent cell culture dishes based on pure binocular vision
CN120374737B
Tunnel lamp identification and spatial positioning method based on binocular vision
CN120526126A
A method for tunnel lamp recognition and spatial positioning based on binocular vision
CN120526126B
Binocular camera target detection and tracking system and method
CN120782821A