An underwater double-end alignment method and system based on visual servo control
By employing a two-stage alignment strategy based on visual servo control and image enhancement processing, combined with feature matching of LED and laser light sources, the problems of low alignment accuracy and poor efficiency in deep-sea environments have been solved, achieving high-precision and high-efficiency underwater alignment suitable for deep-sea environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE PLA NAVY SUBMARINE INST
- Filing Date
- 2026-05-29
- Publication Date
- 2026-06-30
AI Technical Summary
In deep-sea environments, traditional underwater alignment methods suffer from low precision, poor efficiency, and weak adaptability. In particular, under conditions of severe light absorption and scattering, and high requirements for equipment sealing and reliability, it is difficult to achieve high-precision and high-efficiency alignment of dual-end equipment.
A two-stage alignment strategy based on vision servo control is adopted, combining LED coarse alignment and laser fine alignment. Through image enhancement processing and ORB feature matching, combined with incremental PID feedback control, high-precision alignment is achieved.
It achieves high-precision and high-efficiency alignment in deep-sea environments, adapts to low visibility conditions, meets the real-time requirements of engineering applications, and the system design is suitable for alignment accuracy better than 1.0° and rapid alignment within 20 seconds at a water depth of 100 meters.
Smart Images

Figure CN122313245A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of underwater robots and marine engineering technology, specifically relating to an underwater dual-end alignment method and system based on visual servo control. Background Technology
[0002] As marine resource development continues to expand into the deep sea, the demand for high-precision alignment technology is becoming increasingly urgent for applications such as underwater docking, subsea pipeline maintenance, and underwater archaeology. In the deep-sea environment, achieving precise alignment of two-way devices faces numerous technical challenges due to factors such as immense water pressure, extremely low visibility, and severely limited communication.
[0003] Traditional underwater alignment methods mainly rely on acoustic positioning or manual remote control. Acoustic positioning methods are susceptible to multipath effects in underwater acoustic channels and interference from environmental noise, resulting in low accuracy and slow update rates. Manual remote control, on the other hand, is highly dependent on operator experience, inefficient, and risky, making it difficult to meet the automation and high-precision application requirements of modern deep-sea engineering.
[0004] In recent years, visual servo control technology has gradually become a research hotspot in the field of underwater alignment due to its advantages such as non-contact measurement, high precision, and good real-time performance. However, its application in the deep-sea environment still faces the following core challenges: First, the underwater environment suffers from severe light absorption and scattering, leading to significant degradation in image quality, reduced contrast, color distortion, and blurred details, which greatly complicates feature detection and recognition. Second, the high-pressure environment of the deep sea places stringent requirements on the sealing, reliability, and pressure resistance of the equipment. In addition, the alignment process needs to be completed within a limited time, placing high demands on the real-time performance of the algorithm and the stability of the control system.
[0005] Therefore, there is an urgent need for a method and system that can adapt to the complex environment of the deep sea and achieve high-precision, high-efficiency automatic alignment. Summary of the Invention
[0006] The purpose of this invention is to propose an underwater dual-end alignment method based on visual servo control to solve the problems of low alignment accuracy, poor efficiency, and weak adaptability of dual-end devices in deep-sea environments.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: An underwater dual-end alignment method based on visual servo control includes the following steps: Step 1. Start the system and enter standby mode, then turn on the LED and laser light sources for the current plane and the target plane; Step 2. The underwater camera acquires the original images of the LED and laser light sources that include the target plane. The original images acquired by the underwater camera are then enhanced to obtain the enhanced images. Step 3. Identify and locate feature points of LED and laser light sources from the enhanced image; Step 4. Calculate the attitude deviation angle of the current plane relative to the target plane based on the pixel coordinates of the feature points of the LED light source and the laser light source; Step 5. Generate control commands based on the attitude deviation angle of the current plane relative to the target plane; If the attitude deviation angle is not less than the coarse alignment threshold, then the coarse alignment strategy based on the feature points of the LED light source is executed, and control commands are generated to drive the two-dimensional turntable to rotate until the coarse alignment is completed, and then return to step 2. If the attitude deviation angle is less than the coarse alignment threshold and not less than the fine alignment threshold, a fine alignment strategy based on the feature points of the laser source is executed, generating control commands to drive the two-dimensional turntable to rotate until the attitude deviation angle is less than the fine alignment threshold, at which point the alignment is completed and the system enters the holding state.
[0008] Furthermore, based on the underwater dual-end alignment method based on visual servo control, this invention also proposes a corresponding underwater dual-end alignment system based on visual servo control, the technical solution of which is as follows: An underwater dual-end alignment system based on visual servo control includes a two-dimensional turntable, an alignment plane array, and a watertight box; A two-dimensional turntable is used to provide pitch and yaw motion; The alignment plane array is set on the top of the two-dimensional turntable, and the alignment plane array rotates synchronously with the two-dimensional turntable; The alignment array integrates an underwater camera, an LED light source, and multiple laser light sources. The underwater camera is used to acquire underwater images, the LED light source is used for coarse alignment, and the laser light source is used for fine alignment. The watertight box is located at the bottom of the two-dimensional turntable, and the control board and power supply are encapsulated inside the watertight box; The control board contains a readable storage medium, and when the readable storage medium is executed, it is used to implement the steps of the underwater dual-end alignment method based on visual servo control described above. The watertight box is also equipped with a communication interface for connecting to the alignment plane array and the two-dimensional turntable via watertight cables.
[0009] Furthermore, based on the aforementioned underwater dual-end alignment method based on visual servo control, this invention also proposes a computer-readable storage medium storing a program thereon; when executed by a processor, this program is used to implement the steps of the aforementioned underwater dual-end alignment method based on visual servo control.
[0010] The present invention has the following advantages: As described above, this invention discloses an underwater dual-end alignment method based on visual servo control. This method addresses the problems of low alignment accuracy, poor efficiency, and weak adaptability of dual-end devices in deep-sea environments. It employs a dual-end symmetrical architecture underwater dual-end alignment system and a two-stage alignment strategy. By combining an LED coarse alignment strategy and a laser fine alignment strategy, it significantly improves alignment efficiency while ensuring high alignment accuracy, thus solving the problem of low efficiency in traditional methods. Furthermore, this invention designs an image enhancement process specifically for underwater environments, effectively overcoming underwater image degradation problems, improving the accuracy and robustness of feature detection, and enabling adaptation to low-visibility deep-sea environments. In addition, this invention uses an attitude estimation method based on ORB features and the P3P algorithm, combined with incremental PID feedback control, to achieve real-time, stable, and high-precision alignment, thereby meeting the real-time requirements of engineering applications. Attached Figure Description
[0011] Figure 1 This is a flowchart of an underwater dual-end alignment method based on visual servo control in an embodiment of the present invention.
[0012] Figure 2 This is a schematic diagram of the underwater dual-end alignment system based on visual servo control in an embodiment of the present invention.
[0013] Figure 3 This is a schematic diagram of the alignment planar array in an embodiment of the present invention.
[0014] Figure 4 This is a flowchart of the visual alignment process in an embodiment of the present invention.
[0015] Figure 5 This is a flowchart of the two-stage alignment strategy in an embodiment of the present invention.
[0016] Among them, 100-Alignment Subsystem A, which includes: 101-Gimbal 1, 102-Alignment Planar Array 1, 103-Watertight Box 1.
[0017] 200-Alignment Subsystem B, which includes: 201-Gimbal II, 202-Alignment Planar Array II, 203-Watertight Box II.
[0018] 10 - Underwater camera, 20 - Laser light source, 30 - LED light source. Detailed Implementation
[0019] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: Example 1 like Figure 1 and Figure 4 As shown, an underwater dual-end alignment method based on visual servo control specifically includes the following steps: Step 1. Start the system and enter standby mode, then turn on the LED and laser light sources for the current plane and the target plane.
[0020] Step 2. Image preprocessing: The underwater camera acquires raw images of the LED and laser light sources that include the target plane. The raw images acquired by the underwater camera are then enhanced to improve image quality, resulting in an enhanced image.
[0021] In step 2 of this embodiment, the process of enhancing the original image acquired by the underwater camera is as follows: Step 2.1. Denoising based on nonlocal mean filtering.
[0022] Unlike traditional median and Gaussian filtering, this invention uses non-local mean filtering (NLM) to denoise the raw underwater images captured by the underwater camera, preserving the edge features of LED and laser spots and avoiding the loss of key features. Its core formula is: .
[0023] in, Indicates the target pixel coordinates. This represents the new pixel after NLM estimation. The normalization coefficient is... Indicates the pixel coordinates to be compared. For the search field, For pixel similarity weights, Indicates the underwater image in pixels The pixel value at that location.
[0024] Step 2.2. Color correction processing based on the prior of the underwater dark channel.
[0025] To address the severe red light attenuation and blue-green light color shift issues, the dark channel prior algorithm is improved for underwater environment adaptation. Its core formula is: .
[0026] in, This represents the image after color correction processing. As global background light, This is a transmission image. This is the minimum transmittance threshold.
[0027] Step 2.3. Contrast enhancement processing based on contrast-limited adaptive histogram equalization.
[0028] This invention employs block equalization and contrast threshold limitation, which can avoid noise amplification in uniform areas and improve the light source recognition under low illumination.
[0029] Step 2.3.1. Segment the image.
[0030] The image obtained after color correction in step 2.2 Divided into Local image region .
[0031] in, , .
[0032] Step 2.3.2. Calculate the local histogram.
[0033] For each local image region Calculate its histogram ,in This indicates the pixel grayscale level, typically ranging from 0 to 255.
[0034] Step 2.3.3. Contrast Limitation.
[0035] For each histogram Crop out items in the histogram that exceed a preset threshold. The part where the threshold The settings can be made based on the results and prior knowledge.
[0036] Total number of pixels cropped Evenly distributed across all histogram entries: .
[0037] in, Indicates the first Line 1 The total number of pixels with all gray levels cropped out within a single local image region of a column. Indicates the first The number of remaining pixels for each gray level; This indicates the total number of gray levels in the image; for example, an 8-bit image has a value of 256. This represents the histogram after clipping and equalization.
[0038] Step 2.3.4. Calculate the cumulative distribution function (CDF).
[0039] For each clipped and equalized histogram Calculate its CDF: .
[0040] in, Indicates the first Line 1 The cumulative distribution function values of gray levels from 0 to k in a local image region; Indicates the first Line 1 In the nth local image region, the th The number of pixels in the equalized histogram corresponding to each gray level. .
[0041] Step 2.3.5. Remap grayscale values.
[0042] CDF is used to map the grayscale value of each pixel within a local image region to a new value: .
[0043] in, This represents the pixel grayscale level index, with a value ranging from 0 to... , Indicates the total number of gray levels; The output grayscale value is obtained after mapping the original grayscale level to CLAHE. This represents the cumulative distribution function value corresponding to the original gray level s within the current local image region; It is the minimum non-zero CDF value in the cumulative distribution function of the current single local image region, and each image region is calculated independently; This represents the total number of pixels within a single local image region.
[0044] Step 2.3.6. Bilinear interpolation.
[0045] For the boundaries of each local image region, bilinear interpolation is used to smooth pixel values between adjacent local image regions to avoid blocky effects. The bilinear interpolation formula is as follows: .
[0046] in, This indicates the coordinates after bilinear interpolation smoothing. The pixel grayscale value output at the location; , , These are the interpolation coefficients calculated based on neighboring pixel values. This represents the bilinear interpolation constant offset coefficient, used to characterize the local pixel grayscale reference offset.
[0047] Step 2.4. Defogging treatment.
[0048] This invention estimates the transmission map and background light using green and blue dark channels to eliminate the fogging effect caused by suspended particles in water.
[0049] Step 2.4.1. Estimation of hierarchical fogging light.
[0050] Feature extraction and region ranking: First, feature extraction is performed on the image obtained by contrast enhancement processing in step 2.3, including color channel changes, texture, and edge information. In this embodiment, the extracted quantized features specifically include the color channel variance, texture roughness, and edge gradient magnitude of each local image region.
[0051] The process of extracting quantized features specifically includes: Obtaining the color channel variance RGB: Calculate the variance of the grayscale values of the red, green, and blue channels of each local image region. Take the mean of the variances of the grayscale values of the red, green, and blue channels as the color fluctuation quantization value of the local image region. The color fluctuation quantization value of the local image region is the color channel variance of each local image region, which is used to reflect the degree of color change in the local image region.
[0052] TEX (Texture Roughness): The texture roughness of each local image region is calculated based on the statistical difference of pixel gray levels within each local image region. The larger the texture roughness value, the more complex the texture details of the local image region.
[0053] Obtaining the edge gradient magnitude (EDGE): The gradient operator is used to solve for the gradient magnitude of all pixels in each local image region. The average of the obtained gradient magnitudes is taken and used as the edge gradient magnitude of each local image region to characterize the richness of edge information in the local image region.
[0054] The three quantization features of color channel variance, texture roughness, and edge gradient magnitude of each local image region are normalized, and the values of each quantization feature are uniformly mapped to the [0,1] interval. A weighted fusion method is then used to construct a comprehensive regional evaluation index. : .
[0055] in, This represents the normalized variance of the color channels. This represents the normalized texture roughness. This represents the normalized edge gradient magnitude. , , For the preset weighting coefficients, satisfy Preset weighting coefficients , , The values can be tuned according to the characteristics of the underwater imaging scene to ensure that the three types of quantization features contribute weights reasonably.
[0056] Based on regional comprehensive evaluation indicators The entire image is quantified and scored for all local image regions, and then sorted hierarchically from high to low scores.
[0057] Selecting veiling light candidate regions: Veiling light is a type of fogging light. In this embodiment, the region most likely to be a veiling light is selected by ranking, while avoiding regions containing bright objects (such as bubbles) as much as possible. The ranking obtained by hierarchical sorting distinguishes between bright objects and veiling light regions. For example, a high edge gradient amplitude indicates rich edge information in the region, and a large grayscale variance in the color channel indicates significant color fluctuation in the region. Based on this, interference from bright foreground objects is eliminated, and the distribution region of the veiling light is accurately located, resulting in a hierarchical region division result.
[0058] The bright objects include the foreground LEDs and the laser-lit targets, while the atomized light curtain area is the background area.
[0059] In this embodiment, the bright object specifically refers to a high-brightness target in the foreground, such as a bubble, an LED spot, or a laser spot. The area of the high-brightness target in the foreground does not belong to the area of the fogged light curtain formed by water scattering.
[0060] Step 2.4.2. Transmission estimation and adjustment.
[0061] Transmission map creation: Using a method based on the green and blue dark channels to create a transmission map helps reduce inaccurate estimates caused by insufficient information in the underwater red channel.
[0062] To prevent oversaturation and artifacts: The transmission map is adjusted by setting a lower limit threshold constraint on transmittance to limit the transmission map values to the physically valid range [0,1], and local neighborhood smoothing filtering is used to suppress spurious noise. This prevents oversaturation (i.e., pixel values exceeding the physical range [0,1]) and artifacts in the background area from appearing in the dehazed image. These artifacts may be caused by low transmission values in the background area.
[0063] Step 2.4.3. Superpixel segmentation and clustering.
[0064] Background region localization: Based on the hierarchical region segmentation results obtained in step 2.4.1, the background region of the image is further located through superpixel segmentation and clustering techniques. That is, superpixel segmentation is used to divide the image into several superpixel units with homogeneous color and texture, and then clustering algorithm is used to achieve accurate differentiation and localization of the background region from the foreground LED and laser bright target regions.
[0065] Transmission value adaptation: Without changing the transmission parameters of the foreground target, the transmission values of the identified background area are locally adaptively fine-tuned based on the transmission map created in step 2.4.2 to correct the background transmission estimation deviation and avoid color distortion and unnatural visual defects in the background area after dehazing. Finally, the enhanced image is output.
[0066] Step 3. Feature Detection: Identify and locate feature points of LED and laser light sources from the enhanced image.
[0067] In step 3 of this embodiment, the enhanced image is initially screened for feature points using a combination of brightness detection and morphological processing. Then, the ORB algorithm is used to locate and describe the feature points. The Hamming distance metric and RANSAC algorithm are used to complete feature matching and mismatch removal. Finally, the pixel coordinates of 3 laser points and 1 LED point are obtained, where the laser points are the feature points of the laser light source and the LED points are the feature points of the LED light source.
[0068] Specifically, the method combining brightness detection and morphological processing employs adaptive brightness thresholding and morphological opening operations to perform initial screening of underwater high-brightness spots, including laser points and LED points, eliminating underwater particle noise, suspended impurities, and false bright spots.
[0069] The adaptive brightness threshold preferably employs either Mean Adaptive Thresholding or Gaussian Adaptive Thresholding to adaptively calculate the image bright spot segmentation threshold. The morphological opening operation uses a morphological method of erosion followed by dilation to preferentially filter out isolated small noise points and suspended particle pseudo-bright spots smaller than the structuring element, while preserving the overall morphology of the real LED and laser spots. The closing operation uses dilation followed by erosion, mainly used to fill small voids inside the spot and repair incomplete edges of the target area. In the initial feature screening stage of this invention, morphological opening operations are mainly used to accurately remove interfering pseudo-feature points caused by underwater particle noise and suspended impurities while preserving the features of the real high-brightness spots.
[0070] The ORB algorithm was used to construct an underwater image scale pyramid, and the FAST corner detection threshold was optimized to adapt to low-contrast underwater light spots and ensure stable feature point extraction.
[0071] In this embodiment, a multi-scale image pyramid is first constructed on the enhanced underwater image, and feature detection is carried out simultaneously at different resolution levels to ensure that the light spot can still be stably captured under changes in distance and scale.
[0072] To address the characteristics of low contrast and blurred light spot edges in underwater images, the FAST corner detection and discrimination threshold is adaptively optimized. The invalid edge response threshold is lowered, the response conditions of high-brightness light spot corners are strengthened, invalid corners generated by water texture and weak edges are suppressed, and the stable feature positions of LED and laser light spots are locked first, so as to achieve reliable and stable feature point extraction in low-contrast underwater environments.
[0073] Hamming distance matching binary descriptors are used, and the RANSAC algorithm is combined to remove mismatches caused by underwater dynamic noise. Finally, the pixel coordinates of 3 laser points and 1 LED point are output, which meets the requirements of subsequent P3P solution.
[0074] Step 4. Planar pose estimation: Calculate the pose deviation angle of the current plane relative to the target plane based on the pixel coordinates, i.e., the image coordinates, of the feature points of the LED light source and the laser light source.
[0075] Planar attitude estimation can be used to calculate the deviation of the camera plane (current plane) relative to the target plane. This deviation includes translational and rotational deviations. Translational deviation refers to the linear displacement of the camera plane relative to the target plane, while rotational deviation refers to the angular deflection of the camera plane relative to the target plane. The translational deviation is ensured by the initial placement and installation spacing of the equipment. In this embodiment, a two-dimensional turntable with pitch and yaw degrees of freedom is used to adjust the rotational deviation of the current plane relative to the target plane. The translational deviation is not included in the visual servo closed-loop alignment adjustment range.
[0076] In this embodiment, the attitude deviation angles of the current plane relative to the target plane include pitch angle and yaw angle.
[0077] The feature point localization problem is transformed into a P3P problem. Based on the P3P algorithm, the rotation matrix is solved using the world coordinates and image coordinates of the three laser points. The LED points are used as constraints to eliminate ambiguity. The unique correct solution is selected from multiple solutions, and then the pitch and yaw angles are obtained.
[0078] In this embodiment, step 4 specifically involves: The positions of the laser points on the target plane are fixed, and the positions of the three laser points are as follows: , , , , , This represents the horizontal coordinates of the three laser points in the world coordinate system. , , This represents the horizontal and vertical coordinates of the three laser points in the world coordinate system. , , This represents the vertical depth coordinates of the three laser points in the world coordinate system; assuming the three laser points are evenly distributed on the same alignment plane. , , It is 0.
[0079] For global image coordinates The point is converted into a point in the normalized camera coordinate system. : .
[0080] in, This represents the normalized depth scale of the camera's normalized coordinate system. The value is fixed at 1, and the camera normalized coordinate system is the same as the normalized camera coordinate system; For the camera intrinsic parameter matrix, Represents the column coordinates of image pixels. Represents the row coordinates of image pixels.
[0081] According to the camera imaging principle, the first laser points The world coordinates and normalized camera coordinates have the following relationship: .
[0082] in, ; The scaling factor is unknown. Indicates the first The horizontal and vertical coordinates of each laser point in the camera's normalized coordinate system That is, the first Standardized camera coordinates for each laser point; It is a rotation matrix; For the first The world coordinates of a laser point; This is the translation vector. The equation represents the world coordinate point, i.e., the laser point. After rotation and translation, it should fall on the point passing through the image. On the rays.
[0083] This invention uses Zhang's calibration method to complete camera calibration and obtain the intrinsic parameter matrix. and distortion coefficients, as well as extrinsic parameters including the rotation matrix. Translation vector Establish global image coordinates with world coordinates The mapping relationship.
[0084] Zhang's calibration method is a camera calibration algorithm that estimates the camera's intrinsic parameters (such as focal length, principal point, radial distortion, etc.) and extrinsic parameters (i.e., the camera's position and orientation relative to the checkerboard) from a series of photographs of a checkerboard pattern. The core of Zhang's calibration method is to use the known two-dimensional planar pattern of the checkerboard and infer the camera's parameters by observing the pattern's projection onto the camera. Since the actual dimensions of the checkerboard are known, it can be used as a reference for the world coordinate system to estimate the camera's intrinsic and extrinsic parameters.
[0085] The algorithm steps of Zhang's calibration method are as follows: Image acquisition: Take multiple images of the chessboard pattern from different angles.
[0086] Corner detection: Detect the corners of the checkerboard pattern in each image, where the image coordinates of these corners are known.
[0087] Corner point correspondence: For each corner point of the chessboard, the coordinates in the world coordinate system are... and coordinates in the image coordinate system .
[0088] Camera model: The camera model is represented as: .
[0089] Same-order coordinates: due to Since points on the chessboard are always 0 (the chessboard is a plane), their coordinates can be expressed in their order of 0: .
[0090] Solving for intrinsic and extrinsic parameters: The extrinsic parameters can be solved first by utilizing the linear part of the equation system. Then, the extrinsic parameters are used to solve for the intrinsic parameters. .
[0091] Since there are three laser points, the problem is a P3P problem. For P3P problems with three equations, the method of this invention does not directly solve the linear equation system, but instead seeks a rotation matrix that satisfies all three points. Translation vector This also satisfies the geometric relationships of their projections onto the image plane. This typically involves solving one or more polynomial equations derived from the distance relationships between points, since the distance between points in world coordinates should be equal to the distance along their respective ray directions.
[0092] Typically, the P3P problem has up to four solutions. The method of this invention uses the LED point as the fourth point, which can generate solutions for each problem. and This is applied to this point to check for solutions that match the projection of the fourth point with the actual image coordinates. Optimization algorithms, including RANSAC, can also be used in this process to eliminate outliers, ultimately yielding the rotation matrix. The only correct solution.
[0093] The first World coordinates of a laser point With global image coordinates Substituting the camera model into the problem transforms it into a P3P problem, and at most 4 candidate solutions for the rotation matrix R can be obtained.
[0094] Substitute the world coordinates of the LED points as constraints to calculate the reprojection error.
[0095] The candidate solution of the rotation matrix R with the smallest reprojection error is selected as the only correct solution of the rotation matrix R.
[0096] This invention optimizes camera parameters by minimizing the reprojection error, which is the difference between the actual detected corner position and the corner position calculated based on the current camera parameters. The minimization process can utilize the Levenberg-Marquardt algorithm or other nonlinear optimization methods.
[0097] By decomposing the unique correct solution of the rotation matrix R, we obtain the attitude deviations of the current plane relative to the target plane, namely the pitch and yaw angles, i.e., the rotation angles.
[0098] Step 5. Feedback Control: Generate control commands, i.e., output control signals, based on the attitude deviation angle of the current plane relative to the target plane.
[0099] If the attitude deviation angle is not less than the coarse alignment threshold, then a coarse alignment strategy based on the feature points of the LED light source is executed, and control commands are generated to drive the two-dimensional turntable to rotate until coarse alignment is completed, and then the process returns to step 2.
[0100] If the attitude deviation angle is less than the coarse alignment threshold and not less than the fine alignment threshold, a fine alignment strategy based on the feature points of the laser source is executed, generating control commands to drive the two-dimensional turntable to rotate until the attitude deviation angle is less than the fine alignment threshold, at which point the alignment is completed and the system enters the holding state.
[0101] In this embodiment, the feedback control process uses the PID method. First, it is necessary to determine three parameters of the PID controller: proportional gain. Integral gain Differential gain The determination of these parameters usually requires adjustment based on the dynamic characteristics of the system, or empirical rules such as the Ziegler-Nichols method can be used to initially set them, and then adjusted according to the actual situation of the two-dimensional turntable.
[0102] In this embodiment, a dual closed-loop control strategy is implemented in step 5, such as... Figure 5 As shown, the system employs a two-stage alignment strategy. In the coarse alignment stage, the outer loop controller, based on LED point features, quickly reduces the deviation from its initial value (potentially greater than 10°) to within 5°. Upon entering the fine alignment stage, the inner loop controller switches to laser point features, further converging the deviation to within 1°, thus achieving high-precision alignment.
[0103] The outer ring performs coarse alignment based on LED dot features, while the inner ring performs fine alignment based on laser dot features. Both the outer and inner ring controllers use incremental PID controllers to generate control commands, which drive the two-dimensional turntable, thus controlling the gimbal rotation.
[0104] For each control cycle Calculate the output of the PID controller , These are control commands, such as the motor's drive voltage or drive current, which can be expressed as: .
[0105] in, Represents continuous-time variables of the system; Indicates proportional gain; This indicates the current attitude deviation error, which is the rotational deviation in the two rotational degrees of freedom: pitch angle and yaw angle. Indicates integral gain; This represents the attitude deviation error at any time within the integration interval; This represents the differential gain.
[0106] In discrete control systems, the formula is discretized as follows: .
[0107] in, Indicates the first The PID output control quantity for each control cycle. Indicates the first Current attitude deviation error in each control cycle This indicates the number of cycles from the initial time to the current time. Historical posture deviation value of the step This represents the historical attitude deviation value from the previous control cycle.
[0108] The output of the PID controller The signal is converted into an input signal for the turntable and sent to the 2D turntable. The 2D turntable performs a rotational motion and monitors the actual position via a feedback system. Adjustments are made based on visual feedback until alignment is achieved.
[0109] In step 5 of this embodiment, the process of executing the coarse alignment strategy based on the feature points of the LED light source is as follows: The coarse alignment threshold is 5°. If the attitude deviation of the current plane relative to the target plane in pitch or yaw angle is not less than 5°, the outer loop coarse alignment is triggered. Alignment is performed using LED points as features. The outer loop controller quickly drives the two-dimensional turntable to perform large-scale rapid alignment to shorten the alignment time. After the attitude deviation of the current plane relative to the target plane in pitch and yaw angle is converged to within 5°, the coarse alignment is completed, and the process returns to step 2.
[0110] The process of implementing a fine alignment strategy based on feature points from a laser source is as follows: The fine alignment threshold is 1°. If the attitude deviation of the current plane relative to the target plane in pitch or yaw angle is not less than 1°, and the attitude deviation of the current plane relative to the target plane in pitch and yaw angle is less than 5°, then the inner loop fine alignment is triggered. Alignment is performed using laser points as features. The inner loop controller drives the two-dimensional turntable to move, improves the control gain, and finely adjusts the attitude of the two-dimensional turntable until the attitude deviation of pitch and yaw angle is less than 1°. Fine alignment is then completed, achieving high-precision alignment, and the system enters the holding state.
[0111] Example 2 This embodiment 2 describes an underwater dual-end alignment system based on visual servo control, which is based on the same inventive concept as the underwater dual-end alignment method based on visual servo control in embodiment 1.
[0112] Specifically, such as Figure 2 and Figure 3 As shown, the underwater dual-end alignment system based on visual servo control includes an alignment subsystem A100 and an alignment subsystem B200. Taking the alignment subsystem A100 as an example, it includes a two-dimensional turntable, an alignment plane array 102, and a watertight box 103.
[0113] The two-dimensional turntable is used to provide pitch and yaw motion. In this embodiment, the two-dimensional turntable has two degrees of freedom, pitch and yaw, and the repeatability is better than 0.01°. The two-dimensional turntable in the alignment subsystem A100 is the gimbal-101, and the gimbal-101 uses the Whale YT12061 underwater actuator.
[0114] Alignment plane array 102 is set on the top of the two-dimensional turntable, and alignment plane array 102 rotates synchronously with the two-dimensional turntable.
[0115] The alignment array 102 integrates an underwater camera 10, an LED light source 30, and multiple laser light sources 20. The underwater camera 10 is used to acquire underwater images. In this embodiment, the underwater camera 10 is preferably a high-sensitivity coaxial camera module with a minimum illumination of 0.01 Lux. The LED light source 30 is used for coarse alignment. In this embodiment, the LED light source 30 is preferably a wide-angle underwater-specific LED with a light emission angle greater than 100°. The laser light sources 20 are used for fine alignment. In this embodiment, the laser light source 20 is preferably a blue-green laser with a wavelength of 532 nm and a beam divergence angle of less than 0.5 mrad.
[0116] Specifically, in this embodiment, the underwater camera 10 uses a Hisilicon GC2053 series 2-megapixel coaxial camera module, supports AHD high-definition output, has a minimum illumination of 0.01 Lux, and is equipped with a specially designed underwater lens. The LED light source 30 is preferably a LUXUS Compact underwater-specific LED light source, used in the coarse alignment stage to provide wide-range illumination; the laser light source 20 is preferably a 532nm blue-green laser light source, used in the fine alignment stage to provide high-precision indication.
[0117] In this embodiment, the alignment plane array 102 adopts a coaxial center layout.
[0118] The underwater camera 10 is located at the geometric center of the alignment plane array 102 and coaxially acquires light spot images.
[0119] LED light source 30 is located next to the center of underwater camera 10, providing wide-angle light and covering a large field of view for coarse alignment.
[0120] Three laser light sources 20 are evenly distributed in an equilateral triangle around the underwater camera 10, with parallel beams and high collimation, for precise alignment.
[0121] The watertight box 103 is located at the bottom of the two-dimensional turntable. The watertight box 103 contains a control board and a power supply.
[0122] The control board contains a readable storage medium, which, when executed, is used to implement the steps of the underwater dual-end alignment method based on visual servo control described in Embodiment 1. In this embodiment, the control board uses an embedded processor as its core, integrating video encoding / decoding, network communication, and peripheral control interfaces; specifically, the control board has a main frequency of 1.5GHz and integrates interfaces such as H.264 video encoding / decoding, Gigabit Ethernet, and UART / RS485, responsible for running the core algorithm and controlling the two-dimensional turntable, while also completing the interaction with the alignment planar array.
[0123] The watertight box 103 is also equipped with a communication interface for connecting to the alignment plane array 102 and the two-dimensional turntable via watertight cables.
[0124] Similarly, alignment subsystem B200 includes a two-dimensional turntable, alignment planar array 202, and watertight box 203, wherein the two-dimensional turntable is also known as gimbal 201. Alignment subsystem B200 has the same structure and function as alignment subsystem A100, and will not be described again here.
[0125] The two subsystems operate independently but work together to complete the automatic alignment of the two-end planes.
[0126] In this embodiment, the underwater dual-end alignment system based on visual servo control operates at a depth of no less than 100 meters, an alignment distance of no less than 10 meters, an alignment accuracy better than 1.0°, and a single alignment time of no more than 20 seconds.
[0127] The hardware system of this invention adopts a modular and pressure-resistant design. Key components have been selected and optimized, enabling stable operation in harsh environments such as water depths of up to 100 meters, demonstrating high reliability. Specifically, the pressure resistance is divided into structural pressure resistance and component pressure resistance. The structure uses a 5mm thick stainless steel shell with a watertight design, capable of withstanding 1MPa (100 meters water depth) pressure. All components are non-cavity devices; the underwater camera 10, LED light source 30, laser light source 20, and 2D turntable are all underwater-specific components. The design was verified through pressure testing after completion.
[0128] This invention effectively solves the problems of image degradation, low alignment accuracy, and poor efficiency in deep-sea environments, achieving an alignment accuracy better than 1.0° at a water depth of 100 meters and rapid alignment within 20 seconds. It is suitable for underwater docking, pipeline maintenance, and other scenarios. In a simulated deep-sea docking experiment, this system was deployed on the end effectors of two underwater robots. Initially, the two robots were 10 meters apart with an attitude deviation of approximately 8°. After system startup, it first entered a coarse alignment stage, taking approximately 9 seconds, during which the deviation was reduced to 3°. Then, it entered a fine alignment stage, taking approximately 6 seconds, ultimately stabilizing within a deviation range of 0.8°. The total alignment time was 15 seconds, meeting the design specification of no more than 20 seconds, and the entire process required no manual intervention.
[0129] Example 3 This embodiment 3 describes a computer-readable storage medium storing a program that, when executed by a processor, implements the steps of an underwater dual-end alignment method based on visual servo control.
[0130] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc.
[0131] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. It should be noted that any equivalent substitutions or obvious modifications made by those skilled in the art under the guidance of this specification fall within the scope of this specification and should be protected by the present invention.
Claims
1. An underwater dual-end alignment method based on visual servo control, characterized in that, Includes the following steps: Step 1. Start the system and enter standby mode, then turn on the LED and laser light sources for the current plane and the target plane; Step 2. The underwater camera acquires the original images of the LED and laser light sources that include the target plane. The original images acquired by the underwater camera are then enhanced to obtain the enhanced images. Step 3. Identify and locate feature points of LED and laser light sources from the enhanced image; Step 4. Calculate the attitude deviation angle of the current plane relative to the target plane based on the pixel coordinates of the feature points of the LED light source and the laser light source; Step 5. Generate control commands based on the attitude deviation angle of the current plane relative to the target plane; If the attitude deviation angle is not less than the coarse alignment threshold, then the coarse alignment strategy based on the feature points of the LED light source is executed, and control commands are generated to drive the two-dimensional turntable to rotate until the coarse alignment is completed, and then return to step 2. If the attitude deviation angle is less than the coarse alignment threshold and not less than the fine alignment threshold, a fine alignment strategy based on the feature points of the laser source is executed, generating control commands to drive the two-dimensional turntable to rotate until the attitude deviation angle is less than the fine alignment threshold, at which point the alignment is completed and the system enters the holding state.
2. The underwater dual-end alignment method based on visual servo control according to claim 1, characterized in that, In step 2, the process of enhancing the original images captured by the underwater camera specifically involves: Step 2.
1. Denoising based on nonlocal mean filtering; Non-local mean filtering (NLM) is used to denoise the raw underwater images acquired by the underwater camera, in order to preserve the edge features of the LED and laser spots. ; in, Indicates the target pixel coordinates. This represents the new pixel after NLM estimation. The normalization coefficient is... Indicates the pixel coordinates to be compared. For the search field, For pixel similarity weights, Indicates the underwater image in pixels Pixel value at; Step 2.
2. Color correction processing based on prior knowledge of underwater dark channels; To address the issues of red light attenuation and blue-green light color shift, the dark channel prior algorithm is improved for underwater environment adaptation. ; in, This represents the image after color correction processing. As global background light, This is a transmission image. The minimum transmittance threshold; Step 2.
3. Contrast enhancement processing based on contrast-limited adaptive histogram equalization; Step 2.3.
1. Segment the image; The image obtained after color correction in step 2.2 Divided into Local image region ; in, , ; Step 2.3.
2. Local histogram calculation; For each local image region Calculate its histogram ,in Indicates pixel grayscale level; Step 2.3.
3. Contrast Limitation; For each histogram Crop out items in the histogram that exceed a preset threshold. Part of; Total number of pixels cropped Evenly distributed across all histogram entries: ; in, Indicates the first Line 1 The total number of pixels with all gray levels cropped out in a single local image region of a column; Indicates the first The number of remaining pixels for each gray level; This represents the total number of gray levels in the image; This represents the histogram after clipping and equalization. Step 2.3.
4. Calculate the cumulative distribution function (CDF); For each clipped and equalized histogram Calculate its CDF: ; in, Indicates the first Line 1 The cumulative distribution function values of gray levels from 0 to k in a local image region; Indicates the first Line 1 In the nth local image region, the th The number of pixels in the equalized histogram corresponding to each gray level. ; Step 2.3.
5. Remap grayscale values; CDF is used to map the grayscale value of each pixel within a local image region to a new value: ; in, This represents the pixel grayscale level index, with a value ranging from 0 to... , Indicates the total number of gray levels; The output grayscale value is obtained after mapping the original grayscale level to CLAHE. This represents the cumulative distribution function value corresponding to the original gray level s within the current local image region; It is the minimum non-zero CDF value in the cumulative distribution function of the current single local image region; This represents the total number of pixels within a single local image region. Step 2.3.
6. Bilinear interpolation; For the boundary of each local image region, bilinear interpolation is used to smooth pixel values between adjacent local image regions. The bilinear interpolation formula is as follows: ; in, This indicates the coordinates after bilinear interpolation smoothing. The pixel grayscale value output at the location; , , These are the interpolation coefficients calculated based on neighboring pixel values. This represents the bilinear interpolation constant offset coefficient, used to characterize the local pixel grayscale reference offset; Step 2.
4. Defogging treatment; Step 2.4.
1. Estimation of hierarchical fogging light; Feature extraction is performed on the image obtained from the contrast enhancement process in step 2.
3. The extracted quantized features include the color channel variance, texture roughness, and edge gradient magnitude of each local image region. Calculate the variance of the gray values of the red, green and blue channels of each local image region, and take the mean of the variance of the gray values of the red, green and blue channels as the color fluctuation quantization value of the local image region. The color fluctuation quantization value of the local image region is the color channel variance of each local image region. The texture roughness of each local image region is calculated based on the statistical difference of pixel gray levels within each local image region. The gradient operator is used to solve for the gradient magnitude of all pixels in each local image region. The average value of the solved gradient magnitude is taken and used as the edge gradient magnitude of each local image region. The three quantization features of color channel variance, texture roughness, and edge gradient magnitude of each local image region are normalized, and a weighted fusion method is used to construct a comprehensive regional evaluation index. : ; in, This represents the normalized variance of the color channels. This represents the normalized texture roughness. This represents the normalized edge gradient magnitude. , , Let be the weighting coefficient, satisfying ; Based on regional comprehensive evaluation indicators The entire image is quantified and scored, and then sorted hierarchically from high to low scores. The ranking obtained by the hierarchical sorting distinguishes between bright objects and foggy light curtain areas, resulting in hierarchical region division results. Bright objects include foreground LEDs and laser bright targets, while foggy light curtain areas are the background areas. Step 2.4.
2. Transmission estimation and adjustment; The transmission map is created using a method based on the green, blue and dark channels. The transmission map is then adjusted by setting a lower limit threshold constraint on transmittance to restrict the transmission map values to the range [0,1]. Local neighborhood smoothing filtering is also used to suppress glitch noise. Step 2.4.
3. Superpixel segmentation and clustering; Based on the hierarchical region segmentation results obtained in step 2.4.1, the background region of the image is further located by superpixel segmentation and clustering techniques. That is, superpixel segmentation is used to divide the image into several superpixel units with homogeneous color and texture, and then clustering algorithm is used to distinguish and locate the background region from the foreground LED and laser bright target regions. Without changing the transmission parameters of the foreground target, the transmission values of the identified background region are adjusted based on the transmission map created in step 2.4.2 to correct the background transmission estimation bias, and finally the enhanced image is output.
3. The underwater dual-end alignment method based on visual servo control according to claim 2, characterized in that, In step 3, the enhanced image is initially screened for feature points using a combination of brightness detection and morphological processing. Then, the ORB algorithm is used to locate and describe the feature points. The Hamming distance metric and RANSAC algorithm are used to complete feature matching and mismatch removal. Finally, the pixel coordinates of 3 laser points and 1 LED point are obtained, where the laser points are the feature points of the laser light source and the LED points are the feature points of the LED light source.
4. The underwater dual-end alignment method based on visual servo control according to claim 3, characterized in that, In step 3, the method combining brightness detection and morphological processing uses adaptive brightness threshold and morphological opening operation to perform initial screening of underwater high-brightness spots, including laser points and LED points, and remove underwater particle noise, suspended impurities, and false bright spots. The ORB algorithm was used to construct an underwater image scale pyramid and optimize the FAST corner detection threshold to adapt to low-contrast underwater light spots and ensure stable feature point extraction. Hamming distance matching binary descriptors are used, and the RANSAC algorithm is combined to remove mismatches caused by underwater dynamic noise. Finally, the pixel coordinates of 3 laser points and 1 LED point are output, which meets the requirements of subsequent P3P solution.
5. The underwater dual-end alignment method based on visual servo control according to claim 4, characterized in that, In step 4, the attitude deviation angle of the current plane relative to the target plane includes pitch angle and yaw angle; The feature point localization problem is transformed into a P3P problem. The rotation matrix is solved using the world coordinates and image coordinates of the three laser points. The LED points are used as constraints to select the unique correct solution from multiple solutions, and then the pitch and yaw angles are obtained by decomposition.
6. The underwater dual-end alignment method based on visual servo control according to claim 5, characterized in that, Step 4 specifically involves: The positions of the laser points on the target plane are fixed, and the positions of the three laser points are as follows: , , , , , This represents the horizontal coordinates of the three laser points in the world coordinate system. , , This represents the horizontal and vertical coordinates of the three laser points in the world coordinate system. , , This represents the vertical depth coordinates of the three laser points in the world coordinate system; assuming the three laser points are evenly distributed on the same alignment plane. , , =0; For global image coordinates The point is converted into a point in the normalized camera coordinate system. : ; in, This represents the normalized depth scale of the camera's normalized coordinate system. Setting it to 1 enables the camera coordinate system to be normalized, i.e., the camera coordinate system to be normalized. For the camera intrinsic parameter matrix, Represents the column coordinates of image pixels. Represents the row coordinates of image pixels; According to the camera imaging principle, the first laser points The world coordinates and normalized camera coordinates have the following relationship: ; in, ; It is a scaling factor; Indicates the first The horizontal and vertical coordinates of each laser point in the camera's normalized coordinate system That is, the first Standardized camera coordinates for each laser point; It is a rotation matrix; For the first The world coordinates of a laser point; It is a translation vector; Camera calibration was performed using Zhang's calibration method to obtain the intrinsic parameter matrix. And distortion coefficients, to establish global image coordinates with world coordinates The mapping relationship; The camera model is represented as: ; The first World coordinates of a laser point With global image coordinates Substituting the camera model, the problem is transformed into a P3P problem, and at most 4 candidate solutions for the rotation matrix R are obtained. Substitute the world coordinates of the LED points as constraints to calculate the reprojection error; The candidate solution of the rotation matrix R with the smallest reprojection error is selected as the only correct solution of the rotation matrix R. By decomposing the unique correct solution of the rotation matrix R, we can obtain the attitude deviations of the pitch and yaw angles of the current plane relative to the target plane.
7. The underwater dual-end alignment method based on visual servo control according to claim 6, characterized in that, In step 5, a dual closed-loop control strategy is implemented. The outer loop performs coarse alignment based on LED points, and the inner loop performs fine alignment based on laser points. Both the outer loop controller and the inner loop controller use incremental PID controllers to generate control commands. For each control cycle Calculate the output of the PID controller Represented as: ; in, Represents continuous-time variables of the system; Indicates proportional gain; This indicates the current attitude deviation error, which is the rotational deviation in the two rotational degrees of freedom: pitch angle and yaw angle. Indicates integral gain; This represents the attitude deviation error at any time within the integration interval; Represents differential gain; In discrete control systems, the formula is discretized as follows: ; in, Indicates the first The PID output control quantity for each control cycle. Indicates the first Current attitude deviation error in each control cycle This indicates the number of cycles from the initial time to the current time. Historical posture deviation value of the step This represents the historical attitude deviation value from the previous control cycle.
8. The underwater dual-end alignment method based on visual servo control according to claim 7, characterized in that, In step 5, the process of executing the coarse alignment strategy based on the feature points of the LED light source is as follows: The coarse alignment threshold is 5°. If the attitude deviation of the current plane relative to the target plane in pitch or yaw angle is not less than 5°, the outer ring coarse alignment is triggered. Alignment is performed using LED points as features. The two-dimensional turntable is driven by the outer loop controller to move and the attitude deviation of the pitch and yaw angles of the current plane relative to the target plane is brought to within 5°. After this, coarse alignment is completed and the process returns to step 2. The process of implementing a fine alignment strategy based on feature points from a laser source is as follows: The precision alignment threshold is 1°. If the attitude deviation of the current plane relative to the target plane in pitch or yaw angle is not less than 1°, and the attitude deviation of the current plane relative to the target plane in pitch and yaw angle is less than 5°, then the inner loop precision alignment is triggered. Alignment is performed using laser points as features. The two-dimensional turntable is driven to move by the inner loop controller until the attitude deviation of pitch and yaw angle is less than 1°. The precision alignment is completed and the system enters the holding state.
9. An underwater dual-end alignment system based on visual servo control, characterized in that, Includes a two-dimensional turntable, an alignment planar array, and a watertight box; The two-dimensional turntable is used to provide pitch and yaw motion; The alignment plane array is disposed on the top of the two-dimensional turntable, and the alignment plane array rotates synchronously with the two-dimensional turntable; The alignment array integrates an underwater camera, an LED light source, and multiple laser light sources. The underwater camera is used to acquire underwater images, the LED light source is used for coarse alignment, and the laser light source is used for fine alignment. The watertight box is located at the bottom of the two-dimensional turntable, and the control board and power supply are encapsulated inside the watertight box; The control board contains a readable storage medium, and when the readable storage medium is executed, it is used to implement the steps of the underwater dual-end alignment method based on visual servo control as described in any one of claims 1 to 8. The watertight box is also equipped with a communication interface for connecting to the alignment plane array and the two-dimensional turntable via watertight cables.
10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the underwater dual-end alignment method based on visual servo control as described in any one of claims 1 to 8.