Maneuvering control and imaging optimization method and system of underwater remote control camera

By combining deep learning networks with physical models, an underwater imaging system was developed that solved the problems of color distortion and contrast reduction in turbid waters. It achieved adaptive color correction and high-definition imaging, and improved the system's energy balance and stability through thrust distribution optimization and attitude stabilization control.

CN121418679APending Publication Date: 2026-01-27BROAD VISION (XIAMEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511642564.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing underwater imaging equipment suffers from color distortion and reduced contrast in turbid or deep water environments. Traditional methods cannot effectively correct wavelength selective attenuation, robot thrusters have uneven energy distribution, and attitude control is unstable under external disturbances.

Method used

By combining deep learning networks and physical models, the attenuation coefficient is dynamically calculated through real-time acquisition of environmental parameters, and color compensation and scattering dehazing are performed. A thrust allocation optimization based on quadratic programming is adopted, and attitude stabilization control of the disturbance observer is introduced.

Benefits of technology

It achieves adaptive color correction and high-definition imaging of underwater images, balanced thruster energy consumption, and improved stability of attitude control in complex environments, making it suitable for embedded platform applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121418679A_ABST
    Figure CN121418679A_ABST
Patent Text Reader

Abstract

The invention provides a maneuvering control and imaging optimization method and system for an underwater remote control camera, and relates to the technical field of underwater camera shooting, and the method comprises the steps: collecting an original image and depth, illumination, color temperature and water body type data through an underwater camera; calculating an RGB component attenuation coefficient based on the depth information and the water body type, extracting environment background light and calculating a white balance parameter; a double-path compensation mechanism is adopted, one path inputs original data into a color compensation network with physical constraints, and the other path performs compensation based on a physical model; and finally, calculating a fusion weight according to image quality parameters, and carrying out weighted fusion on two paths of compensation results to obtain a color correction image. According to the method, the advantages of deep learning and a physical model are combined, and high-quality color correction of the underwater image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to underwater camera technology, and more particularly to a method and system for maneuver control and imaging optimization of an underwater remotely operated camera. Background Technology

[0002] Existing underwater imaging equipment commonly suffers from severe color distortion and contrast reduction in turbid or deep water environments. Due to the significant differences in the absorption coefficients of water for different wavelengths of light, red and yellow bands attenuate considerably after a short underwater propagation distance, resulting in a severe blue-green tint in the captured images, which worsens with depth. Forward and backscattering caused by underwater suspended particles contribute to decreased image contrast and blurred details, creating a white fog effect similar to that seen in foggy weather. Traditional white balance adjustment methods can only perform global color temperature compensation and cannot effectively correct for the wavelength-selective attenuation unique to underwater environments. Existing image enhancement algorithms are mostly designed for terrestrial scenes and perform poorly when directly adapted to underwater environments. Physical model-based methods rely on fixed parameters and cannot adapt to dynamic environments, while purely data-driven deep learning methods lack physical constraints, resulting in insufficient generalization ability and susceptibility to artifacts.

[0003] Existing thrust distribution methods for underwater robots mostly employ fixed pseudo-inverse matrix solutions. When there is redundancy in the number of thrusters, they cannot optimize energy consumption allocation, leading to high overall energy consumption and uneven thruster lifespan. Furthermore, they cannot automatically reconfigure the distribution scheme when a thruster fails, severely impacting system fault tolerance and operational continuity. In addition, the underwater environment presents complex external disturbances, including water flow impacts and eddies. Traditional PID attitude controllers only adjust based on attitude errors and cannot actively compensate for external disturbances. In strong current environments, attitude angle fluctuations are large, and stabilization times are long, making stable hovering difficult. While some advanced control methods offer some robustness, they require precise system model parameters and have high computational complexity, making them unsuitable for embedded platforms. They generally lack online estimation and active compensation mechanisms for external disturbances. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for maneuver control and imaging optimization of underwater remotely operated cameras, which can solve the problems in existing technologies.

[0005] A first aspect of the present invention provides a method for maneuver control and imaging optimization of an underwater remotely operated camera, comprising: Control the underwater camera to collect raw image data, depth information, illumination information, color temperature information, and initial values ​​of water body type; Based on the product of the depth information and the initial value of the water body type, and the correction amount of the illuminance information and color temperature information, the attenuation coefficients of the red, green and blue components in the original image data are calculated respectively. Extract pixel color information whose brightness values ​​are within a preset range from the original image data as ambient background light, and calculate white balance adjustment parameters based on pixel distribution; The original image data, attenuation coefficient, ambient background light, illuminance information, and color temperature information are input into a color compensation network. The color compensation network applies a physical constraint based on the attenuation coefficient in the intermediate layer and outputs a first compensated image. After subtracting the ambient background light from the original image data, it is amplified according to an exponential function of the attenuation coefficient and outputs a second compensated image. The fusion weights are calculated based on the image quality parameters of the original image data. The first compensation image and the second compensation image are then weighted and summed using the fusion weights to obtain the color-corrected image.

[0006] Optionally, the steps of calculating the first compensated image and the second compensated image include: The color compensation network adopts an encoder-decoder structure. A constraint layer is set in the middle layer. The constraint layer calculates the pixel-wise product of the original image data and the white balance adjustment parameter, and then the pixel-wise product of the attenuation coefficient exponential function. The calculation result is weighted and fused with the output of the network front layer and then passed to the subsequent layers to output the first compensated image and the first confidence value. The ambient background light component of the corresponding channel is subtracted from each pixel value of the original image data. The subtraction result is multiplied pixel by pixel with the attenuation coefficient exponential function to obtain a preliminary magnified image. Skin color pixel regions in the preliminary magnified image are identified. The white balance adjustment parameters are adjusted and recalculated according to the deviation between the skin color chroma and the standard skin color chroma to obtain a second compensated image and a second confidence value.

[0007] Optionally, after acquiring the raw image data and before inputting it into the color compensation network, a scattering dehazing process is also included: Spatial frequency transformation is performed on the original image data to extract the low-frequency component as the initial value of the backscatter component, and the initial value of the direct light component is obtained by subtracting the initial value of the backscatter component. The initial values ​​of the backscatter component, the initial values ​​of the direct light component, the depth information, and the illuminance information are input into the backscatter estimation network, and the refined backscatter component and transmittance correction coefficient are output. Acquire image data from multiple frames at adjacent time points, calculate the pixel displacement field, transform the backscattered component of the next frame to the coordinate system of the current frame, and adjust the refined backscattered component to minimize the sum of squared differences between the transformed backscattered component and the backscattered component of the current frame, thus obtaining the time-stable backscattered component. The time-stable backscatter component is subtracted from the original image data and multiplied pixel by pixel with the reciprocal of the transmittance correction coefficient to obtain the dehazed image; The dark channel statistics of the dehazed image are calculated as the residual suspended particle density index. When the residual suspended particle density index is higher than the preset index threshold, the reciprocal value of the transmittance correction coefficient is reduced, and the adjusted image is used as the original image data input to the color compensation network.

[0008] Optionally, the steps for controlling the underwater camera include: Generate the desired velocity and desired angular velocity based on the control commands, and calculate the desired force in three orthogonal directions and the desired torque on three orthogonal axes; Based on the installation position coordinates of multiple thrusters in the body coordinate system and the thrust direction unit vector, calculate the force components of each thruster about the three orthogonal directions and the torque components about the three orthogonal axes, and construct the thrust mapping matrix. Under the conditions of satisfying the equality constraints of the thrust mapping matrix, the desired force and the desired torque, and the upper and lower limits of thrust, find the thrust command vector that minimizes the sum of the squares of thrust. The speed and current feedback of each thruster are detected. When the speed feedback is lower than the preset ratio of the command value or the current feedback exceeds the rated value, the thruster is determined to be faulty. The column corresponding to the faulty thruster is deleted from the thrust mapping matrix and the thrust command vector is re-solved. The elements of the thrust command vector are first-order low-pass filtered and then output to each thruster.

[0009] Optionally, the method further includes: Data from gyroscopes, accelerometers, electronic compasses, and depth sensors are collected and fused using an extended Kalman filter to obtain attitude quaternion estimates and angular rate estimates. The estimated value of the external disturbance torque is calculated based on the body inertia matrix, the rate of change of angular rate, the thruster control torque, and the quadratic term of angular rate. The quaternion error between the desired attitude and the estimated attitude is calculated as the attitude error, the difference between the desired angular rate and the estimated angular rate is calculated as the angular rate error, and the negative value of the sum of the proportional term of the attitude error, the differential term of the angular rate error, and the estimated value of the external disturbance torque is used as the attitude control torque. The attitude control torque is superimposed on the desired torque to solve for the thrust command vector.

[0010] Secondly, it provides a maneuvering control and imaging optimization system for underwater remotely operated cameras, including: The first unit is used to control the underwater camera and collect raw image data, depth information, illumination information, color temperature information, and initial values ​​of water body type; The second unit is used to calculate the attenuation coefficients of the red, green and blue components in the original image data respectively, based on the product of the depth information and the initial value of the water body type, superimposed with the correction amount of the illuminance information and the color temperature information. The third unit is used to extract pixel color information with brightness values ​​within a preset range from the original image data as ambient background light, and to calculate white balance adjustment parameters based on pixel distribution. The fourth unit is used to input the original image data, attenuation coefficient, ambient background light, illuminance information and color temperature information into the color compensation network. The color compensation network applies physical constraints based on the attenuation coefficient in the intermediate layer and outputs a first compensated image. After subtracting the ambient background light from the original image data, it is amplified according to the exponential function of the attenuation coefficient and outputs a second compensated image. The fifth unit is used to calculate the fusion weight based on the image quality parameters of the original image data, and to use the fusion weight to perform a weighted summation of the first compensation image and the second compensation image to obtain the color-corrected image.

[0011] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0012] This invention achieves dual optimization of color reproduction by combining an underwater optical physical model with a deep learning network. Attenuation coefficients for each band are dynamically calculated based on real-time environmental parameters collected by depth sensors, illuminometers, and color thermometers, providing physical prior constraints for color compensation. Simultaneously, an adaptive channel gain field is output by using a neural network to learn the nonlinear scattering and attenuation characteristics in complex water bodies. A dual-branch fusion mechanism ensures both physical consistency and improved adaptability to complex scenes. The adaptive fusion strategy adjusts the fusion weights online based on local image contrast, signal-to-noise ratio, and color temperature deviation, avoiding over-compensation or under-compensation. The scattering dehazing process effectively suppresses backscattering caused by suspended particles through the combination of physical priors and deep learning, and eliminates flicker artifacts using multi-frame temporal constraints, significantly reducing color difference and improving resolution and image clarity in turbid water bodies.

[0013] This invention establishes a thrust allocation optimization framework based on quadratic programming. Under the premise of satisfying desired force and torque constraints, it uses the minimization of the sum of squared thrusts as the objective function to achieve balanced energy distribution among the thrusters. Power allocation is dynamically adjusted according to the thruster efficiency characteristics, significantly reducing average power consumption and extending endurance. By monitoring thruster speed and current feedback in real time, the system can automatically detect thruster failure states and remove the column corresponding to the failed thruster from the thrust mapping matrix for resolving, achieving fault-tolerant control. Even with partial thruster failure, it can still maintain the motion capability of the main degrees of freedom, significantly improving operational reliability and mission continuity, and meeting the requirements for long-term operations in complex marine environments.

[0014] This invention introduces an attitude stabilization control method based on a disturbance observer. By fusing data from multiple sensors, including gyroscopes, accelerometers, and electronic compasses, it estimates the external disturbance torque in real time using the body dynamics equations and feeds it forward to compensate the attitude control law. This active disturbance rejection design significantly reduces attitude angle fluctuations and stabilization time in strong water flow environments, greatly improving hovering position accuracy and ensuring camera image stability under complex water flow conditions. The adaptive gain adjustment mechanism of the extended Kalman filter dynamically optimizes the filter parameters according to the disturbance intensity, avoiding tracking lag or oscillation problems under non-stationary disturbances with fixed gain. This method has moderate computational complexity, making it suitable for embedded real-time control. It can be directly integrated into underwater remote-controlled camera systems, providing high-quality visual information support for applications such as port inspection, ship inspection, and marine engineering. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the maneuver control and imaging optimization method for an underwater remotely operated camera according to an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the present invention will be described below with reference to the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0017] Figure 1 Methods for optimizing the maneuvering control and imaging of underwater remotely operated cameras include: Control the underwater camera to collect raw image data, depth information, illumination information, color temperature information, and initial values ​​of water body type; Based on the product of the depth information and the initial value of the water body type, and the correction amount of the illuminance information and color temperature information, the attenuation coefficients of the red, green and blue components in the original image data are calculated respectively. Extract pixel color information whose brightness values ​​are within a preset range from the original image data as ambient background light, and calculate white balance adjustment parameters based on pixel distribution; The original image data, attenuation coefficient, ambient background light, illuminance information, and color temperature information are input into a color compensation network. The color compensation network applies a physical constraint based on the attenuation coefficient in the intermediate layer and outputs a first compensated image. After subtracting the ambient background light from the original image data, it is amplified according to an exponential function of the attenuation coefficient and outputs a second compensated image. The fusion weights are calculated based on the image quality parameters of the original image data. The first compensation image and the second compensation image are then weighted and summed using the fusion weights to obtain the color-corrected image. The product coefficient and correction coefficient in the attenuation coefficient calculation are updated based on the color difference between the color-corrected image and the standard color chart. The deep learning network branch and the physical calculation branch are processed in parallel. The fusion weights are dynamically adjusted based on the local contrast and color temperature deviation to achieve adaptive color correction under different water conditions.

[0018] For example, after the underwater maneuvering imaging system is started, the analog voltage signal from the joystick is converted into standardized velocity commands via the control terminal. The linear velocity component covers a range from -0.8 m / s to +0.8 m / s, and the angular velocity component covers a range from -30° / s to +30° / s. These commands are then filtered by a first-order inertial filter to generate the desired velocity vector and the desired angular velocity vector. The filtering time constant is set to 0.15 s to balance response speed and operational smoothness.

[0019] The main controller calculates the desired forces required in three orthogonal directions based on the difference between the desired and current velocities, combined with the body mass and the added mass matrix. For a prototype with a mass of 3.2 kg, the added mass coefficient is set to 1.15 in the forward direction, 1.35 in the lateral direction, and 1.25 in the vertical direction. The calculation of the desired forces considers the proportional term of the velocity error and the acceleration feedforward term, with the proportional gain set to 25 N / (m / s) in the forward direction and 18 N / (m / s) in the lateral and vertical directions. Simultaneously, based on the difference between the desired and current angular velocities, the desired torques on the three orthogonal axes are calculated using the body inertia matrix and the Coriolis torque term. The diagonal elements of the inertia matrix are calculated using a CAD model and corrected through a water tank suspension experiment; the roll axis inertia is 0.42 kg·m. 2 The pitch axis is 0.51 kg·m. 2 The yaw rate is 0.38 kg·m. 2 .

[0020] The thruster assembly comprises four horizontal thrusters and two vertical thrusters. The horizontal thrusters are installed at four positions on the fuselage (forward, backward, left, and right), with their centers 0.14m from the fuselage's center of gravity. Their thrust direction is parallel to the fuselage's longitudinal or transverse axis. The vertical thrusters are installed at the top and bottom of the fuselage, 0.11m vertically from the center of gravity, with their thrust direction along the fuselage's vertical axis. A six-row, six-column thrust mapping matrix is ​​constructed based on the position coordinates of each thruster in the fuselage coordinate system and the unit vector of its thrust direction. The first three rows of this matrix correspond to the force components in the three directions, and the last three rows correspond to the torque components along the three axes. Taking the forward horizontal thruster as an example, its position vectors are 0.14m longitudinally, 0m laterally, and 0m vertically. Its thrust direction is a longitudinal unit vector. Therefore, its contribution coefficient to the longitudinal force is 1, its contribution coefficients to the lateral and vertical forces are 0, its contribution coefficient to the roll moment is 0, its contribution coefficient to the pitch moment is -0.14 N·m / N, and its contribution coefficient to the yaw moment is 0.

[0021] The thrust allocation problem is formulated as follows: Under the constraints of equality of the thrust mapping matrix and desired force and torque, and inequality constraints between the lower limit (-12N) and upper limit (+12N) of thrust for each thruster, find the thrust command vector that minimizes the sum of the squares of the thrusts of all thrusters. This problem belongs to standard quadratic programming and is solved using a sequential quadratic programming algorithm or the interior-point method. An energy consumption weight matrix is ​​introduced into the cost function. For thrusters with relatively flat efficiency curves, the weight is set to 1.0; for thrusters with significant efficiency drops in the low-thrust region, the weight is set to 1.2, guiding the optimizer to prioritize the use of the efficient operating region. The solver iterates within each control cycle, and the convergence criterion is that the change in the objective function is less than 0.001N. 2 Or the number of iterations reaches 20.

[0022] During operation, the main controller continuously monitors the speed and current feedback of each thruster. Speed ​​feedback is acquired via a Hall effect sensor built into the motor, with a sampling frequency of 1kHz. Current feedback is measured by the bus shunt resistor and sampled by a 16-bit analog-to-digital converter. If a thruster's speed feedback is below 70% of the command value for 50ms consecutively, or its current feedback exceeds 1.4 times the rated value of 3.5A (i.e., 4.9A) for 30ms consecutively, the thruster is deemed to have failed. The column corresponding to the failed thruster is removed from the thrust mapping matrix, forming a reduced-dimensional mapping matrix. Simultaneously, the upper and lower limits of the thrust constraint for that thruster are set to 0. The quadratic programming problem is resolved. If the feasible region is not empty, a fault-tolerant allocation result is output; otherwise, a degraded control strategy is triggered, limiting the desired velocity amplitude of certain degrees of freedom.

[0023] The obtained thrust command vector corresponds to the target thrust of each of the six thrusters, in N. These thrust values ​​are converted into pulse width modulation duty cycle commands for the motor controller using thruster calibration curves. The calibration curves are fitted with a cubic polynomial, covering a thrust range from -12N to +12N, with a root mean square residual of less than 0.3N. To reduce the impact of high-frequency jitter in the thrust commands on the mechanical system, each thrust command is output to the thruster motor controller after passing through a first-order low-pass filter with a cutoff frequency of 8Hz. The filtered thrust commands are updated at a frequency of 200Hz, synchronized with the main control cycle.

[0024] The image acquisition module acquires raw image data via a low-light CMOS sensor with a resolution of 1920×1080 pixels and an output format of 12-bit linear RGB. Depth information is provided by a pressure sensor with an accuracy of 0.01m, converted to water depth via hydrostatic pressure, with a measurement range of 0 to 100m. Illuminance information is measured by a light intensity sensor mounted next to the optical window, with a measurement range of 0.1lx to 10000lx, and the output is logarithmic to cover a wide dynamic range. Color temperature information is estimated using a dual-channel spectral sensor, measuring the intensity ratio of the blue and red light bands respectively, and mapping it to a correlated color temperature value, ranging from 2500K to 10000K. The initial water body type is selected by the operator based on the operating area, including four categories: clear seawater, nearshore water, algae-rich water, and turbid freshwater, each corresponding to different attenuation coefficient reference parameters.

[0025] The attenuation coefficient calculation module calculates the attenuation coefficients for the red, green, and blue channels by superimposing the product of depth information and the initial value of the water body type, along with correction terms for illuminance and color temperature information. The product of depth information and the initial value of the water body type uses a linear relationship. For clear seawater, the coefficient corresponding to the initial value of the water body type is set to 0.35m for the red channel. -1 Green channel takes 0.08m -1 The blue channel is 0.02m. -1 In turbid freshwater: the red channel should be 0.55m. -1 Green channel takes 0.18m -1 The blue channel is 0.12m. -1 The correction for illuminance information uses a logarithmic function. When the illuminance is below 10 lx, the correction is positive to compensate for attenuation estimation bias under low light conditions; when the illuminance is above 1000 lx, the correction is negative to suppress overestimation of attenuation in overexposed areas. The correction for color temperature information is achieved through a lookup table; when the color temperature is below 4000 K, the attenuation coefficient of the red channel is increased by 0.05 μm. -1 When the color temperature is above 7000K, reduce the red channel attenuation coefficient by 0.03m. -1 And increase the attenuation coefficient of the blue channel by 0.02m -1The correction amount is limited to ±0.15m. -1 To avoid drastic fluctuations in parameters.

[0026] The ambient background light estimation module extracts pixel color information of distant water bodies from the original image data. A ring-shaped region with a boundary width of 1 / 10 of the original image width is selected around the image center. The 95th percentile of the red, green, and blue channels of the pixels within this region is calculated as the initial value of the ambient background light. This initial value is then subjected to a time-domain exponential smoothing filter with a smoothing coefficient of 0.85 to suppress background light flicker caused by water particle movement. The white balance adjustment parameters are calculated using the gray-world assumption to normalize the mean values ​​of each channel in the original image data, making the ratio of the three channel means close to 1:1:1. In scenarios where a known color chart exists, the white balance adjustment parameters are further corrected based on the deviation between the measured color of the color chart area and the standard color. The correction amount is calculated using the least squares method to find the gain coefficient that minimizes the sum of squared color differences.

[0027] The color compensation network employs an encoder-decoder structure. The encoder contains four convolutional layers with 32, 64, 128, and 256 kernels respectively, each kernel being 3×3 pixels in size. A stride of 2 is used for downsampling, and the activation function is a modified linear unit (MRU). The decoder uses a symmetrical structure, achieving upsampling through transposed convolutions and introducing skip connections to concatenate features from the encoder and corresponding decoder layers. A physical constraint layer is placed between the third layer of the encoder and the second layer of the decoder. This layer calculates the pixel-wise product of the original image data and the white balance adjustment parameters, and then the pixel-wise product of this product with an exponential function of the attenuation coefficient, where the exponent of the exponential function is the product of the negative attenuation coefficient and depth information. The output of the constraint layer is weighted and fused with the convolutional features of the second layer of the decoder using weights of 0.6 and 0.4, and then passed to subsequent layers. The network output layer contains two branches: one outputs a three-channel first compensated image, and the other outputs a single-channel first confidence value, ranging from 0 to 1, normalized using the sigmoid function.

[0028] The physical calculation branch subtracts the ambient background light component of the corresponding channel from each pixel value of the original image data to obtain a background-reduced image. The background-reduced image is then multiplied pixel-by-pixel by an attenuation coefficient exponential function, where the base of the exponential function is the base of the natural logarithm, e, and the exponent term is the product of the attenuation coefficient and depth information, resulting in a preliminary magnified image. Skin color pixel regions in the preliminary magnified image are identified using the YCbCr chromaticity space thresholding method. Pixels with a Cb component between 77 and 127 and a Cr component between 133 and 173 are identified as skin color pixels. The average chromaticity value of the skin color pixel region is calculated and compared with the standard skin color chromaticity value, which is set to Cb = 113 and Cr = 155. White balance adjustment parameters are adjusted based on chromaticity deviation. When the deviation is greater than 10, the gain of the red channel is increased by (deviation × 0.008) and the gain of the blue channel is decreased by (deviation × 0.006). The background-reducing and magnification processes are recalculated to obtain the second compensated image. The second confidence value is estimated based on the signal-to-noise ratio of the initially magnified image. The confidence value is set to 0.8 when the signal-to-noise ratio is higher than 25dB and 0.3 when the signal-to-noise ratio is lower than 15dB. Linear interpolation is used in the intermediate region.

[0029] The fusion weight calculation module determines the weights based on a comprehensive analysis of the local contrast, signal-to-noise ratio (SNR), and color temperature deviation of the original image data. Local contrast is calculated using a sliding window of 32×32 pixels, with the standard deviation of the pixel grayscale values ​​within each window used as the contrast index. SNR is estimated using the signal mean and noise variance of flat regions in the image; flat regions are defined as connected regions with gradient amplitudes less than a threshold of 5. Color temperature deviation is the absolute difference between the measured color temperature and the target color temperature of 6500K. After normalization, the three indices are weighted and summed using coefficients of 0.4, 0.35, and 0.25, respectively. This sum is then mapped to a fusion weight between 0 and 1 using an S-curve function, with the midpoint of the S-curve set at 0.5 and the slope set to 8 for a smooth transition. The fusion weight is multiplied by the first compensation image, and the value of (1 minus the fusion weight) is multiplied by the second compensation image. These values ​​are then summed pixel-by-pixel to obtain the color-corrected image.

[0030] During operation, standard color chart images are periodically acquired. The color chart contains 24 standard color patches, covering grayscale and common colors. The color difference between the color of the color chart area in the color-corrected image and the color of the standard color chart is calculated. The color difference is calculated using the CIEDE2000 formula, comprehensively considering differences in brightness, chromaticity, and saturation. When the average color difference exceeds a threshold of 3.0, an online update of the attenuation coefficient is triggered. The update uses gradient descent, adjusting the product coefficient of depth information and initial water body type, as well as the correction coefficients of illuminance information and color temperature information, to minimize the average color difference. The learning rate of gradient descent is set to 0.002, and it stops after 10 iterations. The updated coefficients are saved and used for subsequent image processing.

[0031] For example, when photographing a white target in nearshore waters at a depth of 5m and a turbidity of 50 NTU, the original image shows a noticeable blue-green tint, with the red channel's average value being only 0.42 times that of the blue channel. After color correction algorithm processing, the red channel's average value increased to 0.88 times that of the blue channel, and the chromaticity deviation from standard white decreased from 12.6 to 4.1. When photographing a color chart in clear seawater at a depth of 15m, the original image had an average color difference of 8.3, which was reduced to 3.2 after processing, and the MTF50 index increased from 812 lp / ph to 1168 lp / ph.

[0032] Optionally, the steps of calculating the first compensated image and the second compensated image include: The color compensation network adopts an encoder-decoder structure. A constraint layer is set in the middle layer. The constraint layer calculates the pixel-wise product of the original image data and the white balance adjustment parameter, and then the pixel-wise product of the attenuation coefficient exponential function. The calculation result is weighted and fused with the output of the network front layer and then passed to the subsequent layers to output the first compensated image and the first confidence value. The ambient background light component of the corresponding channel is subtracted from each pixel value of the original image data. The subtraction result is multiplied pixel by pixel with the attenuation coefficient exponential function to obtain a preliminary magnified image. Skin color pixel regions in the preliminary magnified image are identified. The white balance adjustment parameters are adjusted and recalculated according to the deviation between the skin color chroma and the standard skin color chroma to obtain a second compensated image and a second confidence value.

[0033] For example, the encoder part of the color compensation network receives raw image data as input, with an image size of 1920×1080 pixels and a three-channel RGB format. The first convolutional layer uses 32 convolutional kernels, each 3×3 pixels in size, with uniform padding and a stride of 2, reducing the output feature map size to 960×540 pixels. A batch normalization layer and a modified linear unit activation function are then applied after the convolution. The second convolutional layer uses 64 convolutional kernels with the same parameter settings as the first layer, reducing the output feature map size to 480×270 pixels. The third and fourth layers use 128 and 256 convolutional kernels respectively, resulting in a final encoder output size of 120×68 pixels and 256 channels.

[0034] The constraint layer is located between the encoder's third-layer output and the decoder's second-layer input. The constraint layer first receives the original image data, downsampled to 240×135 pixels using bilinear interpolation, and receives the white balance adjustment parameter matrix. This matrix is ​​a 3×3 diagonal matrix, with diagonal elements representing the white balance gain coefficients for the red, green, and blue channels. Each channel of the downsampled image is multiplied pixel-by-pixel with its corresponding white balance gain coefficient to obtain the white balance corrected image. The constraint layer also receives attenuation coefficients and depth information, calculating the attenuation coefficient exponential function. The exponent term is the product of the negative attenuation coefficient and the depth information, which is considered constant across the entire image. The red, green, and blue channels each correspond to three attenuation coefficients, resulting in three exponential function values. The exponential function value for each channel is multiplied pixel-by-pixel with the white balance corrected image for that channel to obtain the physical constraint image.

[0035] The physical constraint image is mapped to 128 channels of features, the same as the output channels of the encoder's third layer, through a 1×1 pixel convolutional layer. This feature is fused with the encoder's third-layer output using learnable weights, which are automatically determined during network training, with initial values ​​of 0.6 for the physical constraint features and 0.4 for the encoder features. The fused features are input to the second layer of the decoder, which uses transposed convolutions for upsampling. The second layer has 128 transposed convolution kernels, a kernel size of 4×4 pixels, and a stride of 2, resulting in an output feature map size of 240×135 pixels. The first layer of the decoder has 64 transposed convolution kernels, an output size of 480×270 pixels, and is concatenated with the encoder's second-layer output via skip connections. The final decoder output layer contains two branches: the first branch uses three convolutional kernels to output a three-channel first compensation image, and the second branch uses one convolutional kernel to output a single-channel first confidence value. The confidence value is mapped to a value between 0 and 1 using a sigmoid activation function.

[0036] The network training employs a hybrid loss function, comprising reconstruction loss, perceptual loss, and physical consistency loss. The reconstruction loss is the L1 distance between the first compensated image and the ground truth image, with a weighting coefficient of 1.0. The perceptual loss uses features from the third and fourth layers of the pre-trained VGG16 network, calculating the cosine distance between the feature maps, with a weighting coefficient of 0.3. The physical consistency loss constrains the first compensated image to satisfy the underwater imaging physics model; that is, the product of the first compensated image and (the inverse of the attenuation coefficient exponential function) plus the ambient background light should approximate the original image. This loss term uses the L2 distance, with a weighting coefficient of 0.5. The training dataset contains 5000 pairs of underwater images and their corresponding ground truth images, generated through pool experiments and synthetic data. The optimizer uses the Adam algorithm with an initial learning rate of 0.0001, decaying to 0.00001 after 30 training epochs. The batch size is 8, and the total number of training epochs is 80.

[0037] The physical computation branch is implemented independently of the neural network. First, ambient background light is subtracted from each pixel value of the original image data. Ambient background light is a three-channel vector: red (Br), green (Bg), and blue (Bb). For a pixel's red channel value Ir, the de-illuminated value is calculated as (Ir - Br). If the result is less than 0, it is truncated to 0. The green and blue channels are processed similarly to obtain the de-illuminated image. Each channel of the de-illuminated image is multiplied by its corresponding attenuation coefficient exponential function. The exponential function for the red channel is e raised to the power of (the negative red attenuation coefficient multiplied by the depth), and the same applies to the green and blue channels. Multiplication is performed pixel-by-pixel to obtain a preliminary magnified image.

[0038] The image is initially enlarged and converted to the YCbCr color space, where Y represents luminance and Cb and Cr represent chrominance. The conversion formulas are: Y = (0.299 × R + 0.587 × G + 0.114 × B), Cb = (-0.169 × R - 0.331 × G + 0.5 × B + 128), Cr = (0.5 × R - 0.419 × G - 0.081 × B + 128). Here, R represents the red channel value, G represents the green channel value, and B represents the blue channel value. All pixels in the image are iterated over; pixels with a Cb value between 77 and 127 and a Cr value between 133 and 173 are marked as candidate skin color pixels. The mean values ​​of Cb and Cr are calculated for all candidate skin color pixels, denoted as CbMean and CrMean. The standard skin tone has a Cb value of 113 and a Cr value of 155, with the calculated deviations being (CbMean-113) and (CrMean-155).

[0039] The white balance adjustment parameters are adjusted based on Cb and Cr deviations. These parameters are three-element vectors, corresponding to the gains of the red, green, and blue channels, respectively. The red channel gain is adjusted by (Cr deviation × 0.008), the blue channel gain by (negative Cb deviation × 0.006), and the green channel gain remains unchanged. The adjusted white balance parameters are applied to the original image, and the background light removal and exponential magnification processes are recalculated to obtain the second compensated image. The second confidence value is determined based on the signal-to-noise ratio (SNR) of the initially magnified image. The SNR is calculated as (mean grayscale value of flat areas divided by the noise standard deviation) multiplied by 20. When the SNR is ≥ 25 dB, the second confidence value is set to 0.8; when the SNR is ≤ 15 dB, it is set to 0.3; the intermediate region is calculated using linear interpolation.

[0040] The calculation of the fusion weight combines the first and second confidence values. First, the confidence difference is calculated as (first confidence value - second confidence value), ranging from -1 to +1. This difference is then input into a sigmoid function, which is 1 divided by (1 + e raised to the power of -8 times the difference), with an output ranging from 0 to 1. The output value is the fusion weight, used to weight the first compensation image, and (1 - fusion weight) is used to weight the second compensation image. A weighted summation is performed pixel-by-pixel: each pixel in the first compensation image is multiplied by the fusion weight, and the corresponding pixel in the second compensation image is multiplied by (1 - fusion weight). The sum of these two values ​​yields the pixel value in the final color-corrected image. This operation is performed independently for each of the three color channels.

[0041] This method achieves a complementary fusion of deep learning and physical models. By introducing physical constraints into the intermediate layers of the neural network, it retains the adaptability of deep learning while ensuring physical rationality. Simultaneously, special processing of skin-colored pixel regions improves image quality in portrait scenes. The confidence values ​​output by the two compensation images contribute to more accurate adaptive fusion, significantly improving the color reproduction and naturalness of underwater images under different water conditions.

[0042] Optionally, after acquiring the raw image data and before inputting it into the color compensation network, a scattering dehazing process is also included: Spatial frequency transformation is performed on the original image data to extract the low-frequency component as the initial value of the backscatter component, and the initial value of the direct light component is obtained by subtracting the initial value of the backscatter component. The initial values ​​of the backscatter component, the initial values ​​of the direct light component, the depth information, and the illuminance information are input into the backscatter estimation network, and the refined backscatter component and transmittance correction coefficient are output. Acquire image data from multiple frames at adjacent time points, calculate the pixel displacement field, transform the backscattered component of the next frame to the coordinate system of the current frame, and adjust the refined backscattered component to minimize the sum of squared differences between the transformed backscattered component and the backscattered component of the current frame, thus obtaining the time-stable backscattered component. The time-stable backscatter component is subtracted from the original image data and multiplied pixel by pixel with the reciprocal of the transmittance correction coefficient to obtain the dehazed image; The dark channel statistics of the dehazed image are calculated as the residual suspended particle density index. When the residual suspended particle density index is higher than the preset index threshold, the reciprocal value of the transmittance correction coefficient is reduced, and the adjusted image is used as the original image data input to the color compensation network.

[0043] For example, before the original image data is input into the color compensation network, the backscattering dehazing module preprocesses the image to remove the contrast reduction and visibility degradation caused by suspended particles. The original image is first transformed to the frequency domain using a two-dimensional discrete cosine transform, with a transform block size of 8×8 pixels, dividing the entire image into blocks for processing. For each transform block, the coefficients of the upper left quarter region are extracted as low-frequency components, and the remaining coefficients are zeroed before an inverse transform is performed to obtain the initial value of the backscattering component. The low-frequency component corresponds to the slowly changing background haze in the image, while the high-frequency component corresponds to target details and textures. Subtracting the initial value of the backscattering component from the original image yields the initial value of the direct light component, which contains real scene information but is still affected by residual scattering.

[0044] The backscatter estimation network employs a U-shaped structure. The encoder contains five convolutional layers with 16, 32, 64, 128, and 256 channels per layer, respectively. The kernel size is 5×5 pixels, the stride is 2, and the activation function is leakyReLU with a negative slope of 0.2. The network input consists of four channels: the first three channels are the initial RGB values ​​of the backscatter components, and the fourth channel is a single-channel image with depth information normalized to between 0 and 1. The depth information is constant throughout the image but expanded into a feature map of the same size as the image within the network. The encoder's fifth layer outputs 60×34 pixels with 256 channels. The decoder uses a symmetrical structure, upsampling through transposed convolutions. The number of channels per layer is 128, 64, 32, and 16, respectively. The final output layer contains two branches: one outputs a single-channel refined backscatter component ranging from 0 to 1, and the other outputs a single-channel transmittance correction coefficient ranging from 0.5 to 1.5, linearly mapped after passing through a sigmoid function.

[0045] Illuminance information serves as an auxiliary input to guide backscatter estimation. Illuminance values ​​are first normalized: the logarithmic illuminance measurement is subtracted from the minimum value of 2.3 (corresponding to the logarithm of 10 lx), and then divided by the range 6.9 (corresponding to the logarithmic difference between 10000 lx and 10 lx) to obtain a normalized illuminance value between 0 and 1. This scalar is mapped to a 16-dimensional feature vector through a fully connected layer, and then expanded into a feature map of the same size as the output of the third layer of the encoder, with 16 channels, via a broadcast mechanism. This feature map is concatenated with the output of the third layer of the encoder, resulting in (64 + 16) = 80 channels, which is then input to the fourth layer of the encoder.

[0046] The network training employs supervised learning. The ground truth backscattered components are coarsely estimated using a dark channel prior algorithm and then refined manually. The loss function includes L1 reconstruction loss, total variational regularization, and physical consistency constraints. The L1 loss calculates the mean absolute error between the refined backscattered components and the ground truth, with a weight of 1.0. Total variational regularization penalizes the spatial gradient of the backscattered components, with a weight of 0.01, promoting a smoother backscattered distribution. The physical consistency constraint requires that the original image, after subtracting the backscattered components, divided by the transmittance should approximate the fog-free image; this term uses perceptual loss with a weight of 0.2. The training dataset contains 3000 pairs of foggy and fog-free underwater images. The optimizer is AdamW, with an initial learning rate of 0.0003, weight decay of 0.0001, and 60 training epochs.

[0047] The multi-frame optical flow constraint module acquires image data from adjacent time points, with a time interval of 0.1 seconds, corresponding to adjacent frames when the camera is in slow motion. A dense optical flow field is calculated for the current frame and the next frame using the pyramid LK optical flow algorithm, with 3 pyramid layers and a window size of 15 pixels. The optical flow field includes the horizontal and vertical displacement of each pixel between the two frames. A geometric transformation is performed on the backscattered component of the next frame based on the optical flow field, mapping it to the coordinate system of the current frame using bilinear interpolation. The difference between the transformed backscattered component and the refined backscattered component of the current frame is calculated, and the squared difference is summed across the entire image to obtain the temporal inconsistency loss.

[0048] The temporally stable backscattered component is obtained through optimization. The objective function is (temporal inconsistency loss plus spatial smoothing regularization term) with a weight ratio of 0.7:0.3. The optimization variable is the adjustment amount of the backscattered component in the current frame, with the adjustment range limited to ±0.05. Gradient descent is used for 5 iterations with a learning rate of 0.01. The updated backscattered component is used as the temporally stable backscattered component. This component remains continuous in the time dimension, avoiding flicker artifacts caused by independent frame-by-frame processing.

[0049] The calculation of the dehazed image first involves subtracting the time-stabilized backscatter component from the original image data. For a certain channel value I of a pixel in the original image, the backscatter component is Jb, and the subtraction result is (I-Jb). If it is less than 0, it is truncated to 0. The subtraction result is divided by the reciprocal of the transmittance correction coefficient, i.e., multiplied by the transmittance correction coefficient, to obtain the dehazed pixel value. The transmittance correction coefficient is output by the network and ranges from 0.5 to 1.5. When the water is clear, the coefficient is close to 1.0; when scattering is severe, the coefficient is less than 1.0 to enhance contrast recovery; and when there is a risk of overcompensation, the coefficient is greater than 1.0 to suppress noise amplification. The three channels are processed independently to obtain the dehazed image.

[0050] The quality assessment of dehazed images is achieved through dark channel statistics. The dark channel is defined as the minimum value of the RGB three channels for each pixel. A dark channel image is calculated for the dehazed image. The 10% of pixels with the lowest brightness in the dark channel image are selected, and their average value is calculated as a residual suspended particle density index. This index reflects the degree of haze remaining in the dehazed image, ideally close to 0. When the index exceeds a preset threshold of 0.15, the dehazing intensity is deemed insufficient or there is overcompensation. The reciprocal of the transmittance correction coefficient is reduced, specifically by multiplying the coefficient by 0.9, and the dehazed image is recalculated. The adjusted dark channel index of the dehazed image typically drops below 0.12. At this point, the image is used as the raw image data input to the color compensation network.

[0051] This method accurately estimates the backscatter component by combining frequency domain separation with deep learning, and the temporal stability constraint effectively eliminates the inter-frame flicker problem. The residual suspended particle density assessment mechanism can adaptively adjust the dehazing intensity to prevent image distortion caused by excessive dehazing, and significantly improve the imaging clarity and contrast in turbid water environments while maintaining image naturalness.

[0052] Optionally, the steps for controlling the underwater camera include: Generate the desired velocity and desired angular velocity based on the control commands, and calculate the desired force in three orthogonal directions and the desired torque on three orthogonal axes; Based on the installation position coordinates of multiple thrusters in the body coordinate system and the thrust direction unit vector, calculate the force components of each thruster about the three orthogonal directions and the torque components about the three orthogonal axes, and construct the thrust mapping matrix. Under the conditions of satisfying the equality constraints of the thrust mapping matrix, the desired force and the desired torque, and the upper and lower limits of thrust, find the thrust command vector that minimizes the sum of the squares of thrust. The speed and current feedback of each thruster are detected. When the speed feedback is lower than the preset ratio of the command value or the current feedback exceeds the rated value, the thruster is determined to be faulty. The column corresponding to the faulty thruster is deleted from the thrust mapping matrix and the thrust command vector is re-solved. The elements of the thrust command vector are first-order low-pass filtered and then output to each thruster.

[0053] For example, the body motion control module calculates the desired motion state based on the manipulation commands from the surface control terminal. The manipulation commands include normalized control quantities for six degrees of freedom, corresponding to forward / backward movement, lateral movement, ascent / descent, roll, pitch, and yaw, with each control quantity ranging from -1 to +1. The normalized control quantity is multiplied by the maximum velocity or angular velocity of the corresponding degree of freedom to obtain the desired velocity vector and desired angular velocity vector. The maximum forward / backward velocity is set to 0.8 m / s, the maximum lateral movement velocity to 0.6 m / s, the maximum ascent / descent velocity to 0.5 m / s, the maximum roll and pitch angular velocities to 25° / s, and the maximum yaw angular velocity to 30° / s. The desired velocity vector is denoted as a three-dimensional vector vd, containing the longitudinal velocity u, the lateral velocity v, and the vertical velocity w; the desired angular velocity vector is denoted as a three-dimensional vector ωd, containing the roll angular velocity p, the pitch angular velocity q, and the yaw angular velocity r.

[0054] The current velocity and angular velocity are estimated through fusion of data from an inertial measurement unit (IMU) and a depth sensor. The IMU includes a three-axis accelerometer and a three-axis gyroscope, with a sampling frequency of 200 Hz. The accelerometer measures the specific force in the body coordinate system, and the gyroscope measures the angular velocity. The depth sensor measures water depth via pressure, with a sampling frequency of 50 Hz. An extended Kalman filter fuses the data from the accelerometer, gyroscope, and depth sensor. The state vector contains 13 dimensions, including position, velocity, attitude quaternions, and angular velocity. The filter prediction step updates the attitude quaternion based on the angular velocity and updates the velocity and position based on the acceleration and attitude. The update step corrects the vertical position using depth sensor measurements and uses zero-velocity correction to constrain horizontal velocity drift. The filter outputs the current velocity estimate v and the angular velocity estimate ω.

[0055] The expected force is calculated based on velocity error and acceleration feedforward. The velocity error is (expected velocity vd - current velocity v), yielding the three-dimensional error vector ev. The expected acceleration is estimated using the time derivative of the expected velocity, employing a first-order difference: (current expected velocity - expected velocity from the previous control cycle) divided by the control cycle of 0.005 s. The expected force Fd equals (mass plus additional mass matrix) multiplied by the expected acceleration, plus (proportional gain matrix multiplied by the velocity error). The mass matrix is ​​a diagonal matrix with diagonal elements: longitudinal (mass plus additional mass) 3.68 kg, lateral 4.32 kg, and vertical 4.00 kg. The proportional gain matrix has diagonal elements: longitudinal 25 N / (m / s), lateral 18 N / (m / s), and vertical 18 N / (m / s). The expected force Fd is a three-dimensional vector containing longitudinal force Fx, lateral force Fy, and vertical force Fz.

[0056] The desired torque is calculated based on angular velocity error and angular acceleration feedforward, while compensating for Coriolis torque. The angular velocity error is (desired angular velocity ωd - current angular velocity ω), yielding the three-dimensional error vector eω. The desired angular acceleration is estimated using the time derivative of the desired angular velocity. The inertia matrix is ​​a diagonal matrix, with diagonal elements including: roll axis inertia 0.42 kg·m. 2 Pitch axis 0.51 kg·m 2 Yaw axis 0.38 kg·m 2 The Coriolis torque is calculated as (angular velocity vector cross product of inertia matrix multiplied by angular velocity vector). The desired torque Md equals (inertia matrix multiplied by desired angular acceleration), plus (proportional gain matrix multiplied by angular velocity error), minus the Coriolis torque. The diagonal elements of the proportional gain matrix are: roll axis 3.5 N·m / (rad / s), pitch axis 4.2 N·m / (rad / s), yaw axis 3.0 N·m / (rad / s). The desired torque Md is a three-dimensional vector containing roll torque Mx, pitch torque My, and yaw torque Mz. The desired force Fd and the desired torque Md are combined into a six-dimensional desired generalized torque wd.

[0057] The thrust mapping matrix B is constructed based on the installation positions and thrust directions of the six thrusters. With the aircraft's center of gravity as the origin, the longitudinal axis is the positive x-axis pointing forward, the transverse axis is the positive y-axis pointing to the right, and the vertical axis is the positive z-axis pointing downwards. The position of the front horizontal thruster is (x=0.14m, y=0m, z=0m), and the thrust direction is a positive unit vector along the x-axis. The position of the rear horizontal thruster is (x=-0.14m, y=0m, z=0m), and the thrust direction is a negative unit vector along the x-axis. The position of the right horizontal thruster is (x=0m, y=0.14m, z=0m), and the thrust direction is a positive unit vector along the y-axis. The position of the left horizontal thruster is (x=0m, y=-0.14m, z=0m), and the thrust direction is a negative unit vector along the y-axis. The top vertical thruster is located at (x=0m, y=0m, z=-0.11m), and its thrust direction is a negative unit vector along the z-axis. The bottom vertical thruster is located at (x=0m, y=0m, z=0.11m), and its thrust direction is a positive unit vector along the z-axis.

[0058] The thrust mapping matrix B is a 6×6 matrix, with each column corresponding to the contribution of a thruster to the six-dimensional generalized moment. For each thruster, the thrust contribution to the longitudinal force is the thrust direction x-component, to the lateral force is the y-component, and to the vertical force is the z-component. The thrust contribution to the roll moment is (position y-coordinate × thrust z-component - position z-coordinate × thrust y-component), to the pitch moment is (position z-coordinate × thrust x-component - position x-coordinate × thrust z-component), and to the yaw moment is (position x-coordinate × thrust y-component - position y-coordinate × thrust x-component). The actual thruster arrangement needs adjustment to generate the yaw moment. The fore and aft thrusters are laterally offset. Assuming the fore thruster is laterally offset by 0.05m, its position is corrected to (0.14m, 0.05m, 0m), while the thrust direction remains (1, 0, 0). Then the yaw moment is (0.14 × 0 - 0.05 × 1), which equals -0.05 N·m / N. The rear thruster is laterally offset by -0.05m, located at (-0.14m, -0.05m, 0m), with a thrust direction of (-1, 0, 0) and a yaw moment of (-0.14×0 - (-0.05)×(-1)) equal to -0.05 N·m / N. With this configuration, both the front and rear thrusters can generate yaw control. The thrust mapping matrix B is constructed accordingly, with six columns corresponding to the six thrusters.

[0059] The thrust allocation problem seeks to find a thrust command vector *t* such that *Bt* equals *wd* (i.e., the thrust mapping matrix *B* multiplied by the thrust command vector *t* equals the desired generalized torque vector *wd*), and each element of *t* is between -12N and +12N. Simultaneously, the problem aims to minimize the sum of squares of *t*. The sum of squares represents the energy cost, and a weighted sum of squares is used. The weight matrix *We* is a diagonal matrix, with diagonal elements representing the reciprocals of the efficiency weights of each thruster. Thrusters with flat efficiency curves have a weight of 1.0, while those with decreasing efficiency in the low-thrust region have a weight of 1.2. The objective function is (0.5 × t transpose × *We* × t). The constraints are that *Bt* equals *wd*, the lower bound of *t* is a six-element vector of -12N, and the upper bound is a six-element vector of +12N. This problem is a convex quadratic programming problem, solved using an interior-point solver with a convergence tolerance of 0.001 and a maximum of 20 iterations. The solver is invoked in each control cycle, with a cycle of 0.005s (200Hz).

[0060] Thruster failure detection is achieved through speed and current feedback. Each thruster motor controller transmits a speed signal, measured by a Hall effect sensor with a resolution of 1 rpm. Current feedback is sampled via a shunt resistor and a 16-bit analog-to-digital converter, with a resolution of 0.01 A. The thrust command is converted into a target speed command; the conversion relationship is determined by the thruster calibration curve, which shows a linear relationship between thrust and the square of the speed, with a coefficient of 0.0032 N / (rpm). 2 The target rotational speed is the square root of (thrust command divided by 0.0032). Positive thrust corresponds to positive rotational speed, and negative thrust corresponds to negative rotational speed.

[0061] The speed feedback is compared with the target speed. If the absolute value of the speed feedback is lower than 70% of the absolute value of the target speed for 50 ms consecutively, the thruster response is considered insufficient. The current feedback is compared with the rated current of 3.5A. If the current feedback exceeds (1.4 × rated current), i.e., 4.9A, for 30 ms consecutively, the thruster is considered overloaded. Both cases are considered thruster failures. The column corresponding to the failed thruster is deleted from the thrust mapping matrix B, forming a reduced-dimensional matrix B' with one fewer column. The elements of the failed thruster in the thrust command vector t are set to 0 and fixed. The optimization problem becomes solving for the thrust of the remaining thrusters, and the constraint becomes (B' × t') equals wd, where t' is the thrust vector of the remaining thrusters. If the number of remaining thrusters is ≥ 6 and matrix B' is full rank, the problem still has a solution, and the solver outputs the fault-tolerant allocation result. If matrix B' is not full rank or the number of remaining thrusters is less than 6, the problem is not feasible, triggering a degradation strategy to limit certain desired velocity amplitudes, such as prohibiting roll motion, setting the desired roll angular velocity to 0, and recalculating the desired torque before solving.

[0062] The obtained thrust command vector t is smoothed by a first-order low-pass filter before being output. The filter transfer function is (1 divided by 1 plus s multiplied by the filter time constant), with a time constant of 0.02s, corresponding to a cutoff frequency of approximately 8Hz. The filter is implemented using a difference equation, where the current output equals (the previous output + the quotient of the control cycle divided by the time constant plus the control cycle × the difference between the current input and the previous output). After filtering, the thrust command is converted into a duty cycle command for the motor controller via a calibration curve, with a duty cycle range of 5% to 95%, corresponding to a thrust of -12 to +12N. The duty cycle command is transmitted to each thruster motor controller via the CAN bus at a baud rate of 500kbit / s, with a transmission period of 0.005s.

[0063] This method minimizes energy consumption through an optimized thrust allocation algorithm, and the thruster failure detection and reconstruction mechanism significantly improves the system's fault tolerance. Thrust smoothing effectively reduces control jitter, ensuring the stability of the camera platform. Even in complex water flow environments or with partial thruster failure, it can maintain precise position and attitude control, providing stable platform support for high-quality underwater imaging.

[0064] Optionally, the method further includes: Data from gyroscopes, accelerometers, electronic compasses, and depth sensors are collected and fused using an extended Kalman filter to obtain attitude quaternion estimates and angular rate estimates. The estimated value of the external disturbance torque is calculated based on the body inertia matrix, the rate of change of angular rate, the thruster control torque, and the quadratic term of angular rate. The quaternion error between the desired attitude and the estimated attitude is calculated as the attitude error, the difference between the desired angular rate and the estimated angular rate is calculated as the angular rate error, and the negative value of the sum of the proportional term of the attitude error, the differential term of the angular rate error, and the estimated value of the external disturbance torque is used as the attitude control torque. The attitude control torque is superimposed on the desired torque to solve for the thrust command vector.

[0065] For example, the attitude stabilization control module adds active compensation for external disturbances to the body motion control, ensuring attitude stability during water flow impacts or asymmetrical thruster output. A gyroscope measures three-axis angular velocity with a range of ±500° / s and a resolution of 0.01° / s. An accelerometer measures three-axis specific force with a range of ±16g and a resolution of 0.005g. An electronic compass measures three-axis magnetic field strength with a range of ±4 Gauss and a resolution of 0.001 Gauss. A depth sensor measures hydrostatic pressure converted to water depth with a range of 0 to 100m and a resolution of 0.01m. Data from these four sensors are input to an extended Kalman filter for fusion.

[0066] The extended Kalman filter's state vector contains an attitude quaternion q, with four elements: a scalar part q0, vector parts q1, q2, and q3, and an angular rate vector omega, containing three elements p, q, and r, for a total of seven dimensions. The prediction step updates the attitude quaternion based on the angular rate measured by the gyroscope. The quaternion differential equation is (the derivative of q equals 0.5 times the quaternion multiplication matrix multiplied by the angular rate extension vector), and the angular rate extension vector is (0, p, q, r). After discretization, the quaternion is updated as (current quaternion + 0.5 × control period × quaternion multiplication matrix × angular rate), and the updated quaternion is normalized to maintain unit length. The angular rate prediction uses the current angular rate, assuming a uniform velocity model with zero angular acceleration. The prediction covariance matrix is ​​updated based on the process noise covariance matrix Q, a diagonal matrix whose diagonal elements are the gyroscope noise variance of 0.0001 (rad / s). 2 .

[0067] The update process utilizes measurements from the accelerometer, electronic compass, and depth sensor to correct the state estimate. The accelerometer measures the gravity vector under stationary or uniform motion conditions; the accelerometer output is predicted through attitude quaternion rotation, and the difference between the predicted and measured values ​​is used to correct the attitude. The electronic compass measures the geomagnetic vector; the compass output is predicted through attitude quaternion rotation, and the difference is used to correct the yaw angle. The depth sensor measures the vertical position, directly correcting the vertical position state. The measurement noise covariance matrix R is a diagonal matrix; the standard deviation of the accelerometer noise is 0.02g, the standard deviation of the compass noise is 0.05 Gauss, and the standard deviation of the depth sensor noise is 0.02m. The Kalman gain K is calculated as (prediction covariance × transpose of the observation matrix × inverse of [observation matrix × prediction covariance × transpose of the observation matrix + measurement noise covariance]). The state update is (predicted state + Kalman gain × measurement residual), and the covariance update is ([identity matrix - Kalman gain × observation matrix] × prediction covariance). The filter outputs attitude quaternion estimates q_est and angular rate estimates omega_est.

[0068] The external disturbance torque is estimated based on the body dynamics equations. Angular acceleration is equal to (the inverse of the inertia matrix multiplied by [thruster control torque + external disturbance torque - Coriolis torque]). Angular acceleration is calculated using the time derivative of the estimated angular rate omega_est, employing a first-order difference: (current angular rate - previous cycle angular rate) divided by the control cycle of 0.005s. The thruster control torque is (the last three rows of the thrust mapping matrix multiplied by the thrust command vector), yielding the three-dimensional torque vector tau_cmd. The Coriolis torque is (the cross product of the angular rate vector and the inertia matrix multiplied by the angular rate vector). Rearranging the dynamics equations, the external disturbance torque tau_dist is equal to (inertia matrix × angular acceleration - thruster control torque + Coriolis torque). The inertia matrix is ​​a diagonal matrix with diagonal elements of 0.42, 0.51, and 0.38 kg·m. 2 The external disturbance moment tau_dist is a three-dimensional vector that includes roll disturbance, pitch disturbance, and yaw disturbance.

[0069] Attitude error is calculated as the quaternion difference between the desired attitude and the estimated attitude. The desired attitude is given by control commands or waypoint planning and is represented as a quaternion q_des. The attitude error quaternion q_err is equal to (the conjugate of the desired quaternion q_des and the estimated quaternion q_est), where the conjugate quaternions have the same scalar part and the negative vector part. The vector part of the attitude error quaternion is extracted as a three-dimensional attitude error vector e_q, with elements q_err1, q_err2, and q_err3, corresponding to the roll, pitch, and yaw errors, in approximately radians. The angular rate error e_omega is (desired angular rate omega_des - estimated angular rate omega_est).

[0070] The attitude control torque tau_att is calculated as (proportional term of attitude error + differential term of angular rate error - estimated external disturbance torque). The proportional term is (proportional gain matrix Kp × attitude error e_q), where Kp is a diagonal matrix with diagonal elements of: roll axis 5 N·m / rad, pitch axis 6 N·m / rad, yaw axis 4 N·m / rad. The differential term is (differential gain matrix Kd × angular rate error e_omega), where Kd is a diagonal matrix with diagonal elements of: roll axis 2 N·m / (rad / s), pitch axis 2.5 N·m / (rad / s), yaw axis 1.8 N·m / (rad / s). The attitude control torque tau_att equals (-Kp×e_q-Kd×e_omega+tau_dist), with the negative sign indicating that the control torque is inversely proportional to the error.

[0071] The proportional gain Kp is adaptively adjusted based on the magnitude of the external disturbance torque estimate tau_dist. The magnitude of the disturbance torque is the square root of the sum of the squares of the three elements of tau_dist. When the magnitude is < 0.5 N·m, the gain remains at the baseline value. When the magnitude is ≥ 0.5 N·m, the gain increases proportionally by a factor of (1 + 0.3 × [magnitude - 0.5]), with an upper limit of twice the baseline value. For example, if the baseline gain for the roll axis is 5 N·m / rad, and the disturbance magnitude is 1.5 N·m, the adjusted gain is (5 × [1 + 0.3 × 1]) equal to 6.5 N·m / rad. The gain adjustment is updated every 0.1 s using exponential smoothing with a smoothing coefficient of 0.7 to avoid abrupt changes in gain.

[0072] The attitude control torque tau_att is superimposed on the desired torque Md to form the last three elements of the comprehensive desired torque wd. The first three elements are the desired force Fd, which remains unchanged. The comprehensive desired generalized torque wd is input to the thrust distribution module to solve for the thrust command vector. The thrust command is filtered and output to the thruster, forming a unified closed loop for attitude stabilization and motion control.

[0073] This method improves attitude estimation accuracy by fusing multi-sensor data using an extended Kalman filter, and the external disturbance torque estimation and compensation mechanism effectively addresses water flow interference. The PID control structure based on quaternion errors avoids the singularity problem of Euler angle representation, ensuring stable control across the entire attitude range and significantly enhancing the attitude stability and anti-interference capability of the underwater camera platform in complex environments.

[0074] Secondly, it provides a maneuvering control and imaging optimization system for underwater remotely operated cameras, including: The first unit is used to control the underwater camera and collect raw image data, depth information, illumination information, color temperature information, and initial values ​​of water body type; The second unit is used to calculate the attenuation coefficients of the red, green and blue components in the original image data respectively, based on the product of the depth information and the initial value of the water body type, superimposed with the correction amount of the illuminance information and the color temperature information. The third unit is used to extract pixel color information with brightness values ​​within a preset range from the original image data as ambient background light, and to calculate white balance adjustment parameters based on pixel distribution. The fourth unit is used to input the original image data, attenuation coefficient, ambient background light, illuminance information and color temperature information into the color compensation network. The color compensation network applies physical constraints based on the attenuation coefficient in the intermediate layer and outputs a first compensated image. After subtracting the ambient background light from the original image data, it is amplified according to the exponential function of the attenuation coefficient and outputs a second compensated image. The fifth unit is used to calculate the fusion weight based on the image quality parameters of the original image data, and to use the fusion weight to perform a weighted summation of the first compensation image and the second compensation image to obtain the color-corrected image.

[0075] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

Claims

1. A method for maneuver control and imaging optimization of an underwater remotely operated camera, characterized in that, include: Control the underwater camera to collect raw image data, depth information, illumination information, color temperature information, and initial values ​​of water body type; Based on the product of the depth information and the initial value of the water body type, and the correction amount of the illuminance information and color temperature information, the attenuation coefficients of the red, green and blue components in the original image data are calculated respectively. Extract pixel color information whose brightness values ​​are within a preset range from the original image data as ambient background light, and calculate white balance adjustment parameters based on pixel distribution; The original image data, attenuation coefficient, ambient background light, illuminance information, and color temperature information are input into a color compensation network. The color compensation network applies a physical constraint based on the attenuation coefficient in the intermediate layer and outputs a first compensated image. After subtracting the ambient background light from the original image data, it is amplified according to an exponential function of the attenuation coefficient and outputs a second compensated image. The fusion weights are calculated based on the image quality parameters of the original image data. The first compensation image and the second compensation image are then weighted and summed using the fusion weights to obtain the color-corrected image.

2. The method according to claim 1, characterized in that, The steps for calculating the first compensated image and the second compensated image include: The color compensation network adopts an encoder-decoder structure. A constraint layer is set in the middle layer. The constraint layer calculates the pixel-wise product of the original image data and the white balance adjustment parameter, and then the pixel-wise product of the attenuation coefficient exponential function. The calculation result is weighted and fused with the output of the network front layer and then passed to the subsequent layers to output the first compensated image and the first confidence value. The ambient background light component of the corresponding channel is subtracted from each pixel value of the original image data. The subtraction result is multiplied pixel by pixel with the attenuation coefficient exponential function to obtain a preliminary magnified image. Skin color pixel regions in the preliminary magnified image are identified. The white balance adjustment parameters are adjusted and recalculated according to the deviation between the skin color chroma and the standard skin color chroma to obtain a second compensated image and a second confidence value.

3. The method according to claim 1, characterized in that, After acquiring the raw image data and before inputting it into the color compensation network, scattering dehazing processing is also included: Spatial frequency transformation is performed on the original image data to extract the low-frequency component as the initial value of the backscatter component, and the initial value of the direct light component is obtained by subtracting the initial value of the backscatter component. The initial values ​​of the backscatter component, the initial values ​​of the direct light component, the depth information, and the illuminance information are input into the backscatter estimation network, and the refined backscatter component and transmittance correction coefficient are output. Acquire image data from multiple frames at adjacent time points, calculate the pixel displacement field, transform the backscattered component of the next frame to the coordinate system of the current frame, and adjust the refined backscattered component to minimize the sum of squared differences between the transformed backscattered component and the backscattered component of the current frame, thus obtaining the time-stable backscattered component. The time-stable backscatter component is subtracted from the original image data and multiplied pixel by pixel with the reciprocal of the transmittance correction coefficient to obtain the dehazed image; The dark channel statistics of the dehazed image are calculated as the residual suspended particle density index. When the residual suspended particle density index is higher than the preset index threshold, the reciprocal value of the transmittance correction coefficient is reduced, and the adjusted image is used as the original image data input to the color compensation network.

4. The method according to claim 1, characterized in that, The steps for controlling an underwater camera include: Generate the desired velocity and desired angular velocity based on the control commands, and calculate the desired force in three orthogonal directions and the desired torque on three orthogonal axes; Based on the installation position coordinates of multiple thrusters in the body coordinate system and the thrust direction unit vector, calculate the force components of each thruster about the three orthogonal directions and the torque components about the three orthogonal axes, and construct the thrust mapping matrix. Under the conditions of satisfying the equality constraints of the thrust mapping matrix, the desired force and the desired torque, and the upper and lower limits of thrust, find the thrust command vector that minimizes the sum of the squares of thrust. The speed and current feedback of each thruster are detected. When the speed feedback is lower than the preset ratio of the command value or the current feedback exceeds the rated value, the thruster is determined to be faulty. The column corresponding to the faulty thruster is deleted from the thrust mapping matrix and the thrust command vector is re-solved. The elements of the thrust command vector are first-order low-pass filtered and then output to each thruster.

5. The method according to claim 4, characterized in that, The method further includes: Data from gyroscopes, accelerometers, electronic compasses, and depth sensors are collected and fused using an extended Kalman filter to obtain attitude quaternion estimates and angular rate estimates. The estimated value of the external disturbance torque is calculated based on the body inertia matrix, the rate of change of angular rate, the thruster control torque, and the quadratic term of angular rate. The quaternion error between the desired attitude and the estimated attitude is calculated as the attitude error, the difference between the desired angular rate and the estimated angular rate is calculated as the angular rate error, and the negative value of the sum of the proportional term of the attitude error, the differential term of the angular rate error, and the estimated value of the external disturbance torque is used as the attitude control torque. The attitude control torque is superimposed on the desired torque to solve for the thrust command vector.

6. A maneuver control and imaging optimization system for an underwater remotely operated camera, used to implement the method of any one of claims 1-5, characterized in that, include: The first unit is used to control the underwater camera and collect raw image data, depth information, illumination information, color temperature information, and initial values ​​of water body type; The second unit is used to calculate the attenuation coefficients of the red, green and blue components in the original image data respectively, based on the product of the depth information and the initial value of the water body type, superimposed with the correction amount of the illuminance information and the color temperature information. The third unit is used to extract pixel color information with brightness values ​​within a preset range from the original image data as ambient background light, and to calculate white balance adjustment parameters based on pixel distribution. The fourth unit is used to input the original image data, attenuation coefficient, ambient background light, illuminance information and color temperature information into the color compensation network. The color compensation network applies physical constraints based on the attenuation coefficient in the intermediate layer and outputs a first compensated image. After subtracting the ambient background light from the original image data, it is amplified according to the exponential function of the attenuation coefficient and outputs a second compensated image. The fifth unit is used to calculate the fusion weight based on the image quality parameters of the original image data, and to use the fusion weight to perform a weighted summation of the first compensation image and the second compensation image to obtain the color-corrected image.

7. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 5.