Underwater robot adaptive vision enhancement system based on multi-modal feature fusion
The underwater robot adaptive vision enhancement system, which integrates multimodal feature fusion, solves the problem of poor underwater imaging quality, achieves high-quality transmittance maps and color calibration, and can accurately restore image details and colors in complex underwater environments.
Patent Information
- Application Number
- CN202511503307.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-13
AI Technical Summary
The complex underwater optical environment leads to poor image quality, and traditional methods are insufficient to effectively correct the diverse color distortion problems of underwater silt.
An underwater robot adaptive vision enhancement system based on multimodal feature fusion is adopted. The descattered image is processed by a color restoration model and a color classification network. The system combines the multi-scale Retinex algorithm, the underwater dark channel prior algorithm and the adaptive loss function to achieve illumination equalization, scattering optimization and color calibration of the image.
Achieving high-quality transmittance maps and color calibration in complex underwater environments improves image clarity and color fidelity, enabling accurate reproduction of the true geological and ecological features of underwater scenes.
Smart Images

Figure CN121329787A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual enhancement technology, specifically to an adaptive visual enhancement system for underwater robots based on multimodal feature fusion. Background Technology
[0002] Underwater robots are key equipment for marine exploration, resource development, underwater engineering, and scientific research, and their operational performance largely depends on the perception capabilities of their vision systems. However, the underwater optical environment is extremely complex, posing a severe challenge to image quality. First, light undergoes severe absorption and scattering in water. Light absorption is wavelength-selective, with red light attenuating the fastest, resulting in underwater images generally exhibiting a blue-green tint and severely distorted color information. Second, suspended particles in the water cause light scattering, which not only blurs image details but also forms a hazy background light, greatly reducing image contrast and clarity.
[0003] In terms of image enhancement, traditional methods such as histogram equalization, white balance, and wavelet transform can improve the subjective visual effect of images to a certain extent; however, in terms of color correction, a single color correction model is difficult to cope with the diverse colors of underwater silt due to its chemical composition and biological adhesion.
[0004] To address this, an adaptive vision enhancement system for underwater robots based on multimodal feature fusion is proposed. Summary of the Invention
[0005] The purpose of this invention is to provide an adaptive vision enhancement system for underwater robots based on multimodal feature fusion. The system identifies a set of candidate restoration images by using a color restoration model to identify descattered images and ambient light layers; it then identifies the descattered images by using a color classification network to obtain a silt color probability map; and finally, it weights and fuses the candidate restoration images based on the silt color probability map to obtain a color calibration image.
[0006] To achieve the above objectives, an adaptive vision enhancement system for underwater robots based on multimodal feature fusion is proposed, including:
[0007] The equalization optimization module performs histogram analysis on the underwater silt image to obtain the silt brightness distribution; based on the silt brightness distribution, it performs illumination equalization optimization on the silt image to obtain an equalized image.
[0008] The image perception module performs texture analysis on the equalization image to assess the structural complexity of the silt; based on the structural complexity, it configures the underwater dark channel prior algorithm to parse the equalization image and obtain the ambient light vector and initial transmittance map of the water body.
[0009] The scattering optimization module constructs a scattering optimization model including an adaptive loss function based on the structural complexity; the scattering optimization model optimizes the initial transmittance map based on the equalization image to obtain an optimized transmittance map; and performs inverse imaging operations based on the equalization image, the optimized transmittance map, and the ambient light vector to obtain a descattered image.
[0010] The color calibration module extends the ambient light vector into an ambient light layer. It then uses a color restoration model to identify the descattered image and the ambient light layer, resulting in a set of candidate restoration images. A color classification network is used to identify the descattered image, resulting in a silt color probability map. Based on the silt color probability map, the candidate restoration images are weighted and fused to obtain the color calibration image.
[0011] In the illumination equalization optimization of silt images, a multi-scale Retinex algorithm is used, including:
[0012] The brightness distribution of silt in the silt image is obtained, and the Gaussian filter kernel parameters of the multi-scale Retinex algorithm are dynamically adjusted according to the silt brightness distribution. When the silt brightness is lower than a preset threshold, the parameter combination of the large-scale Gaussian filter kernel is used; when the silt brightness is greater than or equal to the preset threshold, the parameter combination of the small-scale Gaussian filter kernel is used.
[0013] The process of evaluating the structural complexity of silt includes: calculating the gradient magnitude map of the equalization image. The structural complexity of the current processing position refers to the average value of the gradient magnitude map within a preset neighborhood centered on the current processing position.
[0014] The process of configuring the underwater dark channel prior algorithm based on structural complexity includes:
[0015] Establish an inverse mapping model between structural complexity and the computation window size of the prior algorithm for underwater dark channels;
[0016] When parsing the equalized image, the structural complexity of the current processing position is calculated; the calculated structural complexity is input into the inverse mapping model to calculate the computation window size for the current processing position.
[0017] The process of obtaining the ambient light vector and initial transmittance map of the water body includes:
[0018] Global color histogram analysis is performed on the balanced image to preliminarily determine the dominant silt color type in the image. Based on the determined silt color type, the color channel on which the calculation of the underwater dark channel map depends is adaptively selected. The silt color types include iron oxide type, sulfide type, and biological type.
[0019] When the silt color type is sulfide type and iron oxide type, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the blue and green color channels within the calculation window of each pixel position;
[0020] When the silt color type is biological, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the red and blue color channels within the calculation window at each pixel position;
[0021] Based on the underwater dark channel map obtained by adaptive calculation, the pixels with the highest brightness value in the map are selected according to a preset ratio, and the corresponding positions are found in the equalized image. The average value of the red, green and blue channels is calculated to determine the ambient light vector. Then, the initial transmittance map is calculated using the underwater dark channel map and the ambient light vector.
[0022] The total loss function of the scattering optimization model is defined as a weighted sum of two sub-loss functions, which include a smoothing loss term and a gradient matching loss term.
[0023] Establish a weight mapping model between the structural complexity and the weight coefficients of the sub-loss function;
[0024] When optimizing the initial transmittance map, the structural complexity of the current processing region is obtained, and the weight coefficients are dynamically identified for the total loss function through the weight mapping model.
[0025] The color restoration model is a collection of multiple parallel and independent conditional generative adversarial networks (GANs), each of which is over-trained for a specific type of silt color. The processing steps include:
[0026] The ambient light vector is expanded into an ambient light layer of the same size as the descattering image;
[0027] The descattered image is stitched together with the ambient light layer to form a conditional input, and the conditional input is simultaneously fed into each conditional generative adversarial network in the color restoration model set to generate a set of candidate restoration images with different color restoration tendencies in parallel.
[0028] The sludge color types include iron oxide type, sulfide type, and biological type.
[0029] The process of acquiring the color calibration image includes:
[0030] The descattered image is subjected to pixel-level soft classification using a pre-trained color classification network to obtain a multi-channel mud color classification probability map that represents the probability of each pixel belonging to different mud color categories.
[0031] Using the silt color classification probability map as weights, the set of candidate restoration images is weighted and summed at the pixel level. The calculation method is that the color value of each final pixel is equal to the sum of the product of the color value of each candidate restoration image at that pixel and the probability of the corresponding silt color category.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0033] 1. This invention uses a neural network model to learn complex nonlinear mapping relationships from data. It not only captures the macroscopic trend of higher complexity and smaller windows, but also learns subtle changes within a specific complexity range, making window size adjustments smoother and more precise. This data-driven approach greatly improves the model's generalization ability and adaptive accuracy, thus obtaining high-quality transmittance maps in various complex underwater terrains.
[0034] 2. This invention identifies and classifies the color types of underwater silt, generating multiple high-quality candidate results with different professional restoration directions. Compared with the average restoration effect that a single model may produce, it can provide richer and more targeted color assumptions and achieve high-fidelity color calibration.
[0035] 3. This solution achieves pixel-level fine-grained color calibration, enabling the correct color restoration of different types of silt within a single image. For example, it can make a rust-rich area appear reddish-brown, while an area covered with algae appears green. This spatially adaptive fusion method makes the final output image extremely realistic and natural, maximizing the restoration of the true geological and ecological features of the underwater scene. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the underwater robot adaptive vision enhancement system based on multimodal feature fusion according to the present invention.
[0037] Figure 2 This is a schematic diagram of the descattering image acquisition process of the present invention;
[0038] Figure 3 This is a schematic diagram of the color calibration image acquisition process of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] Example 1:
[0041] This invention proposes an adaptive vision enhancement system for underwater robots based on multimodal feature fusion. The structure of the system is as follows: Figure 1 As shown, it includes: equalization optimization module, image perception module, scattering optimization module, and color calibration module;
[0042] The equalization optimization module performs histogram analysis on the underwater silt image to obtain the silt brightness distribution; based on the silt brightness distribution, it performs illumination equalization optimization on the silt image to obtain an equalized image.
[0043] In the illumination equalization optimization of silt images, a multi-scale Retinex algorithm is used, including:
[0044] The brightness distribution of silt in the silt image is obtained, and the Gaussian filter kernel parameters of the multi-scale Retinex algorithm are dynamically adjusted according to the silt brightness distribution. When the silt brightness is lower than a preset threshold, the parameter combination of the large-scale Gaussian filter kernel is used; when the silt brightness is greater than or equal to the preset threshold, the parameter combination of the small-scale Gaussian filter kernel is used.
[0045] There is a correlation between illumination characteristics and spatial frequency. Illumination unevenness in underwater dark areas is usually a large-scale, low-frequency gradual change, requiring a large-scale Gaussian kernel to accurately estimate and remove this slowly changing illumination. In contrast, in areas with sufficient brightness, details are more important. In this case, using a large-scale kernel can easily produce halo artifacts, while using a small-scale kernel can better preserve local contrast and details.
[0046] After inputting the original underwater silt image, its brightness components are first calculated to obtain a global brightness histogram or a local brightness mean, i.e., the silt brightness distribution. A brightness threshold is set (e.g., 60 within the range of 0-255). When processing each region of the image, its average brightness is judged. If it is below 60, it is identified as a dark area, and the system calls a preset large-scale Gaussian filter kernel parameter combination to execute the multi-scale Retinex algorithm; if the brightness is greater than or equal to 60, it is identified as a normal or highlight area, and the system switches to a small-scale Gaussian filter kernel parameter combination. The algorithm finally outputs a balanced image after illumination equalization.
[0047] Among them, the multi-scale Retinex algorithm decomposes the image into illumination and reflection components by simulating the color constancy of the human visual system, and removes the influence of the illumination component.
[0048] This invention uses brightness recognition in underwater silt images to perform adaptive illumination equalization processing. It can effectively brighten dark areas and restore details while accurately maintaining the natural contrast of bright areas, significantly improving the overall dynamic range and visual quality of the image, and providing high-quality input for subsequent descattering and color correction.
[0049] The image perception module performs texture analysis on the equalization image to assess the structural complexity of the silt; based on the structural complexity, it configures an underwater dark channel prior algorithm to parse the equalization image and obtain the ambient light vector and initial transmittance map of the water body.
[0050] The process of evaluating the structural complexity of silt includes: calculating the gradient magnitude map of the equalization image. The structural complexity of the current processing position refers to the average value of the gradient magnitude map within a preset neighborhood centered on the current processing position.
[0051] Image gradient is a metric that measures the rate of change of pixel values. Regions with rich texture and sharp edges exhibit rapid changes in pixel values, resulting in correspondingly higher gradient magnitudes; while regions such as water bodies and flat mud exhibit gradual changes in pixel values, leading to lower gradient magnitudes. Therefore, the average gradient magnitude within a neighborhood can intuitively and effectively quantify the texture details and structural information of that region.
[0052] The input equalized image is processed using gradient operators such as Sobel or Prewitt to calculate the gradient maps in the x and y directions. Then, the gradient magnitude of each pixel is calculated to form a gradient magnitude map. For any current processing position in the image, a predefined neighborhood (e.g., a 15×15 pixel window) is defined centered on that position, and the average gradient magnitude of all pixels within that window is calculated and defined as the structural complexity of the current position. A structural complexity map of the same size as the original image is then generated.
[0053] This solution provides an efficient image content quantization method that transforms abstract and complex concepts into concrete quantized data, enabling the system to perceive the local characteristics of the image. This is the key basis for achieving subsequent adaptive processing, such as dynamically adjusting the window size and dynamically configuring the loss function.
[0054] The process of configuring the underwater dark channel prior algorithm based on structural complexity includes:
[0055] Establish an inverse mapping model between structural complexity and the computation window size of the prior algorithm for underwater dark channels;
[0056] When parsing the equalized image, the structural complexity of the current processing position is calculated; the calculated structural complexity is input into the inverse mapping model to calculate the computation window size for the current processing position.
[0057] In this embodiment, the inverse mapping model is a pre-trained lightweight neural network, such as a multilayer perceptron with two hidden layers. The network structure is as follows:
[0058] Input layer: 1 neuron, receiving a single scalar value of structural complexity.
[0059] Hidden layers: 2 hidden layers, each containing 16 neurons, using non-linear activation functions such as ReLU.
[0060] Output layer: 1 neuron, outputting a floating-point value representing the size of the prediction calculation window.
[0061] Offline training phase of the model:
[0062] Dataset Construction: Collect a large number of underwater images with different silt landform features. Divide these images into a large number of overlapping image blocks (e.g., 64×64 pixels).
[0063] Feature and Label Generation: For each image patch, its structural complexity is first calculated (as an input feature of the model). Then, the "optimal window size" for that image patch (as the ground truth label of the model) is determined through an optimization process or by annotation by domain experts. For example, multiple different window sizes can be tried for each image patch, and the best-performing size can be selected as the label based on metrics such as the sharpness and contrast of the restored image.
[0064] Model training: Using the constructed "feature-label" dataset, the MLP network is trained through backpropagation and gradient descent optimizer so that it can accurately predict the optimal window size based on the structural complexity of the input.
[0065] The final result is a calculation window size map, where the value of each pixel represents the specific window side length to be used when performing dark channel calculations at that location.
[0066] This invention utilizes a neural network model to learn more complex and nonlinear mapping relationships from data. It not only captures the macroscopic trend of higher complexity and smaller windows, but also learns subtle changes within specific complexity ranges, making window size adjustments smoother and more precise. This data-driven approach significantly improves the model's generalization ability and adaptive accuracy, enabling the acquisition of high-quality transmittance maps in various complex underwater terrains.
[0067] The process of obtaining the ambient light vector and initial transmittance map of the water body includes:
[0068] Global color histogram analysis is performed on the balanced image to preliminarily determine the dominant silt color type in the image. Based on the determined silt color type, the color channel on which the calculation of the underwater dark channel map depends is adaptively selected. The silt color types include iron oxide type, sulfide type, and biological type.
[0069] When the silt color type is sulfide type and iron oxide type, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the blue and green color channels within the calculation window of each pixel position;
[0070] When the silt color type is biological, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the red and blue color channels within the calculation window at each pixel position;
[0071] Based on the underwater dark channel map obtained by adaptive calculation, the pixels with the highest brightness value in the map are selected according to a preset ratio, and the corresponding positions are found in the equalized image. The average value of the red, green and blue channels is calculated to determine the ambient light vector. Then, the initial transmittance map is calculated using the underwater dark channel map and the ambient light vector.
[0072] The sludge color types include iron oxide type, sulfide type, and biological type.
[0073] Iron oxide type: This type of sludge usually forms in oxidizing environments, that is, areas in the water body with relatively abundant dissolved oxygen. Ferrous ions in the water are oxidized to ferric ions and deposited in the form of ferric hydroxide or ferric oxide, giving the sludge a warm color ranging from yellow and brown to reddish brown. Its physical form is mostly clay or silt.
[0074] Because water absorbs long-wavelength red and yellow light most intensely, the characteristic color information of iron oxide silt attenuates most rapidly during propagation. In underwater cameras, its true reddish-brown color is easily lost, "submerged" by the blue-green of ambient light, causing severe color distortion and posing significant challenges to color-based geological composition analysis.
[0075] Sulfide type: This type of sludge is a typical sign of anoxic or anaerobic environments and is commonly found in estuaries, deep pools or the bottom of lakes polluted by organic matter. Under anoxic conditions, microorganisms produce sulfides when they decompose organic matter. The sulfides react with metal ions in the water and sediment to form black ferrous sulfide and other metal sulfides, making the sludge appear dark gray to pure black.
[0076] Black sulfide sludge has extremely low light reflectivity, and most of the incident light is absorbed by it. This results in extremely weak effective signals received by the imaging sensor, and the overall image is dark with a low signal-to-noise ratio. In such images, key details such as the minute textures and topographical undulations on the sludge surface are almost completely submerged in noise, and conventional enhancement algorithms tend to amplify the noise rather than restore the details.
[0077] Biofilms: Common biofilms are formed by photosynthetic organisms such as algae, cyanobacteria, or benthic diatoms. The color of this type mainly comes from the biological community covering the surface of the muddy substrate, rather than the mud itself. Depending on the dominant species, the color can be bright green, blue-green, or yellowish-brown. The color of the biofilm (especially green) is very close in hue to the blue-green hue of the water itself, resulting in low color contrast with the surrounding environment, blurred boundaries, and difficulty in precise segmentation. In addition, these biofilms usually have a clear physical boundary with the underlying muddy substrate. Inappropriate image processing (such as smoothing and noise reduction or color blending) can easily blur this important ecological boundary, losing key information for assessing the health of benthic ecosystems.
[0078] For biological scenarios, by actively avoiding the green channel, which is heavily affected by pigments such as chlorophyll, and instead using the relatively "purer" red and blue channels, the accuracy of estimating the optical parameters (ambient light, transmittance) of the real water body can be significantly improved. This results in higher quality front-end input (initial transmittance map) for the entire visual enhancement process, providing more reliable physical quantities for subsequent descattering and color calibration modules. As a result, the sharpness and color fidelity of the final enhanced image are systematically improved, especially when dealing with complex ecological scenarios rich in aquatic plants or algae.
[0079] The input consists of an equalized image and an adaptive computation window. For each pixel, based on its corresponding window size, only the pixel values of the corresponding channel are traversed within that window, and the minimum value is found. This minimum value is used as the pixel's value in the underwater dark channel image. After generating a complete single-channel dark channel image, the pixels with the highest brightness in the image are selected according to a preset ratio. Next, the corresponding positions of these brightest pixels are found in the equalized image, and the average value of the red, green, and blue channel pixels at these positions is calculated. This average value is then determined as the global ambient light vector.
[0080] Finally, based on the underwater imaging model and the calculated ambient light vector, an initial transmittance map is obtained. Further, the calculation of the initial transmittance map is based on a physical model of underwater optical imaging, aiming to quantify the attenuation of light propagating in water. Using a known equalized image, the estimated global ambient light vector, and an underwater dark channel map calculated based on dark channel prior theory, the transmittance describing scene sharpness is solved inversely. Specifically, the transmittance is estimated as the difference between the underwater dark channel map values after ambient light vector normalization. To ensure the naturalness and robustness of the restoration effect, an adjustment factor is introduced in the calculation to retain a small amount of natural fog, and a reasonable lower limit is set for the transmittance to avoid noise and distortion due to numerical instability in dark or distant areas. The resulting initial transmittance map provides pixel-level depth and sharpness information for subsequent image descattering and color calibration steps.
[0081] This scheme significantly improves the reliability of the dark channel map and the accuracy of the ambient light vector estimation by adopting a dual-channel dark channel calculation that conforms to the physical characteristics of underwater. This results in a more accurate initial transmittance map, laying a solid foundation for subsequent descattering and image restoration. It can also more effectively remove the blue-green bias and turbidity in underwater images.
[0082] The scattering optimization module constructs a scattering optimization model including an adaptive loss function based on structural complexity. This optimization model optimizes the initial transmittance map based on the equalization image, resulting in an optimized transmittance map. An inverse imaging operation is then performed using the equalization image, the optimized transmittance map, and the ambient light vector to obtain the descattered image. The process for acquiring the descattered image is as follows: Figure 2 As shown.
[0083] The total loss function of the scattering optimization model is defined as a weighted sum of two sub-loss functions, which include a smoothing loss term and a gradient matching loss term.
[0084] Establish a weight mapping model between the structural complexity and the weight coefficients of the sub-loss function;
[0085] When optimizing the initial transmittance map, the structural complexity of the current processing region is obtained, and the weight coefficients are dynamically identified for the total loss function through the weight mapping model.
[0086] Optimizing transmittance maps requires striking a balance between smoothing and preserving edge details. In areas with flat structures, smoothing should be prioritized; while in areas with complex textures, the focus should be on maintaining the original texture. Figure 1 The proposed solution uses a neural network to learn this complex trade-off, enabling the loss function to intelligently adapt to different image content.
[0087] Specifically, the weight mapping model is a pre-trained neural network, such as a multilayer perceptron (MLP). The network's input is the average structural complexity of the image region, and its output is two weight coefficients: smoothing loss weights and gradient matching loss weights. When optimizing the initial transmissivity map, the image is divided into blocks; for each block, its structural complexity is calculated and input into the weight mapping model to obtain a unique weight combination for that block, thus defining the total loss function for that block; the optimizer iteratively optimizes the transmissivity map based on this dynamically calculated loss function.
[0088] Furthermore, the scattering optimization module classifies water scattering sources into global background scattering and local disturbance scattering. The optimization process includes: identifying locally high-concentration suspended clouds formed by the robot itself and the bottom undercurrent stirring up the silt by analyzing the changes between image frames and the local contrast anomalies of a single frame image; for the suspended clouds, a local transmittance correction layer is superimposed on the global transmittance map. This layer is estimated based on the concentration and thickness of the clouds, and a precise physical inverse operation is performed on the non-uniform local scattering.
[0089] First, an initial transmittance map representing the global water turbidity is obtained through standard underwater dark channel prior calculations. Then, based on a local disturbance detector, regions with characteristics of "sharp edges, blurred internal texture, and extremely low contrast" are identified to determine suspended sediment clouds, and a separate local transmittance map is estimated for each. This estimation is based on the color saturation of the region (sediment clouds typically have lower saturation) or through a small neural network. Finally, the total transmittance is applied to this region.
[0090] This solution directly addresses the most common and visually impactful problem encountered by underwater robots during operations: "stirred sediment." Existing models typically assume uniform turbidity in the water, failing to handle sudden, localized, and extremely turbid conditions. This invention establishes a two-layer scattering model, decoupling background scattering from near-field perturbation scattering, significantly improving adaptability and reconstruction accuracy in dynamic, non-uniformly turbid scenarios. This allows the robot to obtain clear visual feedback even when performing tasks that easily stir up bottom sediment, such as close-range observation and sampling.
[0091] This scheme, based on a weighted mapping model, achieves dynamic adaptation of the loss function, making the optimization process more refined and intelligent. It ensures smooth and noise-free transmittance in flat silt areas, while accurately preserving the contours and texture details in complex areas with organisms or rocks. The resulting optimized transmittance map is of higher quality, and the restored descattered image is clearer and more natural.
[0092] Furthermore, before constructing the scattering optimization model, the scattering optimization module also includes silt ripple structure identification; by performing frequency domain analysis or directional gradient analysis on the equalized image, it detects whether there is a quasi-periodic silt ripple structure formed by water flow in the image; if such a structure is detected, the gradient matching loss term in the scattering optimization model will be replaced by an anisotropic gradient loss term, which will apply a higher weight along the normal direction of the ripple, i.e., the direction of the most drastic gradient change, and a lower weight along the tangent direction of the ripple, i.e., the direction of texture direction.
[0093] Standard gradient matching treats gradients in all directions equally. When dealing with highly directional wavy structures, it is prone to blurring the sharp edges of the wavy lines due to noise interference. This invention, by identifying this specific structure and applying directional optimization constraints, can significantly enhance the three-dimensionality and clarity of the wavy lines while improving image contrast, making the transition between light and dark areas on the backlit and frontlit sides more natural and realistic.
[0094] The color calibration module extends the ambient light vector into an ambient light layer. It then uses a color restoration model to identify the descattered image and the ambient light layer, resulting in a set of candidate restoration images. A color classification network is used to identify the descattered image, resulting in a silt color probability map. Based on the silt color probability map, the candidate restoration images are weighted and fused to obtain the color calibration image.
[0095] The process of acquiring color calibration images is as follows: Figure 3 As shown.
[0096] The color restoration model is a collection of multiple parallel and independent conditional generative adversarial networks (GANs), each of which is over-trained for a specific type of silt color. The processing steps include:
[0097] The ambient light vector is expanded into an ambient light layer of the same size as the descattering image;
[0098] The descattered image is stitched together with the ambient light layer to form a conditional input, and the conditional input is simultaneously fed into each conditional generative adversarial network in the color restoration model set to generate a set of candidate restoration images with different color restoration tendencies in parallel.
[0099] The color of underwater silt varies greatly due to differences in its chemical and biological composition, making it difficult for a single remediation model to handle all situations. This solution employs a hybrid expert strategy, allowing each conditional generative adversarial network to specialize in the remediation of a single color type, thus decomposing the complex global remediation problem into several simpler, more specialized sub-problems. Using ambient light as a conditional input provides the network with crucial prior information about water color and lighting conditions, helping it to better understand and reverse the color decay process.
[0100] The specific process is as follows: The system pre-trains three independent conditional generative adversarial networks (e.g., Pix2Pix architecture). The first conditional generative adversarial network is trained on the data pair of "descattered image - iron oxide sludge true color image", the second on the data pair of sulfide type, and the third on the data pair of biological type.
[0101] In actual processing, the system receives the "descattered image" and copies and expands the obtained "ambient light vector" into an "ambient light layer" of the same size as the image. Then, the descattered image (3 channels) and the ambient light layer (3 channels) are concatenated along the channel dimension to form a 6-channel input tensor. This tensor is simultaneously fed into three conditional generative adversarial networks for parallel processing, each outputting candidate restoration images with specific color tendencies, which together constitute a set of candidate results.
[0102] This invention identifies and classifies the color types of underwater silt, generating multiple high-quality candidate results with different professional restoration directions. Compared with the average restoration effect that a single model may produce, it can provide richer and more targeted color assumptions and achieve high-fidelity color calibration.
[0103] The process of acquiring the color calibration image includes:
[0104] The descattered image is subjected to pixel-level soft classification using a pre-trained color classification network to obtain a multi-channel mud color classification probability map that represents the probability of each pixel belonging to different mud color categories.
[0105] Using the silt color classification probability map as weights, the set of candidate restoration images is weighted and summed at the pixel level. The calculation method is that the color value of each final pixel is equal to the sum of the product of the color value of each candidate restoration image at that pixel and the probability of the corresponding silt color category.
[0106] Furthermore, a pre-trained semantic segmentation network is used as the color classification network, which can classify input image pixels into three categories: "iron oxides," "sulfides," and "biological." The "descattered image" is input into this network, and the network outputs a 3-channel probability map, with each channel representing the probability of each pixel belonging to its corresponding category. Based on this probability map, the three candidate images generated in the previous step are summed at the pixel level to obtain the final color-calibrated image.
[0107] This solution achieves pixel-level fine-grained color calibration, enabling the correct color restoration of different types of silt within a single image. For example, it can make a rust-rich area appear reddish-brown, while an area covered with algae appears green. This spatially adaptive fusion method makes the final output image extremely realistic and natural, maximizing the reproduction of the true geological and ecological features of the underwater scene.
[0108] Furthermore, when the color calibration module performs weighted fusion of candidate restoration images based on the silt color probability map, it also introduces a boundary sharpening constraint, which includes:
[0109] First, an edge detection algorithm is run on the brightness channel of the descattered image to identify high-contrast contours. If a contour highly coincides with the transition boundary between "biological type" and "non-biological type" in the color classification probability map, the system will sharpen the fusion weights when fusing pixels near the boundary. That is, the weights are forcibly normalized to 0 or 1 to achieve a "hard" transition between the biological film area and the bare silt substrate, rather than a blurry color gradient.
[0110] Before performing pixel-level weighted summation, the Canny edge detection operator is first applied to the descattered image to obtain a binary edge map. Simultaneously, the "biological type" channel in the color classification probability map is thresholded to obtain a binary mask of "biological overlay." The boundary of this mask is then ANDed with the Canny edge map. If a boundary exists in both maps at a certain pixel location, the system considers this a "real physical boundary" that needs protection. Subsequently, during weighted fusion, the system modifies the fusion weights for these marked boundary pixels and their neighborhoods; for example, if the biological probability is >0.5, the weight is forced to 1.0; if <0.5, it is forced to 0.0. This results in a clear switch between colors generated by the biological inpainting conditional generative adversarial network and colors generated by other conditional generative adversarial networks, rather than a smooth transition.
[0111] This solution addresses the ecological characteristic of underwater sediment surfaces often covered with "biofilms" such as algae and bacterial mats. These biofilms typically have a clear physical boundary with the underlying sediment substrate; traditional smoothing weighted fusion blurs this boundary, producing unrealistic color transitions and losing ecological information. This invention, by actively detecting and forcibly sharpening the color transitions of this geoecological boundary zone, can extremely realistically restore the true morphology and coverage of biofilms, preserving boundary information crucial for assessing the health of benthic ecosystems. This results in enhanced outcomes that achieve a new level of fidelity not only visually but also scientifically.
[0112] Example 2: This invention proposes an adaptive vision enhancement system for underwater robots based on multimodal feature fusion, comprising:
[0113] The equalization optimization module performs histogram analysis on the underwater silt image to obtain the silt brightness distribution; based on the silt brightness distribution, it performs illumination equalization optimization on the silt image to obtain an equalized image.
[0114] The global average brightness of the silt image is calculated, and a brightness threshold is set to distinguish between dark and bright areas. In this embodiment, the preset threshold is set to 70 (in a gray level of 0-255), and the preferred range of the threshold is [50, 80].
[0115] Then, the multi-scale Retinex algorithm is applied based on the regional brightness. The specific parameter combination of the Gaussian filter kernel is as follows:
[0116] When the average brightness of a region is below 70, it is identified as a dark area, and a parameter combination of large-scale Gaussian filter kernels is used. This combination contains three Gaussian kernels with scales (standard deviations σ) of 30, 150, and 300, respectively.
[0117] When the average brightness of a region is greater than or equal to 70, it is classified as a normal or highlight area, and a parameter combination of small-scale Gaussian filter kernels is used. This combination contains three Gaussian kernels with scales of 15, 80, and 220, respectively.
[0118] In both combinations, the weighting coefficients for the three scale components are all set to 1 / 3. After processing, the output is an equalized image with illumination equalization.
[0119] The image perception module performs texture analysis on the equalization image to assess the structural complexity of the silt; based on the structural complexity, it configures an underwater dark channel prior algorithm to parse the equalization image and obtain the ambient light vector and initial transmittance map of the water body.
[0120] To further evaluate the structural complexity of the silt, the Sobel operator was used to calculate the gradient magnitude map of the equalization image. When calculating the structural complexity at the current processing location, a square window of size 11×11 pixels was used as the preset neighborhood.
[0121] Next, an inverse mapping model is established between structural complexity and the computation window size of the underwater dark channel prior algorithm. In this embodiment, the model is a multilayer perceptron (MLP) containing two hidden layers (16 neurons per layer). The training process of this model is as follows:
[0122] Dataset construction: Collect images containing different underwater landforms and crop them into overlapping image patches of 64×64 pixels.
[0123] Feature label generation:
[0124] Features: For each image patch, calculate its average gradient magnitude in an 11×11 neighborhood, which is used as the input feature for “structural complexity”.
[0125] Label: To determine the "optimal window size" for each image patch, an automated optimization method was employed. Specifically, the underwater dark channel prior algorithm was executed on each image patch using a series of odd-sized windows ranging from 3×3 to 21×21 (with a step size of 2), and the Underwater Color Image Quality Evaluation (UCIQE) score was calculated for each restored result. The window size that yielded the highest UCIQE score was used as the ground truth label for the "optimal window size" of that image patch.
[0126] Model training: Using the constructed “feature label” dataset, the MLP network was trained for 100 epochs with the Adam optimizer, a learning rate of 1e-4 and a mean squared error (MSE) loss function, until the model converged.
[0127] The process of obtaining the ambient light vector and initial transmittance map of the water body includes:
[0128] Global color histogram analysis is performed on the balanced image to preliminarily determine the dominant silt color type in the image. Based on the determined silt color type, the color channel on which the calculation of the underwater dark channel map depends is adaptively selected. The silt color types include iron oxide type, sulfide type, and biological type.
[0129] When the silt color type is sulfide type and iron oxide type, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the blue and green color channels within the calculation window of each pixel position;
[0130] When the silt color type is biological, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the red and blue color channels within the calculation window at each pixel position;
[0131] Based on the underwater dark channel map obtained by adaptive calculation, the pixels with the highest brightness value in the map are selected according to a preset ratio, and the corresponding positions are found in the equalized image. The average value of the red, green and blue channels is calculated to determine the ambient light vector. Then, the initial transmittance map is calculated using the underwater dark channel map and the ambient light vector.
[0132] When acquiring the ambient light vector and initial transmittance map of the water body, the underwater dark channel map is first calculated. Then, a preset proportion of pixels with the highest brightness values in the underwater dark channel map are selected. In this embodiment, this preset proportion is set to 0.2%. Finally, the ambient light vector is determined based on the average R / G / B values of these pixels at their corresponding positions in the equalized image, and the initial transmittance map is calculated.
[0133] The scattering optimization module constructs a scattering optimization model including an adaptive loss function based on the structural complexity; the scattering optimization model optimizes the initial transmittance map based on the equalization image to obtain an optimized transmittance map; and performs inverse imaging operations based on the equalization image, the optimized transmittance map, and the ambient light vector to obtain a descattered image.
[0134] The system receives the equalized image, initial transmittance map, and ambient light vector, and establishes a weight mapping model between the structural complexity and the weight coefficients of the sub-loss function. This model is also an MLP (Multi-Level Processing). Its training process is as follows:
[0135] Dataset construction: Use the same image patches and their corresponding "structural complexity" features as in the previous module.
[0136] Label Generation: A grid search method is used to determine the optimal weight coefficients (smoothing loss weights w and gradient matching loss weights) for each image patch. For each weight combination, the initial transmittance map of the image patch is optimized. The optimization results are evaluated through a joint evaluation function, which aims to reward gradient matching, the structural similarity (SSIM) of the gradient maps of the restored image patch and the original image patch, while penalizing noise in flat regions and the variance of transmittance in flat regions. The weight combination that scores the highest in this joint evaluation function is used as the ground truth label for the "optimal weight coefficients" of that image patch.
[0137] Model training: Training is performed using parameters similar to those of the inverse mapping model described above. After obtaining the optimized transmittance map, the descattered image is obtained through inverse imaging operations.
[0138] The color calibration module extends the ambient light vector into an ambient light layer. It then uses a color restoration model to identify the descattered image and the ambient light layer, resulting in a set of candidate restoration images. A color classification network is used to identify the descattered image, resulting in a silt color probability map. Based on the silt color probability map, the candidate restoration images are weighted and fused to obtain the color calibration image.
[0139] The color restoration model is a collection of multiple parallel and independent conditional generative adversarial networks (GANs). This embodiment employs three Pix2Pix architecture GANs, corresponding to three types of sludge: iron oxide, sulfide, and biological sludge. The training dataset is constructed as follows:
[0140] Generating paired training data: Using real silt color images as input, simulation software based on the Jaffe-McGlamery underwater optical model is used to render a large number of images with blue-green tint and blurring effects by setting different water attenuation and scattering coefficients. These simulated images serve as input to a conditional generative adversarial network, while the original clear water tank images serve as its target output, thus forming a paired "input-output" training dataset.
[0141] The descattered image is subjected to pixel-level soft classification using a pre-trained color classification network.
[0142] Finally, using the silt color probability map output by the color classification network as weights, the candidate restoration images generated by the three cGANs are summed at the pixel level to obtain the final color calibration image.
[0143] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An adaptive vision enhancement system for underwater robots based on multimodal feature fusion, characterized in that, include: The equalization optimization module performs histogram analysis on underwater silt images to obtain the silt brightness distribution. Based on the brightness distribution of silt, illumination equalization optimization is performed on the silt image to obtain a balanced image; The image perception module performs texture analysis on the balanced image to assess the structural complexity of the silt. Based on the structural complexity, an underwater dark channel prior algorithm is configured to analyze the equalization image and obtain the ambient light vector and initial transmittance map of the water body. The scattering optimization module constructs a scattering optimization model, including an adaptive loss function, based on the structural complexity. The scattering optimization model optimizes the initial transmittance map based on the equalization image to obtain an optimized transmittance map; and performs inverse imaging operations based on the equalization image, the optimized transmittance map, and the ambient light vector to obtain a descattered image. The color calibration module extends the ambient light vector into an ambient light layer. It then uses a color restoration model to identify the descattered image and the ambient light layer, resulting in a set of candidate restoration images. A color classification network is used to identify the descattered image, resulting in a silt color probability map. Based on the silt color probability map, the candidate restoration images are weighted and fused to obtain the color calibration image.
2. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 1, characterized in that: In the illumination equalization optimization of silt images, a multi-scale Retinex algorithm is used, including: The brightness distribution of silt in the silt image is obtained, and the Gaussian filter kernel parameters of the multi-scale Retinex algorithm are dynamically adjusted according to the silt brightness distribution. When the silt brightness is lower than a preset threshold, the parameter combination of the large-scale Gaussian filter kernel is used; when the silt brightness is greater than or equal to the preset threshold, the parameter combination of the small-scale Gaussian filter kernel is used.
3. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 1, characterized in that, The process of evaluating the structural complexity of silt includes: calculating the gradient magnitude map of the equalization image. The structural complexity of the current processing position refers to the average value of the gradient magnitude map within a preset neighborhood centered on the current processing position.
4. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 1, characterized in that: The process of configuring the underwater dark channel prior algorithm based on structural complexity includes: Establish an inverse mapping model between structural complexity and the computation window size of the prior algorithm for underwater dark channels; When parsing the equalized image, the structural complexity of the current processing position is calculated; the calculated structural complexity is input into the inverse mapping model to calculate the computation window size for the current processing position.
5. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 4, characterized in that: The process of obtaining the ambient light vector and initial transmittance map of the water body includes: Global color histogram analysis is performed on the balanced image to preliminarily determine the dominant silt color type in the image. Based on the determined silt color type, the color channel on which the calculation of the underwater dark channel map depends is adaptively selected. The silt color types include iron oxide type, sulfide type, and biological type. When the silt color type is sulfide type and iron oxide type, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the blue and green color channels within the calculation window of each pixel position; When the silt color type is biological, the calculation of the underwater dark channel map is based on obtaining the minimum value of all pixel values of the red and blue color channels within the calculation window at each pixel position; Based on the underwater dark channel map obtained by adaptive calculation, the pixels with the highest brightness value in the map are selected according to a preset ratio, and the corresponding positions are found in the equalized image. The average value of the red, green and blue channels is calculated to determine the ambient light vector. Then, the initial transmittance map is calculated using the underwater dark channel map and the ambient light vector.
6. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 1, characterized in that: The total loss function of the scattering optimization model is defined as a weighted sum of two sub-loss functions, which include a smoothing loss term and a gradient matching loss term. Establish a weight mapping model between the structural complexity and the weight coefficients of the sub-loss function; When optimizing the initial transmittance map, the structural complexity of the current processing region is obtained, and the weight coefficients are dynamically identified for the total loss function through the weight mapping model.
7. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 1, characterized in that: The color restoration model is a collection of multiple parallel and independent conditional generative adversarial networks (GANs), each of which is over-trained for a specific type of silt color. The processing steps include: The ambient light vector is expanded into an ambient light layer of the same size as the descattering image; The descattered image is stitched together with the ambient light layer to form a conditional input, and the conditional input is simultaneously fed into each conditional generative adversarial network in the color restoration model set to generate a set of candidate restoration images with different color restoration tendencies in parallel. The sludge color types include iron oxide type, sulfide type, and biological type.
8. The underwater robot adaptive vision enhancement system based on multimodal feature fusion according to claim 1, characterized in that, The process of acquiring the color calibration image includes: The descattered image is subjected to pixel-level soft classification using a pre-trained color classification network to obtain a multi-channel mud color classification probability map that represents the probability of each pixel belonging to different mud color categories. Using the silt color classification probability map as weights, the set of candidate restoration images is weighted and summed at the pixel level. The calculation method is that the color value of each final pixel is equal to the sum of the product of the color value of each candidate restoration image at that pixel and the probability of the corresponding silt color category, thus obtaining the color calibration image.
Citation Information
Cited By
Titanium scrap impurity sorting monitoring method based on image processing
CN122265984A