An adaptive underwater image enhancement method based on multi-modal evaluation

By using multimodal evaluation and closed-loop collaboration of an unsupervised operator library, and dynamically selecting combinations of enhancement operators, the adaptability problem of underwater image enhancement methods under different water conditions and depths is solved, achieving adaptive and interpretable underwater image enhancement effects.

CN122492476APending Publication Date: 2026-07-31KEXI RIEMANN INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN Β· China
Patent Type
Applications(China)
Current Assignee / Owner
KEXI RIEMANN INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-04-28
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods have poor adaptability under different water conditions and depths, lack adaptability, are difficult to handle complex degradation, and rely on paired training data or reference images, lacking real-time diagnosis and iterative optimization capabilities.

Method used

An adaptive underwater image enhancement method based on multimodal evaluation is adopted. Through the closed-loop collaboration of an unsupervised operator library and a large multimodal model, the combination of enhancement operators is dynamically selected, image quality analysis and iterative optimization are performed, and the optimal enhancement scheme is generated.

Benefits of technology

It achieves adaptive underwater image enhancement without the need for paired training data, adapts to different water conditions and depths, has automatic deviation correction capabilities, reduces data acquisition and model training costs, and has highly interpretable outputs, making it suitable for underwater robots and online monitoring systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492476A_ABST
    Figure CN122492476A_ABST
Patent Text Reader

Abstract

This invention relates to an adaptive underwater image enhancement method based on multimodal evaluation, comprising the following steps: acquiring the underwater image to be enhanced and calculating several key indicators; mapping each key indicator to the [0,1] interval using an adaptive normalization algorithm to obtain the severity of each key indicator; selecting the problem type corresponding to the key indicator with the highest severity as the primary problem type; constructing an unsupervised operator library and a candidate pipeline library, and processing the solutions in the candidate pipeline library separately and in parallel to obtain the enhancement result image corresponding to each candidate enhancement solution; defining the evaluation dimension and the weight of the dimension according to the primary problem type, and obtaining the current optimal candidate solution through a multimodal large model; preset reward value threshold and iteration threshold, determining whether the current optimal candidate solution is the final optimal solution, and then obtaining and outputting the final optimal solution. This invention is applicable to image enhancement processing of various underwater images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater image enhancement technology, and in particular to an adaptive underwater image enhancement method based on multimodal evaluation. Background Technology

[0002] Underwater imaging is affected by multiple factors, including light absorption, scattering, occlusion by suspended particles, spectral attenuation at different depths, and non-uniform illumination. This results in images that commonly exhibit complex degradation phenomena such as turbidity, low contrast, blue-green tint, insufficient brightness, and loss of detail. Because the imaging mechanisms differ across sea areas, lakes, reservoirs, nearshore sedimentary waters, and deep-sea low-light environments, the same enhancement algorithm often performs significantly differently in different scenarios.

[0003] Traditional underwater image enhancement methods are mainly divided into three categories, all of which have obvious limitations: (1) Physical model-based methods, such as dark channel prior, underwater dark channel prior, fusion methods, etc. These methods have complex parameter estimation, long computation time, and poor adaptability to different water environments; (2) Non-physical model-based methods, such as adaptive histogram equalization, Retinex theory, white balance, etc. Most of these methods focus on a single problem and lack the ability to comprehensively process complex degradation scenes. They are effective in some waters, but may introduce side effects such as over-enhancement, local overexposure, color distortion, noise amplification, or artifacts in other types of waters; (3) Deep learning-based methods, such as Water-Net, UEBNet, etc. These methods rely on a large number of pairs of underwater degradation images and clear reference images, which are costly to obtain and lack the ability to generalize across waters and interpretability.

[0004] Furthermore, existing automatic enhancement schemes typically only determine a fixed model in the offline stage, lacking the ability to perform real-time diagnosis, dynamically select operators, conduct closed-loop evaluation, and iterative optimization based on the current input image quality status. They also lack intelligent judgment on whether the side effects of enhancement truly affect target recognition, making it difficult to balance sharpness enhancement, contrast enhancement, and color naturalness.

[0005] Therefore, how to achieve an underwater image adaptive enhancement method that does not require paired training data, can adapt to different aquatic environments and depth conditions, handles complex degradation, and achieves adaptive enhancement has become a key problem that urgently needs to be solved in the current technical field. Summary of the Invention

[0006] In view of the above-mentioned problems in the existing technology, the technical problem to be solved by the present invention is: how to achieve adaptive enhancement of images in complex underwater environments.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0008] An adaptive underwater image enhancement method based on multimodal evaluation includes the following steps:

[0009] S1: Acquire underwater images to be enhanced and calculate Several key indicators are used to map each key indicator to the [0,1] interval through an adaptive normalization algorithm to obtain the normalized value of each key indicator. The normalized value of each key indicator is used as the severity of each key indicator, and the severity of each key indicator corresponds to a problem type.

[0010] S2: Select the problem type corresponding to the key indicator with the highest severity as... The main types of problems;

[0011] S3: Construct an unsupervised operator library and a candidate pipeline library. The unsupervised enhancement operator library includes M' enhancement operators, and the candidate pipeline library includes N' candidate enhancement schemes. Each candidate enhancement scheme is composed of one or more operators connected in sequence. Each candidate enhancement scheme is defined as a candidate pipeline.

[0012] S4: Utilize N' candidate enhancement schemes separately and in parallel... Enhancement processing is performed to obtain the enhanced result image corresponding to each candidate enhancement scheme, resulting in a total of N' enhanced result images. The enhancement processing expression is as follows:

[0013]

[0014] in, This represents the enhanced image corresponding to the k-th candidate enhancement scheme. This represents the i-th enhancement operator among the k-th candidate schemes. This indicates that the total number of candidate solutions in the k-th candidate solution is... An enhancement operator;

[0015] S5: Define the evaluation dimensions and their weights based on the main problem types. The corresponding N' augmentation result images are input into the multimodal large model. The multimodal large model analyzes each augmentation result image according to the dimension weight to obtain the comprehensive reward value, remaining problems, improvement directions and analysis report for each augmentation result image. All augmentation result images are sorted in descending order according to the comprehensive reward value, and the candidate augmentation scheme corresponding to the highest comprehensive reward value is taken as the current optimal candidate scheme.

[0016] S6: Preset reward value threshold and iteration threshold. If the reward value of the current best candidate solution is greater than or equal to the reward value threshold, then the current best candidate solution is taken as the final preferred solution and S8 is executed. Otherwise, it is determined whether the current iteration round is greater than the iteration threshold.

[0017] If the current iteration round is greater than the iteration threshold, the current optimal candidate solution is selected as the final preferred solution and S8 is executed; otherwise, based on the comprehensive reward value, remaining problems, improvement directions, and analysis report corresponding to the current optimal candidate solution, the multimodal large model selects enhancement operators from M' to form a new enhancement solution, and then uses the method described in S4 to apply the new enhancement solution to... Perform enhancement processing to obtain a new enhanced image, and proceed to the next step;

[0018] S7: Keep the current dimension weights unchanged, input the newly enhanced image, the remaining problems and improvement directions corresponding to the current best candidate solution into the multimodal large model, obtain the comprehensive reward value, remaining problems, improvement directions and analysis report corresponding to the newly enhanced image, and then use the new enhancement solution as the current best candidate solution and return to S6 while keeping the same reward value threshold and iteration threshold unchanged;

[0019] S8: Output the enhanced image and analysis report corresponding to the final optimal solution.

[0020] Preferably, in S1 Several key metrics include clarity Contrast ,brightness and color cast The calculation formula is as follows:

[0021]

[0022] Wherein, 𝑀 represents The length, 𝑁 represents width, express grayscale image pixels, Represents the Laplace operator. Representing an image The pixel mean after processing with the Laplacian operator, expressed as follows:

[0023]

[0024]

[0025] in, Represents a grayscale image The pixel mean, expressed as follows:

[0026]

[0027]

[0028]

[0029] in, ,and This represents the average pixel value of the red, green, and blue channels in a grayscale image.

[0030] As a preferred option, the middle Clarity Contrast ,brightness and color cast The corresponding severity calculation formulas are as follows:

[0031]

[0032]

[0033]

[0034]

[0035]

[0036] in, This indicates the severity level corresponding to the clarity. This indicates the severity corresponding to the contrast. This indicates the severity corresponding to the brightness. This indicates the severity of the color cast. This represents a piecewise linear normalized function. This represents the key indicator value to be normalized. This indicates the lower threshold corresponding to the indicator. This indicates the upper threshold corresponding to the indicator.

[0037] Preferably, in step S5, defining the weights of the evaluation dimensions also includes the following formula for calculating the weights of each evaluation dimension when an image simultaneously exhibits multiple major problem types:

[0038]

[0039] in, Indicates the first Weights of each dimension The dimension representing the weight. Represents the set of quality dimensions Any dimension index variable in, , indicating the i-th detected problem, with a total of K problems. This represents the i-th question type. Indicate the problem Dimension Influence coefficient, Indicates the severity of the corresponding problem. A set representing quality dimensions.

[0040] In step S5, when the dimension weights are sharpness, contrast, and color, the formula for calculating the comprehensive reward value of the candidate enhancement scheme is as follows:

[0041]

[0042] in, Indicates the first The combined reward value of each candidate enhancement scheme, This represents the weighting coefficients that are dynamically adjusted by the multimodal large model according to the problem type; These represent the evaluation scores of the multimodal large model for the candidate enhancement results in terms of sharpness, contrast, and color, respectively. represents the artifact side effect term; πœ† represents the penalty coefficient.

[0043] Compared with the prior art, the present invention has at least the following advantages:

[0044] First, this invention does not rely on clear reference images or paired training data. It can complete the augmentation decision through the closed-loop collaboration of traditional unsupervised augmentation operators and multimodal large models, which significantly reduces the cost of data acquisition and model training.

[0045] Second, this invention does not use a single enhancement algorithm, but adaptively selects a combination of operators based on the turbidity, blurriness, low contrast, and color cast of the current input image. Therefore, it can simultaneously adapt to underwater images under different water conditions, depths, and lighting conditions.

[0046] Third, the present invention enables the enhancement process to have automatic correction capabilities through a closed-loop iterative optimization mechanism consisting of candidate scheme generation, enhancement execution, multimodal evaluation, problem feedback, and regeneration, which can avoid the under-enhancement and over-enhancement problems commonly found in traditional fixed pipelines.

[0047] Fourth, this invention introduces dynamic weights and side effect tolerance principles in the evaluation stage, making the scoring objectives closer to the actual task requirements, that is, focusing on improving the visibility and recognizability of the objectives, rather than just pursuing a single visual indicator.

[0048] Fifth, the present invention outputs problem diagnosis results, weight allocation results, candidate ranking results, and optimal algorithm combination, which have strong interpretability and engineering reproducibility, making it easy to apply directly in underwater robots, online monitoring, and edge deployment systems. Attached Figure Description

[0049] Figure 1 This is a schematic diagram of the overall framework of the adaptive underwater image enhancement method based on multimodal evaluation of the present invention.

[0050] Figure 2 This is a schematic diagram of the evaluation steps in a specific implementation of the method of the present invention.

[0051] Figure 3 This is a schematic diagram of the LLM intelligent evaluation process in this method.

[0052] Figure 4 This is a schematic diagram comparing the performance metrics of the present invention with those of traditional unsupervised enhancement methods.

[0053] Figure 5 The following are comparison images of the enhancement effects of the method of the present invention and various traditional unsupervised enhancement methods in a laboratory water tank scenario: (a) original image, (b) AncutiFusion, (c) IBLA, (d) Rank-OnePrior, (e) Water-Net, (f) UVE-Net, (g) UIEC2Net, (h) CCL-Net, (i) Spectroformer, (j) Phaseformer, (k) AutomatedMSRCR, (l) the method of the present invention.

[0054] Figure 6 Comparison of the enhancement effects of the method of this invention with various traditional unsupervised enhancement methods in river / port scenes: (a) original image, (b) AncutiFusion, (c) IBLA, (d) Rank-OnePrior, (e) Water-Net, (f) UVE-Net, (g) UIEC2Net, (h) CCL-Net, (i) Spectroformer, (j) Phaseformer, (k) AutomatedMSRCR, (l) the method of this invention;

[0055] Figure 7 Comparison of the enhancement effects of the method of this invention and various traditional unsupervised enhancement methods in deep-sea scenarios: (a) Original image, (b) AncutiFusion, (c) IBLA, (d) Rank-OnePrior, (e) Water-Net, (f) UVE-Net, (g) UIEC2Net, (h) CCL-Net, (i) Spectroformer, (j) Phaseformer, (k) AutomatedMSRCR, (l) The method of this invention;

[0056] Figure 8 The images shown are screenshots of the visualization system interface of this invention and comparison diagrams before and after enhancement.

[0057] Figure 9 This is the original image used in the embodiments of the present invention.

[0058] Figure 10This is an image after enhancement of the original image in an embodiment of the present invention. Detailed Implementation

[0059] The present invention will now be described in further detail.

[0060] This invention discloses an adaptive underwater image enhancement method based on multimodal evaluation. This method combines the stability of traditional unsupervised enhancement operators with the intelligent decision-making ability of multimodal large models. Through the dual mechanism of rapid measurement by computer vision rules and accurate judgment by multimodal large models, the large model simultaneously undertakes the roles of generating enhancement schemes and evaluating enhancement effects. After multiple rounds of closed-loop optimization, the optimal combination of enhancement algorithms is output.

[0061] See Figures 1-10 An adaptive underwater image enhancement method based on multimodal evaluation includes the following steps:

[0062] S1: Acquire underwater images or video frames to be enhanced. and calculate Several key indicators are used, and each key indicator is mapped to the [0,1] interval using an adaptive normalization algorithm to obtain a normalized value for each key indicator. The normalized value of each key indicator is then used as its severity, and the severity of each key indicator corresponds to a problem type. For example, when the key indicators to be calculated are sharpness, contrast, brightness, and color cast, the severity of these four indicators corresponds to four main problem types: blurriness, low contrast, excessive darkness, and color cast. These main problem types are then grouped into... ,Right now ;

[0063] In S1 Several key metrics include clarity Contrast ,brightness and color cast The Laplacian variance of the grayscale image is used to characterize sharpness, the standard deviation of grayscale is used to characterize contrast, the mean of grayscale is used to characterize brightness, and the standard deviation of the means of the three RGB channels is used to characterize the degree of color cast. The calculation formulas are as follows:

[0064]

[0065] Wherein, 𝑀 represents The length, 𝑁 represents width, express grayscale image pixels, Represents the Laplace operator. Representing an image The pixel mean after processing with the Laplacian operator, expressed as follows:

[0066]

[0067]

[0068] in, Represents a grayscale image The pixel mean, expressed as follows:

[0069]

[0070]

[0071]

[0072] in, ,and This represents the pixel mean of the red, green, and blue channels in a grayscale image. The statistics of the above key indicators can be quickly calculated without relying on a reference ground truth image, thus eliminating the influence of some external factors.

[0073] The middle Clarity Contrast ,brightness and color cast The corresponding severity calculation formulas are as follows:

[0074]

[0075]

[0076]

[0077]

[0078]

[0079] in, This indicates the severity level corresponding to the clarity. This indicates the severity corresponding to the contrast. This indicates the severity corresponding to the brightness. This indicates the severity of the color cast. This represents a piecewise linear normalization function used to map input metrics to severity values ​​in the [0,1] interval. This represents the key indicator value to be normalized. This indicates the lower threshold corresponding to the indicator. This indicates the upper threshold corresponding to the indicator.

[0080] S2: Select the problem type corresponding to the key indicator with the highest severity as... The main problem types are identified; assessments are dynamically generated based on the main problem type and the severity of each problem. If the main problem is blurriness, the weight of sharpness is increased; if the main problem is low contrast, the weight of contrast is increased; if the main problem is color cast, the weight of color naturalness is increased.

[0081] S3: Construct an unsupervised operator library and a candidate pipeline library. The unsupervised enhancement operator library includes M' enhancement operators, some of which are existing technologies. Table 1 below contains 18 commonly used unsupervised enhancement operators. The candidate pipeline library includes N' candidate enhancement schemes, and each candidate enhancement scheme is composed of one or more operators connected in sequence. Each candidate enhancement scheme is defined as a candidate pipeline.

[0082] Table 1 Unsupervised Augmentation Operator Library

[0083]

[0084] A set of candidate enhancement schemes is generated from a pre-built library of unsupervised enhancement operators. The enhancement scheme can be a single operator or a pipeline formed by combining multiple operators in sequence. The combination generally includes at least one or more of the following: dehazing, deblurring, color cast correction, local contrast enhancement, white balance, channel compensation, and brightness adjustment.

[0085] S4: Utilize N' candidate enhancement schemes separately and in parallel... Enhancement processing is performed to obtain the enhanced result image corresponding to each candidate enhancement scheme, totaling N' enhanced result images. The candidate schemes are executed in parallel to shorten the multi-scheme polling time. The enhancement results are associated with and saved with the corresponding operator name, parameters, and output path. The parallel execution module can encode the input image into a unified format and distribute it to multiple worker threads to execute different candidate enhancement schemes respectively. The enhancement results are then aggregated to improve polling efficiency. The enhancement processing expression is as follows:

[0086]

[0087] in, This represents the enhanced image corresponding to the k-th candidate enhancement scheme. This represents the i-th enhancement operator among the k-th candidate schemes. This indicates that the total number of candidate solutions in the k-th candidate solution is... An enhancement operator;

[0088] S5: Define the evaluation dimensions and their weights based on the main problem types. The weights for each dimension include sharpness weight, contrast weight, and color weight. This corresponds to the case where an image has only one main problem. The corresponding N' enhanced result images are input into a multimodal large model. The multimodal large model adopts existing technologies such as Qwen3-VL-Plus. The multimodal large model analyzes each enhanced result image according to the dimensional weights to obtain the comprehensive reward value, remaining problems, improvement directions and analysis report for each enhanced result image. All enhanced result images are sorted in descending order according to the comprehensive reward value, and the candidate enhancement scheme with the highest comprehensive reward value is selected as the current optimal candidate scheme. The multimodal large model evaluation adopts the evaluation principle that side effects are only deducted when they actually affect target recognition. This can avoid excessive punishment of enhancement results that significantly improve target visibility due to slight noise, slight color cast or slight artifacts.

[0089] according to The evaluation weights for each dimension are dynamically allocated, and the specific weight allocation strategy is shown in Table 2:

[0090] Table 2. Correspondence between Problem Types and Dynamic Evaluation Weights

[0091]

[0092] In step S5, defining the weights of the evaluation dimensions also includes the following formula for calculating the weights of each evaluation dimension when an image simultaneously exhibits multiple major problem types:

[0093]

[0094] in, Indicates the first Weights of each dimension Dimensions representing weights, such as sharpness, contrast, and color. Represents the set of quality dimensions Any dimension index variable in the table is used for normalized summation. , indicating the i-th detected problem, with a total of K problems. This represents the i-th question type. Indicate the problem Dimension Influence coefficient, Indicates the severity of the corresponding problem. This represents a set of quality dimensions. When multiple main problem types appear in an image, it is necessary to comprehensively consider the adjustment weight of each evaluation dimension. At this time, it is necessary to calculate the weight value of each dimension in a balanced manner. This formula achieves adaptive weight allocation for multi-problem multi-dimensional quality evaluation by weighting and accumulating the mapping relationship between problem severity and quality dimensions and normalizing it globally.

[0095] In step S5, when the dimension weights are sharpness, contrast, and color, the formula for calculating the comprehensive reward value of the candidate enhancement scheme is as follows:

[0096]

[0097] in, Indicates the first The combined reward value of each candidate enhancement scheme, This indicates that the multimodal large model is determined according to the problem type. Dynamically adjusted weighting coefficients; These represent the evaluation scores of the multimodal large model for the candidate enhancement results in terms of sharpness, contrast, and color, respectively. The term "artifact side effect" is obtained by comparing and evaluating the input image and the corresponding enhancement result using a multimodal large model. This term takes a non-zero value when the model determines that there are artifacts in the enhancement result that affect the stability of the target detection or recognition result; otherwise, it takes zero. "Punishment coefficient" is used to adjust the degree of influence of side effects on the overall score. The comprehensive reward value is composed of the weighted average of the indicators corresponding to the various dimensions, plus the side effect penalty. In this method, three dimensions are used for calculation; in actual use, the calculation of the comprehensive reward value can be increased or decreased as needed.

[0098] S6: Preset reward value threshold and iteration threshold. If the reward value of the current best candidate solution is greater than or equal to the reward value threshold, then the current best candidate solution is taken as the final preferred solution and S8 is executed. Otherwise, it is determined whether the current iteration round is greater than the iteration threshold.

[0099] If the current iteration round is greater than the iteration threshold, the current optimal candidate solution is selected as the final preferred solution and S8 is executed; otherwise, based on the comprehensive reward value, remaining problems, improvement directions, and analysis report corresponding to the current optimal candidate solution, the multimodal large model selects enhancement operators from M' to form a new enhancement solution, and then uses the method described in S4 to apply the new enhancement solution to... The enhancement process is performed to obtain a new enhanced image, and then the next step is executed. The iteration here means selecting a set of enhancement schemes from M'. If the selected set of new enhancement schemes does not meet the requirements, then random combination is performed again. The iteration threshold represents the total number of times that enhancement schemes can be randomly combined from M'.

[0100] S7: Keeping the current dimension weights unchanged, input the newly enhanced image, the remaining problems corresponding to the current best candidate solution, and the improvement directions into the multimodal large model to obtain the comprehensive reward value, remaining problems, improvement directions, and analysis report corresponding to the newly enhanced image. Then, use the new enhancement solution as the current best candidate solution and return to S6 while keeping the same reward value threshold and iteration threshold unchanged. After returning to S6, the comparison is with the new enhancement solution, that is, the new enhancement solution obtained in this round is used as the new current best candidate solution for comparison.

[0101] S8: Output the final optimal solution, i.e., the enhanced image and analysis report corresponding to the strongest operator combination. The analysis report includes parameter configuration, problem diagnosis results, ranking report, and reward report, etc., for online deployment, result traceability, and subsequent experience reuse.

[0102] Example:

[0103] Taking an underwater image acquired in a river / port setting as an example, the specific implementation process of the method of this invention is explained. This input image is affected by water scattering, suspended particle occlusion, and non-uniform lighting, resulting in obvious turbidity, blurring, and low contrast. The rock outlines in the target area are unclear, and texture details are submerged, making it a typical degraded image from complex water areas. See [link / reference]. Figure 9 .

[0104] First, the underwater images collected in the river / port scene were taken as the original images to be enhanced as input, and the problem analysis of the input original images to be enhanced was performed. Through image quality analysis, the main degradation problem of the image was identified as blurring.

[0105] Based on the identified problem typeβ€”ambiguityβ€”the dynamic weight values ​​for the dimensions of this evaluation are determined, with sharpness, contrast, and color naturalness weighted at 0.75, 0.25, and 0.0, respectively. Therefore, in this embodiment, improving sharpness is the primary evaluation objective, and improving contrast is the secondary evaluation objective.

[0106] Multiple candidate enhancement schemes are read from a pre-built library of unsupervised enhancement operators, and each candidate scheme is executed in parallel in a single round to obtain the corresponding candidate enhancement result image. In practice, based on the enhancement effect of each round, several enhancement schemes that have been tested will be formed. These enhancement schemes can be used directly without having to go through the step of repeatedly extracting and forming schemes from enhancement operators. In this embodiment, the candidate enhancement schemes (and the enhancement operators contained in each scheme) participating in the evaluation include:

[0107] 1) SimplestColorBalance

[0108] 2) FastMSRCR

[0109] 3) Automated MSRCR

[0110] 4)CompensateChannels->WhiteBalance->CLAHE

[0111] 5)FastMSRCR->SimplestColorBalance

[0112] 6)Defogging->FastMSRCR->SimplestColorBalance

[0113] 7)FastMSRCR->DCP->CompensateChannels->WhiteBalance

[0114] 8) RankOnePlus

[0115] After the candidate enhancement schemes are executed, the original image and all candidate enhancement result images are input into the Qwen3-VL-Plus multimodal large model evaluation module. The multimodal large model then performs a unified ranking of all candidate results. The multimodal large model uses the evaluation principle of "prioritizing primary targets and tolerating side effects" to comprehensively judge each candidate result in terms of target sharpness, contrast, and the impact of side effects, and outputs the corresponding scores and ranking results.

[0116] In this embodiment, the unified evaluation results of the multimodal large model show that:

[0117] 1. The solution Defogging->FastMSRCR->SimplestColorBalance has the highest reward value, at 9.45;

[0118] 2. The reward value for the FastMSRCR scheme is 8.92;

[0119] 3. The reward value for the FastMSRCR->DCP->CompensateChannels->WhiteBalance scheme is 8.37;

[0120] 4. The reward value for the RankOnePlus scheme is 7.84.

[0121] Therefore, the Defogging->FastMSRCR->SimplestColorBalance scheme has a reward value of 9.45, which is greater than the threshold of 9.0. The system will determine it as the optimal enhancement algorithm combination for this input image.

[0122] The specific processing steps for the selected optimal enhancement algorithm are as follows: First, the Defogging operator is used to dehaze and de-turbidify the original image to reduce the fog effect caused by scattering and particle occlusion, thereby improving the separability between the target and the background. Second, the FastMSRCR operator is used to perform multi-scale Retinex enhancement on the dehazed image to improve local contrast, strengthen edge information, and restore target texture details. Finally, the SimplestColorBalance operator is used to perform color channel clipping and normalization on the enhancement result to suppress extreme brightness distributions and improve overall grayscale levels and visual stability.

[0123] After the above processing, the output image is shown below. Figure 10 The main outline of the rock is clear, the boundaries are sharp, and the surface texture details are significantly restored. The separation effect between the target area and the background is significantly better than that of the original image. Although there is a slight purplish-blue tint and a small amount of noise in the output results, these side effects do not substantially interfere with target recognition, and since color naturalness is not the primary evaluation target in this embodiment, they do not affect the fact that the proposed scheme is determined to be the optimal enhancement scheme.

[0124] Therefore, in complex water environments such as rivers and ports, the method of this invention can automatically select the optimal operator combination from multiple unsupervised enhancement candidate schemes to address the specific degradation problem of the input image, thereby achieving adaptive enhancement of complex underwater images. It has good engineering applicability and practical application value.

[0125] Experimental content and results

[0126] The underwater images processed by this invention cover a variety of typical scenarios, including artificially constructed turbid environments, naturally turbid water bodies, and deep-sea artificial light source imaging. Specifically:

[0127] The first category consists of laboratory water tank scenes. These images were acquired in laboratory water tanks, representing artificially constructed controlled turbidity imaging scenarios. By adjusting the concentration of milky particles in the water, underwater scattering effects are simulated, primarily resulting in a moderate decrease in contrast and a degradation in sharpness. These images can be used to verify the enhancement performance of the algorithm under known degradation conditions. 185 images of this type were collected.

[0128] The second category is river / port scenes. These images were collected in natural aquatic environments, such as rivers, ports, or nearshore areas. The water contains silt, organic matter, and other suspended particles, and the ambient lighting is complex and unstable, with a relatively short visibility distance. These images belong to natural, highly turbid water imaging scenes, characterized by strong scattering, low visibility, and non-uniform degradation. They represent a complex and challenging scenario for underwater image enhancement, placing high demands on the algorithm's descattering and detail recovery capabilities. 288 images from this category were collected.

[0129] The third category is deep-sea scenes. These images are acquired in deep-sea environments, typically by underwater robots (such as AUVs or ROVs) equipped with gimbal systems. Since natural light cannot reach this depth, imaging relies entirely on artificial lighting. Due to the absorption and scattering of light by the water, the images exhibit a noticeable blue-green color cast, accompanied by reduced contrast and overall fogging, blurring details and indistinct boundaries of distant targets. Furthermore, the reflection and scattering of light by numerous suspended particles in the water results in significant particle noise and stray bright spots in the images. Additionally, due to the limited range of artificial light sources, the image brightness exhibits an uneven illumination characteristic, with a brighter center and gradually diminishing brightness at the edges. 224 images of this type were collected.

[0130] The aforementioned scenarios correspond to different degradation mechanisms, including turbidity degradation dominated by scattering, extremely low visibility degradation characterized by multiple scattering, and deep-sea imaging degradation dominated by uneven light absorption and illumination. These different degradation mechanisms manifest in images as various combinations of problems such as decreased contrast, reduced sharpness, and color distortion. This invention achieves adaptive processing of multiple degradation types through dynamic weight adjustment and candidate enhancement pipeline optimization mechanisms.

[0131] For each scene image, the optimal enhancement combination selected by the present invention is used. To verify the effectiveness of the present invention's method, representative underwater image enhancement methods are selected for comparison, including methods based on Retinex theory, image fusion methods, physical prior-based restoration methods, and deep learning-based enhancement methods. Among these, the deep learning methods cover convolutional neural network models and state-of-the-art Transformer-based models. Table 3 presents the average unreferenced metrics UIQM and UCIQE for the original image, the SOAT algorithm enhancement results, and the processing results of the optimal enhancement scheme of the present invention on the three scene image types. UIQM mainly reflects the image's sharpness, contrast, and color restoration capability, while UCIQE mainly measures the image's overall color quality and contrast level; a higher value indicates better image quality.

[0132]

[0133] in , , These represent the evaluation items for color quality, sharpness, and contrast, respectively. , , These are weighting coefficients, typically taking the following values: , , .

[0134] ,in, , , The mean, Standard deviation;

[0135] The image was divided into K small blocks. Let x be the weight coefficient of the x-th image patch. , This indicates the maximum pixel value in the block. This represents the minimum pixel value in the block.

[0136] ,in , This is used to avoid numerical instability issues when the denominator is zero or during logarithmic calculations.

[0137]

[0138] in Indicates the standard deviation of colorimetry. Indicates brightness contrast. This represents the mean saturation. , , These are weighting coefficients, and their values ​​are... , , .

[0139] Table 3 Comparison of evaluation results of our method and other methods in different aquatic scenarios

[0140]

[0141] As shown in Table 3, the method of this invention achieved superior objective evaluation results in different aquatic scenarios. Regarding the UIQM index, the method of this invention achieved the highest values ​​in laboratory pools, rivers / ports, and deep-sea scenarios (3.1122, 2.7437, and 3.1122, respectively), demonstrating stable overall performance and a leading advantage, indicating its good effect in image sharpness and structural information recovery.

[0142] Regarding the UCIQE index, the method of this invention achieves high levels (0.2585 and 0.2100) in laboratory pool and deep-sea scenarios, outperforming most comparison methods overall and demonstrating good comprehensive color restoration capabilities. Although some methods achieve high values ​​for a single index in specific scenarios, they often suffer from problems such as excessive contrast enhancement or color distortion.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An adaptive underwater image enhancement method based on multimodal evaluation, characterized in that: Includes the following steps: S1: Acquire underwater images to be enhanced and calculate Several key indicators are used to map each key indicator to the [0,1] interval through an adaptive normalization algorithm to obtain the normalized value of each key indicator. The normalized value of each key indicator is used as the severity of each key indicator, and the severity of each key indicator corresponds to a problem type. S2: Select the problem type corresponding to the key indicator with the highest severity as... The main types of problems; S3: Construct an unsupervised operator library and a candidate pipeline library. The unsupervised enhancement operator library includes M' enhancement operators, and the candidate pipeline library includes N' candidate enhancement schemes. Each candidate enhancement scheme is composed of one or more operators connected in sequence. Each candidate enhancement scheme is defined as a candidate pipeline. S4: Utilize N' candidate enhancement schemes separately and in parallel... Enhancement processing is performed to obtain the enhanced result image corresponding to each candidate enhancement scheme, resulting in a total of N' enhanced result images. The enhancement processing expression is as follows: in, This represents the enhanced image corresponding to the k-th candidate enhancement scheme. This represents the i-th enhancement operator among the k-th candidate schemes. This indicates that the total number of candidate solutions in the k-th candidate solution is... An enhancement operator; S5: Define the evaluation dimensions and their weights based on the main problem types. The corresponding N' augmentation result images are input into the multimodal large model. The multimodal large model analyzes each augmentation result image according to the dimension weight to obtain the comprehensive reward value, remaining problems, improvement directions and analysis report for each augmentation result image. All augmentation result images are sorted in descending order according to the comprehensive reward value, and the candidate augmentation scheme corresponding to the highest comprehensive reward value is taken as the current optimal candidate scheme. S6: Preset reward value threshold and iteration threshold. If the reward value of the current best candidate solution is greater than or equal to the reward value threshold, then the current best candidate solution is taken as the final preferred solution and S8 is executed. Otherwise, it is determined whether the current iteration round is greater than the iteration threshold. If the current iteration round is greater than the iteration threshold, the current optimal candidate solution is selected as the final preferred solution and S8 is executed; otherwise, based on the comprehensive reward value, remaining problems, improvement directions, and analysis report corresponding to the current optimal candidate solution, the multimodal large model selects enhancement operators from M' to form a new enhancement solution, and then uses the method described in S4 to apply the new enhancement solution to... Perform enhancement processing to obtain a new enhanced image, and proceed to the next step; S7: Keep the current dimension weights unchanged, input the newly enhanced image, the remaining problems and improvement directions corresponding to the current best candidate solution into the multimodal large model, obtain the comprehensive reward value, remaining problems, improvement directions and analysis report corresponding to the newly enhanced image, and then use the new enhancement solution as the current best candidate solution and return to S6 while keeping the same reward value threshold and iteration threshold unchanged; S8: Output the enhanced image and analysis report corresponding to the final optimal solution.

2. The adaptive underwater image enhancement method based on multimodal evaluation as described in claim 1, characterized in that: In S1 Several key metrics include clarity Contrast ,brightness and color cast The calculation formula is as follows: Wherein, 𝑀 represents The length, 𝑁 represents width, express grayscale image pixels, Represents the Laplace operator. Representing an image The pixel mean after processing with the Laplacian operator, expressed as follows: in, Represents a grayscale image The pixel mean, expressed as follows: in, ,and This represents the average pixel value of the red, green, and blue channels in a grayscale image.

3. The adaptive underwater image enhancement method based on multimodal evaluation as described in claim 2, characterized in that: The middle Clarity Contrast ,brightness and color cast The corresponding severity calculation formulas are as follows: in, This indicates the severity level corresponding to the clarity. This indicates the severity corresponding to the contrast. This indicates the severity corresponding to the brightness. This indicates the severity of the color cast. Represents a piecewise linear normalized function. This represents the key indicator value to be normalized. This indicates the lower threshold corresponding to the indicator. This indicates the upper threshold corresponding to the indicator.

4. The adaptive underwater image enhancement method based on multimodal evaluation as described in claim 3, characterized in that: In step S5, defining the weights of the evaluation dimensions also includes the following formula for calculating the weights of each evaluation dimension when an image simultaneously exhibits multiple major problem types: in, Indicates the first Weights of each dimension The dimension representing the weight. Represents the set of quality dimensions Any dimension index variable in, , indicating the i-th detected problem, with a total of K problems. This represents the i-th question type. Indicate the problem Dimension Influence coefficient, Indicates the severity of the corresponding problem. A set representing quality dimensions.

5. The adaptive underwater image enhancement method based on multimodal evaluation as described in claim 4, characterized in that: In step S5, when the dimension weights are sharpness, contrast, and color, the formula for calculating the comprehensive reward value of the candidate enhancement scheme is as follows: in, Indicates the first The combined reward value of each candidate enhancement scheme, This represents the weighting coefficients that are dynamically adjusted by the multimodal large model according to the problem type; These represent the evaluation scores of the multimodal large model for the candidate enhancement results in terms of sharpness, contrast, and color, respectively. Indicates the side effects of artifacts; This represents the penalty coefficient.