Underwater image quality evaluation model based on multi-modal contrast learning and semantic guidance

Through multimodal contrast learning and semantic guidance underwater image quality evaluation model, the problem of relying on manual annotation in the existing technology is solved, unsupervised precise quality evaluation and cross-domain generalization capabilities are achieved, and it is suitable for underwater image quality evaluation.

CN120298315APending Publication Date: 2025-07-11DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510297837.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing underwater image quality evaluation methods rely on a large number of manual annotations, making it difficult to accurately quantify the higher-order visual properties of underwater images, especially in extreme turbid or deep-sea scenes, and the enhanced results may introduce artifacts.

Method used

Unsupervised learning multimodal contrast learning and semantic guidance methods are adopted, and the alignment strategy of quality comparison learning and visual language is combined with feature hierarchical comparison of embedded semantic prompt information to achieve accurate docking and sorting of images and quality texts.

Benefits of technology

It realizes accurate quality evaluation of underwater images under unsupervised conditions, improves cross-domain generalization capabilities and evaluation accuracy, and can make quality predictions in zero samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298315A_ABST
    Figure CN120298315A_ABST
Patent Text Reader

Abstract

According to the underwater image quality evaluation method based on multi-modal contrast learning and semantic guidance, an existing underwater image quality evaluation method generally depends on a large amount of manual annotations to train a target quality evaluation model, unsupervised learning is mainly adopted in the underwater image quality evaluation method, and the underwater image quality evaluation method aims at automatically capturing various features of an underwater image; and accurate quality representation is formed. Through quality contrast learning and a visual language alignment strategy and in combination with multi-sample perception similarity measurement of a mixing ratio, accurate docking between an image and a quality text is realized. In addition, the model further introduces feature hierarchical comparison of embedded semantic prompt information, images are grouped according to quality hierarchy by means of quality-related positive and negative text prompts, and the loss is compared through relatively ranked groups, so that efficient quality sorting is ensured. Finally, the statistical features of the image guide the unsupervised visual language model, and the unsupervised visual language model is combined with the visual language model, so that the quality prediction of the zero sample is successfully realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater image quality assessment. Specifically, it particularly relates to an underwater image quality assessment method based on multi-modal contrast learning and semantic guidance. Background Art

[0002] Approximately 71% of the Earth's surface is covered by the ocean, which contains huge potential in mineral resources, biological resources, energy, etc. However, limited by the complex underwater environment and technological level, the current exploration of the ocean by humans only covers about 5% of the sea area, and the actual development and utilization rate is less than 1%. In key fields such as marine resource exploration, environmental monitoring, underwater navigation, and security, high-quality underwater images are the core data sources supporting scientific research and engineering applications. However, the underwater imaging environment poses inherent challenges. The selective absorption of light by water bodies and the scattering of suspended particles result in serious color distortion and fogging effects in images. The limitations of artificial light sources cause local overexposure or shadow occlusion, and dynamic light changes further exacerbate the degradation of imaging quality. In different depth, turbidity, and biological activity scenarios, degradation modes such as blurring, low contrast, and noise exhibit non-linear coupling characteristics. To address the above problems, underwater image enhancement technology has significantly improved image visibility through physical model-driven and data-driven methods. However, with the rapid development of UIE technology, enhancement results may introduce over-enhancement artifacts or oversaturated areas. Existing metrics are mostly based on statistical features and are difficult to quantify high-order visual attributes such as structural integrity and semantic consistency. In extremely turbid or deep-sea scenarios, the correlation between metric scores and human subjective perception significantly decreases. Therefore, constructing an assessment method that can accurately quantify the true quality of underwater images and predict their performance in supporting downstream tasks has become an urgent need in the field of marine intelligent perception. Summary of the Invention

[0003] To address the above-mentioned technical problems, an underwater image quality assessment model based on multi-modal contrast learning and semantic guidance is proposed. Existing underwater image quality assessment methods usually rely on a large number of manual annotations to train the objective quality assessment (IQA) model. In contrast, the present invention mainly uses unsupervised learning, aiming to automatically capture various features of underwater images and form an accurate quality representation. Through quality contrast learning and visual language alignment strategies, combined with a multi-sample perceptual similarity metric with a mixing ratio, we achieve precise docking between images and quality texts. In addition, the model also introduces feature hierarchical contrast with embedded semantic hint information. With positive and negative text hints related to quality, images are grouped according to quality levels (high, medium, low), and through the group contrast loss of relative ranking, efficient quality ranking is ensured. Finally, the statistical features of the images guide the unsupervised visual language model, which, combined with the visual language model, successfully realizes zero-shot quality prediction.

[0004] The technical means adopted by the present invention are as follows:

[0005] Underwater image quality assessment method based on multi-modal contrast learning and semantic guidance, comprising the following steps:

[0006] Step 1: Obtain multiple low-quality images from multiple underwater image datasets, and for each low-quality image, obtain a high-quality reference image after image enhancement processing; the high-quality reference image is an image defined according to objective evaluation indicators or subjective scores;

[0007] Step 2: According to the low-quality images and high-quality reference images obtained in Step 1, linearly mix each pair of low-quality images and high-quality reference images with a mixing ratio of [0, 0.5, 1.0] to generate linearly mixed images; the linearly mixed images include multiple quality levels; where 0 represents a completely low-quality image, 1.0 represents a completely high-quality image, and 0.5 represents an equal mixture of a low-quality image and a high-quality reference image;

[0008] Step 3: Through quality contrast learning and visual language alignment strategy, embed the linearly mixed images generated in Step 2 into quality semantic labels to achieve the docking of image features and text information;

[0009] Step 4: Process image information through an image encoder; in the image encoder, use a multi-view weighted method to extract the features of low-quality images and their corresponding high-quality reference images, and measure the perceptual similarity of images with different qualities through a quality contrast loss function;

[0010] The extracted image features include: low-level texture, edge features, and high-level semantic information;

[0011] Step 5: Process text information through a text encoder; in the text encoder, according to the image perception results generated by the image encoder, assign text prompts to each image; group according to the quality levels of the images, and achieve accurate quality ranking through a group contrast loss function of relative ranking;

[0012] Step 6: Extract the texture features, edge features, and semantic features of the high-quality reference images, and conduct a comparative analysis of the texture features, edge features, and semantic features with the same type of features extracted from the low-quality images in Step 4; by calculating the distance between the features of the high-quality reference image and the low-quality image, quantify the gap between the low-quality image and the high-quality reference image, and perceive the image quality level;

[0013] Step 7: Guide the unsupervised vision-language model CLIP for adaptive training based on the statistical features extracted from low-quality images; the unsupervised vision-language model learns the hierarchical structure of underwater image quality through the fusion of statistical features extracted from high-quality images and vision-language information, and uses objective metrics for verification to ensure the accuracy of quality assessment. Finally, the trained model is used to generate corresponding quality assessment scores, and the image quality is judged according to the scores.

[0014] Further, in Step 2, an image dataset with different quality levels is generated. By constructing training samples with continuous quality changes, the calculation formula is:

[0015] I mix = αI low +(1 - α)I ref ;

[0016] where, I mix represents the image obtained by linear mixing, I low represents the low-quality image, and I ref represents the high-quality reference image. When α is 0, the features of the high-quality reference image are completely retained. When α is 0.5, the features of medium-degraded images are simulated. When α is 1.0, the features of the original low-quality image are retained.

[0017] Further, the joint alignment of image quality and text in Step 3 includes the following steps:

[0018] Step 31: Construct semantic descriptions corresponding to quality texts, and the semantic information is expressed as Perfect, Fair, and Bad.

[0019] Step 32: Initialize the dual encoder using the contrastive language-image pre-training CLIP pre-training model. By optimizing the objective function, the model can align the features of images and texts with predefined semantic labels, so as to achieve the effective fusion of image and text information and embed the corresponding semantic information. The calculation formula is:

[0020]

[0021] where, sin represents calculating the similarity of two vectors, E img represents the feature vector output by the image encoder, E text represents the feature vector output by the text encoder, T i and T j represent text descriptions of different qualities respectively, and L align represents the alignment degree between image and text features.

[0022] Further, the multi-view feature extraction and comparison in Step 4 include the following steps:

[0023] Step 41: The image encoder adopts a feature fusion strategy of ResNet50 and a multi-scale attention module, and the calculation formula is:

[0024]

[0025] where F fused is the fused feature representation, represents the feature map of the k-th layer of ResNet-50, W k represents the learnable weight parameter vector, and Softmax represents normalizing the weights to a probability distribution;

[0026] Step 42: Perform dynamic triplet sampling within the same α group, select the corresponding anchor, positive sample, and negative sample, and the calculation formula of the loss function is:

[0027]

[0028] where is used to calculate the similarity difference between the anchor, positive sample, and negative sample through the contrastive learning method, d represents the cosine distance, F a , F p and F n represent the anchor, positive sample, and negative sample features respectively, α represents the boundary threshold, The loss controls the distance difference between the positive sample and the negative sample, ensuring that the distance between the negative sample and the anchor is greater than the distance gap between the positive sample and the anchor.

[0029] Furthermore, the hierarchical quality ranking in step 5 includes the following steps:

[0030] Define the ranking constraint between quality groups through the group contrast ranking loss, and the calculation formula is:

[0031]

[0032] where represents the ranking loss, N g represents the number of quality groups, and S(q) represents the predicted score value of the model for the quality group q.

[0033] Furthermore, step 6 includes the following steps:

[0034] Step 61: Extract the statistical feature vector from the high-quality reference image set and generate the feature space using K-means clustering;

[0035] Step 62: Calculate the Mahalanobis distance between the mixed image feature generated in step 61 and the nearest prototype feature, and the calculation formula is:

[0036]

[0037] Among them, d Mahalanobis represents the distance between the feature vector of the image to be evaluated and the mean value of the feature space, and F mix represents the feature vector of the image to be evaluated, and μ k represents the mean value of the feature space, represents the inverse matrix of the covariance matrix, which is used to eliminate feature correlation.

[0038] Furthermore, in the step 7, the CLIP is guided by statistical features for multimodal fusion evaluation, and the calculation formula is:

[0039]

[0040] Among them, Q final represents the final image quality evaluation score, Q semantic represents the original score output by the semantic model, Q stat represents the weighted score of statistical features, μ stat and σ stat respectively represent the mean value and variance of the statistical feature scores, β represents the fusion weight, and Sigmoid represents the normalization process.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] 1. To solve the influence of relying on a large number of manual annotations to train the target quality evaluation model. The present invention uses quality contrast learning and visual language alignment strategies, and adopts a multi-sample perceptual similarity metric with a mixed ratio to achieve the precise docking of images and quality texts, optimize the alignment between vision and language, and thus enhance the expressiveness and accuracy of quality evaluation.

[0043] 2. Feature hierarchical contrast with embedded semantic prompt information. With the help of positive and negative text prompts related to quality, we group images according to quality levels (high, medium, low), and use the group contrast loss of relative ranking to accurately capture the semantic quality information of images and achieve precise quality ranking.

[0044] 3. An unsupervised vision-language model guided by image statistical features. By combining the statistical features of images with the vision-language model and through unsupervised learning methods, it can achieve zero-shot quality prediction on multiple IQA datasets, significantly improving the cross-domain generalization ability.

[0045] For the above reasons, the present invention is applicable to underwater image quality evaluation, especially in scenarios where unsupervised zero-shot prediction scoring is required. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0047] Figure 1 It is the model flowchart of the present invention.

[0048] Figure 2 It is a comparison diagram of the effects of the present invention in multi-modal visual quality and text correspondence. Among them, (a), (b), and (c) are three groups of pictures. (a)-1, (b)-1, and (c)-1 all indicate that Bad is an underwater low-quality image. (a)-2, (b)-2, and (c)-2 all indicate that Fair is the corresponding intermediate-level underwater image. (a)-3, (b)-3, and (c)-3 all indicate that Perfect corresponds to a high-quality underwater reference image. It can be seen from left to right that there is an obvious enhancement in the effect.

[0049] Figure 3 It shows the performance of the present invention in multi-modal visual quality and text correspondence prediction scores. (a) represents the original underwater low-quality image, with obvious color deviation and low contrast; (b) represents the image after preliminary processing, with the color shift improved but still relatively blurred; (c) represents the further enhanced image, with the clarity improved; (d) and (e) represent continuing to optimize in terms of color and contrast to make the visual quality closer to the real underwater scene; (f) to (j) represent the enhancement results of different methods or different degrees. (j) represents the high-quality underwater reference image, with the best visual quality and natural color. From (a) to (j), the image gradually enhances from the underwater low-quality image to the high-quality underwater reference image, and during the enhancement process, the prediction score also continuously increases with the improvement of the image quality. Detailed implementation manners

[0050] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] As Figure 1 shown, the present invention provides an underwater image quality assessment method based on multi-modal contrast learning and semantic guidance, including the following steps:

[0053] Step 1: Obtain multiple low-quality images from multiple underwater image datasets, and for each low-quality image, obtain a high-quality reference image after image enhancement processing; the high-quality reference image is an image defined according to objective evaluation indicators or subjective scores;

[0054] Step 2: According to the low-quality images and high-quality reference images obtained in Step 1, linearly mix each pair of low-quality images and high-quality reference images with a mixing ratio [0, 0.5, 1.0] to generate linearly mixed images; the linearly mixed images include multiple quality levels; among them, 0 represents a completely low-quality image, 1.0 represents a completely high-quality image, and 0.5 represents an equal mixture of a low-quality image and a high-quality reference image.

[0055] As a preferred embodiment, in the present application, Step 2 generates an image dataset with different quality levels. By constructing training samples with continuous quality changes, the calculation formula is:

[0056] I mix = αI low +(1 - α)I ref ;

[0057] wherein, I mix represents the image obtained by linear mixing, I low represents the low-quality image, I ref represents the high-quality reference image. When α is 0, the features of the high-quality reference image are completely retained. When α is 0.5, the features of a medium-degraded image are simulated. When α is 1.0, the features of the original low-quality image are retained.

[0058] Step 3: Through quality contrast learning and visual-language alignment strategy, embed the linear mixed images generated in Step 2 into quality semantic labels to achieve the docking of image features and text information.

[0059] As a preferred implementation, the joint alignment of image quality and text in Step 3 of this application includes the following steps:

[0060] Step 31: Construct the corresponding semantic descriptions of quality texts, and the semantic information is expressed as Perfect, Fair, and Bad.

[0061] Step 32: Initialize the dual encoder using the Contrastive Language-Image Pretraining (CLIP) pre-trained model. By optimizing the objective function, the model can align the features of images and texts with predefined semantic labels, thereby achieving the effective fusion of image and text information and embedding the corresponding semantic information. The calculation formula is:

[0062]

[0063] where, sin represents calculating the similarity of two vectors, E img represents the feature vector output by the image encoder, E text represents the feature vector output by the text encoder, T i and T j respectively represent text descriptions of different qualities, and L align represents the alignment degree between image and text features.

[0064] Step 4: Process the image information through the image encoder; in the image encoder, adopt a multi-view weighted method to extract the features of low-quality images and their corresponding high-quality reference images, and measure the perceptual similarity of images of different qualities through the quality contrast loss function; the extracted image features include: low-level texture, edge features, and high-level semantic information.

[0065] Preferably, the multi-view feature extraction and contrast in Step 4 include the following steps:

[0066] Step 41: The image encoder adopts the feature fusion strategy of ResNet50 and the multi-scale attention module. The calculation formula is:

[0067]

[0068] where, F fused is the fused feature representation, represents the k-th layer feature map of ResNet-50, W k represents the learnable weight parameter vector, and Softmax represents normalizing the weights to a probability distribution;

[0069] Step 42: Dynamically sample triples within the same α group, select the corresponding anchor, positive sample, and negative sample, and the calculation formula of the loss function is:

[0070]

[0071] Among them, is used to calculate the similarity difference between the anchor, positive sample, and negative sample through the contrastive learning method, d represents the cosine distance, and F a 、F p and F n represent the anchor, positive sample, and negative sample features respectively, α represents the boundary threshold, The loss controls the distance difference between the positive sample and the negative sample, ensuring that the distance between the negative sample and the anchor is greater than the distance gap between the positive sample and the anchor.

[0072] Step 5: Process the text information through a text encoder; in the text encoder, according to the image perception result generated by the image encoder, assign a text prompt to each image; group according to the quality level of the image, and achieve accurate quality ranking through the group contrast loss function of relative ranking.

[0073] As a preferred implementation manner, the hierarchical quality ranking in Step 5 includes the following steps:

[0074] Define the sorting constraint between quality groups through the group contrast sorting loss, and the calculation formula is:

[0075]

[0076] Among them, represents the sorting loss, N g represents the number of quality groups, and S(q) represents the predicted score value of the model for the quality group q.

[0077] Step 6: Extract the texture features, edge features, and semantic features of the high-quality reference image, and conduct a comparative analysis of the texture features, edge features, and semantic features with the same type of features extracted from the low-quality image in Step 4; by calculating the distance between the high-quality reference image and the low-quality image features, quantify the gap between the low-quality image and the high-quality reference image, and perceive the image quality level;

[0078] As a preferred implementation manner, Step 6 includes the following steps:

[0079] Step 61: Extract the statistical feature vector from the high-quality reference image set and generate a feature space using K-means clustering;

[0080] Step 62: Calculate the Mahalanobis distance between the mixed image features generated in Step 61 and the nearest prototype feature, and the calculation formula is:

[0081]

[0082] where d Mahalanobis represents the distance between the feature vector of the image to be evaluated and the mean of the feature space, and F mix represents the feature vector of the image to be evaluated, and μ k represents the mean of the feature space, represents the inverse matrix of the covariance matrix, which is used to eliminate feature correlation.

[0083] Step 7: Guide the unsupervised vision-language model CLIP for adaptive training according to the statistical features extracted from the low-quality images; the unsupervised vision-language model learns the hierarchical structure of underwater image quality through the fusion of statistical features extracted from high-quality images and vision-language information, and uses objective metrics for verification to ensure the accuracy of quality assessment, and finally uses the trained model to generate the corresponding quality assessment score, and judges the image quality according to the score.

[0084] Preferably, in Step 7, guide CLIP for multimodal fusion evaluation according to the statistical features, and the calculation formula is:

[0085]

[0086] where Q final represents the final image quality assessment score, Q semantic represents the original score output by the semantic model, Q stat represents the weighted score of the statistical features, μ stat and σ stat respectively represent the mean and variance of the statistical feature scores, β represents the fusion weight, and Sigmoid represents the normalization process.

[0087] Example 1:

[0088] To verify the generalization of the present invention for underwater different scene evaluations, multiple underwater image datasets were selected as the test set, and at the same time, the experimental results of the UIQM (Human-visual-system-inspired underwater image quality measures) algorithm, UCIQE (An underwater color image quality evaluation metric) algorithm, CCF (An imaging-inspired no-reference underwater color image quality assessment metric), NUIQ (Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective Metric), Twice Mixing (Twice mixing: a rank learning based quality assessment approach for underwater image enhancement), UIQI (Uiqi: A comprehensive quality evaluation index for underwater images), FDUM (A reference-free underwater image quality assessment metric in frequency domain) algorithms were compared and analyzed qualitatively and quantitatively.

[0089] As Figure 2 shown, the quality assessment method proposed by the present invention constructs a multi-level quality comparison and verification system in underwater scenes (E1 to E3). By dividing the images into three levels of Bad, Fair, and Perfect according to the degree of quality degradation, it can be clearly seen that in the E1 high turbidity scene, the Bad-level images show serious fogging and color deviation, while the Perfect-level images restore the coral texture details and blue-green spectral characteristics after being processed by the algorithm; in the E2 low illumination scene, the Fair-level images reconstruct the biological contours in the dark areas through this method; in particular, the objective quality score prediction error in the E3 dynamic scattering scene is less than 0.08, verifying the robustness of the algorithm under complex degradation coupling conditions. This method can not only accurately quantify the quality process of images from Bad to Perfect, but also significantly outperform traditional metrics.

[0090] As Figure 3As shown in (a)-(j), the quality assessment algorithm proposed by the present invention constructs a ten-level progressive quality verification system. The quality evolution process of the image from severe degradation to ideal enhancement can be clearly observed. In the first row of images, the law of decreasing turbidity is shown, and the outline gradually reveals the calcified structure from being blurred; in the second row of images, the color restoration characteristics are revealed, and the green spectral characteristics reach the best saturation. In particular, the predicted score of the algorithm of the present invention matches the true quality, accurately corresponding to the visually recognizable detail mutation points, which proves the scoring advantage of the evaluation model in this scenario.

[0091] In this embodiment, the experimental results of different algorithms are compared from four objective indicators: SROCC, KROCC, PLCC, and RMS. When SROCC takes 1, it indicates that the algorithm performance is very good, and -1 indicates very poor. The closer the value is to 1, the better the performance of the IQA algorithm. The larger the value of KROCC, the better the correlation between the two data, and the better the performance of the IQA algorithm; the smaller the value, the worse the correlation. In contrast, SROCC focuses on calculating the correlation degree between two vectors, while KROCC tends to evaluate the dependence strength of two vectors. PLCC describes the linear correlation between two sets of data, and its value range is -1 to 1. When the value of PLCC is zero, it means that the two sets of numbers are irrelevant. The closer the RMSE (root mean square error) is to 0, the better the performance of the IQA algorithm. From the data in Table 1, it can be seen that the indicators of this method are significantly improved in image quality assessment. The present invention uses image classification for underwater image quality assessment, adopts an objective mixing ratio to replace the subjective score, and fits the probability obtained from image classification. Therefore, the present invention has a large improvement in SROCC, KROCC, PLCC, and RMSE, and is superior to other underwater image restoration algorithms.

[0092] Table 1 Comparison of the processing results of the model of the present invention and other advanced algorithms on UIEB

[0093]

[0094] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An underwater image quality assessment method based on multimodal contrast learning and semantic guidance, characterized in that Including the following steps: Step 1: Obtain multiple low-quality images from multiple underwater image datasets, and for each low-quality image, obtain a high-quality reference image after image enhancement processing; the high-quality reference image is an image defined according to objective evaluation indicators or subjective scores; Step 2: According to the low-quality images and high-quality reference images obtained in Step 1, linearly mix each pair of low-quality images and high-quality reference images using a mixing ratio [0, 0.5, 1.0] to generate linearly mixed images; the linearly mixed images include multiple quality levels; where 0 represents a completely low-quality image, 1.0 represents a completely high-quality image, and 0.5 represents an equal mixture of a low-quality image and a high-quality reference image; Step 3: Through quality contrast learning and visual language alignment strategy, embed the linearly mixed images generated in Step 2 into quality semantic labels to achieve the docking of image features and text information; Step 4: Process image information through an image encoder; in the image encoder, use a multi-view weighted method to extract the features of low-quality images and their corresponding high-quality reference images, and measure the perceptual similarity of different quality images through a quality contrast loss function; The extracted image features include: low-level texture, edge features, and high-level semantic information; Step 5: Process text information through a text encoder; in the text encoder, according to the image perception results generated by the image encoder, assign text prompts to each image; group according to the quality levels of the images, and achieve accurate quality ranking through a group contrast loss function of relative ranking; Step 6: Extract the texture features, edge features, and semantic features of the high-quality reference images, and conduct a comparative analysis of the texture features, edge features, and semantic features with the same type of features extracted from the low-quality images in Step 4; by calculating the distance between the high-quality reference image and the low-quality image features, quantify the gap between the low-quality image and the high-quality reference image, and perceive the image quality level; Step 7: Guide the unsupervised vision-language model CLIP for adaptive training according to the statistical features extracted from the low-quality images; the unsupervised vision-language model learns the hierarchical structure of underwater image quality through the fusion of statistical features extracted from high-quality images and vision-language information, and uses objective indicators for verification to ensure the accuracy of quality assessment, and finally uses the trained model to generate corresponding quality assessment scores, and judge the image quality according to the scores; 2. The underwater image quality assessment method based on multi-modal contrast learning and semantic guidance according to claim 1, wherein The Step 2 generates image datasets with different quality levels, and by constructing training samples with continuous quality changes, the calculation formula is: I mix = αI low + (1 - α)I ref ; Among them, I mix represents the image obtained by linear mixing, I low represents the low-quality image, I ref represents the high-quality reference image. When α is 0, the features of the high-quality reference image are completely retained. When α is 0.5, the features of the moderately degraded image are simulated. When α is 1.0, the features of the original low-quality image are retained.

3. The underwater image quality assessment method based on multi-modal contrast learning and semantic guidance according to claim 1, characterized in that The joint alignment of image quality and text in Step 3 includes the following steps: Step 31: Construct a semantic description corresponding to quality text, and the semantic information is expressed as Perfect, Fair, and Bad; Step 32: Initialize the dual encoder using the contrastive language-image pre-training CLIP pre-training model. By optimizing the objective function, enable the model to align the features of images and texts with predefined semantic labels, thereby achieving the effective fusion of image and text information and embedding corresponding semantic information. The calculation formula is as follows: Among them, sin represents calculating the similarity between two vectors, E img represents the feature vector output by the image encoder, E text represents the feature vector output by the text encoder, T i and T j respectively represent text descriptions of different qualities, L align represents the alignment degree between the image and text features.

4. The underwater image quality assessment method based on multi-modal contrast learning and semantic guidance according to claim 1, characterized in that The multi-view feature extraction and contrast in step 4 include the following steps: Step 41: The image encoder adopts a feature fusion strategy of ResNet50 and a multi-scale attention module. The calculation formula is as follows: Among them, F fused is the fused feature representation, represents the feature map of the k-th layer of ResNet-50, W k represents the learnable weight parameter vector, and Softmax represents normalizing the weights to a probability distribution; Step 42: Perform dynamic triplet sampling within the same α group, select the corresponding anchor points, positive samples, and negative samples. The calculation formula for the loss function is as follows: Among them, used to calculate the similarity difference between the anchor, positive sample and negative sample through the contrastive learning method, d represents the cosine distance, F a 、F p and F n respectively represent the anchor, positive sample and negative sample features, α represents the boundary threshold, The loss controls the distance difference between the positive sample and the negative sample, ensuring that the distance between the negative sample and the anchor is greater than the distance gap between the positive sample and the anchor.

5. The underwater image quality assessment method based on multi-modal contrast learning and semantic guidance according to claim 1, characterized in that The hierarchical quality ranking in step 5 includes the following steps: Define the ranking constraint between quality groups through the group contrast ranking loss. The calculation formula is as follows: Among them, represents the sorting loss, and N g represents the number of quality groups, and S(q) represents the predicted score value of the model for the quality group q.

6. The underwater image quality assessment method based on multi-modal contrast learning and semantic guidance according to claim 1, wherein Step 6 includes the following steps: Step 61: Extract the statistical feature vectors from the high-quality reference image set and generate the feature space using K-means clustering; Step 62: Calculate the Mahalanobis distance between the mixed image features generated in step 61 and the nearest prototype features. The calculation formula is as follows: Among them, d Mahalanobis represents the distance between the feature vector of the image to be evaluated and the mean of the feature space, F mix represents the feature vector of the image to be evaluated, μ k represents the mean of the feature space, represents the inverse matrix of the covariance matrix, which is used to eliminate feature correlation.

7. The underwater image quality assessment method based on multi-modal contrast learning and semantic guidance according to claim 1, characterized in that In step 7, conduct the multi-modal fusion evaluation by guiding CLIP based on the statistical features. The calculation formula is as follows: Among them, Q final represents the final image quality assessment score, Q semantic represents the original score output by the semantic model, Q stat represents the weighted score of statistical features, μ stat and σ stat respectively represent the mean and variance of the statistical feature scores, β represents the fusion weight, and Sigmoid represents the normalization process.