Underwater image enhancement method and system based on edge feature attention fusion
Through the method based on attention fusion based on edge features, the problem of underwater image enhancement in the prior art is solved, and a higher quality underwater image enhancement effect is achieved.
Patent Information
- Application Number
- CN202510144126.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
Existing underwater image enhancement methods are difficult to fully cope with multiple interferences in complex underwater environments, resulting in unstable image quality and difficult to meet the needs of high-quality underwater images.
The underwater image enhancement model is constructed through multi-level feature extraction, edge detection and attention mechanism fusion, step-by-step amplification module and multi-dimensional visual loss function.
Improves the image color accuracy, contrast and edge clarity, and enhances the model's adaptability and image enhancement effect to complex underwater environments.
Smart Images

Figure CN120070219A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an underwater image enhancement method, specifically an underwater image enhancement method and system based on edge feature attention fusion, belonging to the technical field of image processing. Background Art
[0002] Generally speaking, underwater images refer to image data of the underwater environment captured by devices such as remotely operated vehicles (ROVs). They have important application values in fields such as ocean exploration, underwater archaeology, and fishery monitoring. However, the underwater imaging environment is extremely complex. The light reflection and scattering caused by high pressure, terrain, and organisms on the seabed, the interference of seabed impurities on optical properties, and the obstruction of light propagation by sea ice and bubbles in low-temperature sea areas all result in serious problems in underwater images, such as severe color distortion, reduced contrast, and blurred details. In addition, the absorption and scattering of light by water further exacerbate the degradation of underwater images, leading to phenomena such as similar colors between objects and the background and difficult-to-distinguish edges, making it difficult for underwater images to meet the requirements of precise research in multiple fields in terms of key indicators such as color accuracy, contrast, and edge sharpness.
[0003] In the prior art, an underwater image enhancement method and device disclosed in Publication No. CN118710528A include: obtaining an underwater image to be enhanced, and encoding the underwater image to be enhanced through a first image enhancement encoder to obtain a low-dimensional feature vector to be enhanced, where the first image enhancement encoder is obtained by optimizing according to first encoding and decoding optimization data; inputting the low-dimensional feature vector to be enhanced into a target mapping network to obtain a low-dimensional feature vector of target quality, where the target mapping network is trained by using the data of the first encoding and decoding process as input data and the data of the second encoding and decoding process as output data for a neural network model based on an attention mechanism; decoding the low-dimensional feature vector of target quality through a second image enhancement encoder to obtain an enhanced underwater image, where the second image enhancement encoder is obtained by optimizing according to second encoding and decoding optimization data. The underwater image enhancement technology aims to improve the quality of underwater images by eliminating their unique degradation features and making the images closer to the authenticity and clarity under normal lighting conditions. However, traditional underwater image enhancement methods usually rely on single feature extraction or simple image processing techniques and are difficult to comprehensively cope with multiple interferences brought by the complex underwater environment. For example, color correction methods may not be able to effectively restore edge details, while contrast enhancement-based technologies may ignore the accuracy of color information. At the same time, due to the diversity and complexity of the underwater environment, single enhancement methods are often difficult to meet the requirements of different scenarios, resulting in unstable enhancement effects and making it difficult to meet the urgent need for high-quality underwater images in practical applications. Summary of the Invention
[0004] The object of the present invention is to provide an underwater image enhancement method and system based on edge feature attention fusion in order to solve at least one of the above technical problems.
[0005] The present invention realizes the above object through the following technical solutions: An underwater image enhancement method based on edge feature attention fusion, the underwater image enhancement method includes the following steps:
[0006] S1. By extracting multi-level features of the underwater image, multi-level features of the image are obtained;
[0007] S2. By using an edge detection tool to extract the edge information of the underwater image, and through an attention mechanism, the edge information is fused with the multi-level features to obtain the re-distributed weighted features;
[0008] S3. Through a step-by-step amplification module, the multi-level features and the weighted features are fused to restore the clarity of the underwater image and generate an enhanced underwater image;
[0009] S4. By comparing the enhanced underwater image with the underwater image, the visual difference loss and the contrast loss between the two are calculated;
[0010] S5. According to the visual difference loss and the contrast loss, as well as the outlier loss calculated during the training process, a comprehensive loss function is constructed, and the underwater image enhancement model is trained using this loss function to obtain an underwater image enhancement model based on edge feature attention fusion.
[0011] As a further scheme of the present invention: The extraction of multi-level features of the underwater image refers to gradually extracting features of different sizes from the underwater features of the underwater image data to obtain the representation of the underwater image features in the underwater image captured by the underwater image, specifically including:
[0012] S11. By using a filter to extract features of the underwater image, multi-level detailed feature maps are obtained;
[0013] S12. The multi-level detailed feature maps are divided into multiple small blocks;
[0014] S13. These small blocks are sequentially processed through a feature extraction module to obtain the representation of the multi-level features of the underwater image.
[0015] As a further scheme of the present invention: The feature extraction module uses a transformer module to extract features of the multi-level detailed feature maps to obtain underwater image features, specifically including:
[0016] S131. Through the attention mechanism, linearly process the multi-level detailed feature maps input into the transformer module to obtain the query vectors, key-value vectors, and value vectors of the multi-level detailed feature maps;
[0017] S132. Normalize the query vectors, key-value vectors, and value vectors through a normalization function to obtain the output of the transformer module, and obtain the representation of the underwater image features based on the output.
[0018] As a further solution of the present invention: The fusion of edge information and multi-level features is performed by a multi-feature cross-fusion module, specifically including:
[0019] S21. Extract the edge information of the multi-level features through an edge detection tool;
[0020] S22. Through a cross-channel attention mechanism, fuse the edge information with the multi-level features to enhance the information interaction between different channels;
[0021] S23. Through a cross-space attention mechanism, capture the long-range correlations between the multi-level features to enhance the semantic information in space.
[0022] As a further solution of the present invention: The cross-channel attention mechanism is implemented through the following steps:
[0023] S221. Concatenate the multi-level features and the edge information along the channel direction to generate unified keys and values;
[0024] S222. Generate query vectors through a 1x1 convolution operation;
[0025] S223. Calculate the attention weights through the Softmax function to fuse the feature information of different channels.
[0026] As a further solution of the present invention: The cross-space attention mechanism is implemented through the following steps:
[0027] S231. Concatenate the multi-level features along the channel direction to generate queries and keys;
[0028] S232. Generate value vectors through a 1x1 convolution operation;
[0029] S233. Calculate the spatial attention weights through the Softmax function to enhance the feature correlations at different spatial positions.
[0030] As a further solution of the present invention: The step-by-step amplification module is implemented through the following steps:
[0031] S31. Amplify the output of the last stage of the feature extraction module and input the output into the feature extraction module;
[0032] S32, fusing the output of the multi-feature cross fusion module with the amplification results of each stage of the step-by-step amplification module step by step;
[0033] S33, restoring the clarity of the underwater image by gradually enlarging it, and finally generating an enhanced underwater image.
[0034] As a further solution of the present invention: the calculation of visual difference loss and contrast loss includes edge feature extraction and edge information extraction;
[0035] Among them, edge feature extraction includes:
[0036] Use the Sobel operator to calculate the gradient of the underwater image in the horizontal and vertical directions to generate edge features;
[0037] The edge features are weightedly fused with the multi-level detail feature map output by the transformer module to enhance the details of the underwater image.
[0038] Edge information extraction includes:
[0039] Use edge detection tools to calculate the gradient of underwater images in horizontal and vertical directions to generate edge information;
[0040] The edge information is weightedly fused with the multi-level detail feature map output by the feature extraction module to enhance the details of the image.
[0041] As a further solution of the present invention: the loss function is used to train the underwater image enhancement model, including:
[0042] Visual difference loss: evaluates visual similarity by comparing the feature differences between the predicted image and the real image in the pre-trained model;
[0043] Multi-scale structural similarity loss: Evaluate image quality by multi-scale structural similarity;
[0044] Smoothness loss: Reduce noise and maintain image sharpness through a smooth absolute error loss.
[0045] An underwater image enhancement system based on edge feature attention fusion includes a memory and a processor; wherein the memory is used to store a computer program; and the processor is used to implement an underwater image enhancement method when executing the computer program.
[0046] The beneficial effects of the present invention are:
[0047] 1) By performing multi-level feature extraction and edge feature fusion on underwater images, local details and global context information of the images can be captured simultaneously, thus obtaining a richer and more comprehensive feature representation. Then, through contrastive learning optimization, a visual perception loss is obtained based on the multi-level feature representation to improve the accuracy of image enhancement;
[0048] 2) By guiding the model to learn the optimization objective of image quality through a multi-dimensional visual loss function, the limitations or biases that may exist in a single enhancement method can be overcome, improving the comprehensiveness and reliability of image enhancement. At the same time, it also helps to improve the color accuracy, contrast, and edge sharpness of the image;
[0049] 3) A multi-dimensional visual loss function is constructed through multi-level feature representation to guide the model to learn the optimization objective of image quality. On this basis, the initial model is trained by combining visual perception loss, multi-dimensional structural similarity loss, and Charbonnier loss. The obtained underwater image enhancement model with edge attention fusion can fuse feature information at different levels with spatial context information and use a multi-dimensional perception optimization method to improve the adaptability to complex underwater environments and the image enhancement effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic flow chart of the whole of the present invention;
[0051] Figure 2 is one of the schematic flow charts of the local-global feature extraction of the present invention;
[0052] Figure 3 is the second of the schematic flow charts of the local-global feature extraction of the present invention;
[0053] Figure 4 is the schematic flow chart in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Embodiment 1, as Figure 1 shown, this embodiment provides an underwater image enhancement method based on edge feature attention fusion. The underwater image enhancement method includes the following steps:
[0056] First: By extracting multi-level features (multi-level details) of the underwater image, a multi-level feature representation of the underwater image is obtained.
[0057] The extraction of multi-level features of underwater images refers to gradually extracting features of different sizes from the underwater features of underwater image data to obtain a representation of the underwater image features in the captured underwater images, which specifically includes:
[0058] 1) Extract features from the underwater image through a filter of a fixed size to obtain multi-level detailed feature maps;
[0059] 2) Segment the multi-level detailed feature maps, and the multi-level detailed feature maps are divided into multiple small blocks;
[0060] 3) Process the small blocks sequentially through a three-layer feature extraction module to obtain a representation of the multi-level features of the underwater image;
[0061] Among them, the feature extraction module uses the transformer module to extract features from the multi-level detailed feature maps to obtain a representation of the underwater image features, which includes:
[0062] 31) Through the attention mechanism, linearly process the multi-level detailed feature maps input to the transformer module to obtain the query vector, key-value vector, and value vector of the multi-level detailed feature maps;
[0063] 32) Normalize the query vector, key-value vector, and value vector through a normalization function to obtain the output of the transformer module, and obtain a representation of the underwater image features according to the output.
[0064] Second: Extract the edge information of the underwater image through an edge detection tool, and fuse the edge information with the multi-level features through the attention mechanism to obtain the reallocated weighted features.
[0065] The fusion of edge information and multi-level features is carried out using a multi-feature cross-fusion module (DDEM), which specifically includes:
[0066] 1) Extract the edge information of the multi-level features through an edge detection tool;
[0067] 2) Fuse the edge information with the multi-level features through a cross-channel attention mechanism to enhance the information interaction between different channels;
[0068] Among them, the cross-channel attention mechanism is implemented through the following steps:
[0069] 21) Concatenate the multi-level features and edge information along the channel direction to generate unified keys and values;
[0070] 22) Generate a query vector through a 1x1 convolution operation;
[0071] 23) Calculate the attention weights through the Softmax function to fuse the feature information of different channels;
[0072] 3) Capture the long-range correlations between multi-level features through the cross-space attention mechanism to enhance the semantic information in space;
[0073] 31) Concatenate the multi-level features along the channel direction to generate queries and keys;
[0074] 32) Generate value vectors through 1x1 convolution operations;
[0075] 33) Calculate the spatial attention weights through the Softmax function to enhance the feature correlations at different spatial positions.
[0076] Third: Through the step-by-step amplification module, fuse the multi-level features and weighted features, gradually restore the clarity of the underwater image, and generate the enhanced underwater image.
[0077] Among them, the step-by-step amplification module is implemented through the following steps:
[0078] 1) Amplify the output of the last stage of the feature extraction module and input this output into the feature extraction module;
[0079] 2) Gradually fuse the output of the multi-feature cross-fusion module with the amplification results of each stage of the step-by-step amplification module;
[0080] 3) Restore the clarity of the underwater image through the step-by-step amplification module, and finally generate the enhanced underwater image.
[0081] Fourth: By comparing the enhanced underwater image with the underwater image (original image), calculate the visual difference loss and contrast loss between the two.
[0082] The calculation of the visual difference loss and contrast loss includes edge feature extraction and edge information extraction; among them, edge feature extraction includes:
[0083] Use the Sobel operator to calculate the gradients of the underwater image in the horizontal and vertical directions to generate edge features;
[0084] Weightedly fuse the edge features with the multi-scale features (multi-level detailed feature maps) output by the Transformer module to enhance the details of the underwater image;
[0085] Edge information extraction includes:
[0086] Use an edge detection tool to calculate the gradients of the underwater image in the horizontal and vertical directions to generate edge information;
[0087] Fuse the edge information with the multi-level features output by the feature extraction module to enhance the details of the underwater image.
[0088] Fifth: Construct a comprehensive loss function according to the visual difference loss, contrast loss, and outlier loss calculated during the training process, and use the loss function to train the underwater image enhancement model to obtain an underwater image enhancement model based on edge feature attention fusion.
[0089] The training of the underwater image enhancement model by the loss function includes:
[0090] 1) Visual difference loss: Evaluate the visual similarity by comparing the feature differences between the predicted image and the real image in the pre-trained model.
[0091] 2) Multi-scale structural similarity loss: Evaluate the image quality through multi-scale structural similarity.
[0092] 3) Smooth loss: Reduce noise and maintain the sharpness of the image through smooth absolute error loss.
[0093] Embodiment 2 provides an underwater image enhancement system based on edge feature attention fusion. The underwater image enhancement system includes a memory and a processor. Among them, the memory is used to store computer programs, and the processor is used to implement the underwater image enhancement method based on edge feature attention fusion in Embodiment 1 when executing the computer program.
[0094] In this embodiment, the processor trains the initial classification model according to the perceptual loss, multi-scale structural similarity loss, and Charbonnier loss. Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory; volatile memory can include random access memory (RAM) or external cache memory.
[0095] By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0096] Embodiment 3, as Figure 4 shown, this embodiment provides an underwater image enhancement method based on edge feature attention fusion.
[0097] In this embodiment, by extracting multi-level features of the underwater image, a representation of the multi-level features of the underwater image is obtained. Combining Figure 4 shown, the Transformer module is used to extract spatial information and low-level feature information through convolutions of different sizes, optimize and reduce the size of the multi-level detail feature maps, and then perform a splitting operation based on convolutional splitting, so as to obtain a representation of the multi-level features of the underwater image.
[0098] In this embodiment, the edge information of the underwater image is extracted through an edge detection tool, and the edge information is fused with the multi-level features through an attention mechanism to obtain the reallocated weighted features. This step applies a feature dispersion method, and the feature representation of the underwater image can be extracted separately through a deep encoder network, and the deep encoder network is a Transformer module based on a convolutional strategy. Combining Figure 4 shown, the feature network is used to transform the feature representation, calculate the structural similarity, and construct a loss by comparing the images before and after enhancement, and the network is constrained to obtain the underwater feature representation of the image, and a Charbonnier loss is obtained.
[0099] By extracting multi-level features of the underwater image, a representation of the multi-level features of the underwater image is obtained, enabling the underwater image enhancement model to simultaneously capture local details and global context information in the underwater image, thus obtaining a richer and more comprehensive underwater feature representation; then, the change before and after image enhancement is calculated through visual perception comparison to obtain a perceptual loss, and the perceptual loss emphasizes the feature changes of different underwater features within the same image, thereby enhancing the distinction of feature regions and improving the effectiveness of the underwater image enhancement model.
[0100] In this embodiment, by performing multi-level feature extraction on the underwater image, a representation of the multi-level features of the underwater image is obtained, including: extracting features from the underwater image data through a convolution kernel of a preset size to obtain a multi-scale feature map (a multi-level detailed feature map); splitting the multi-scale feature map to divide the underwater image data into multiple feature maps; sequentially passing the feature maps through three Transformer modules for feature extraction to obtain a multi-level feature representation; where, combined with Figure 2 As shown in, for the feature extraction of the Transformer module based on the convolution splitting strategy, first obtain a multi-scale feature map through the convolution splitting strategy, then split the multi-scale feature map, and input the generated feature map blocks after splitting into the subsequent Transformer module. This strategy uses a convolutional neural network to increase the ability to extract low-level feature information, thereby reducing the size of the feature map, reducing the operation parameters, and including spatial information and low-level features. Combined with Figure 3 As shown in, through the stacking of three Transformer modules, the fusion and representation of features can be further deepened, enabling the underwater image enhancement model to more precisely understand and represent underwater features. Among them, the Transformer module consists of a multi-layer perceptron and a multi-head attention mechanism for extracting global features.
[0101] Through the convolution operation of a preset size, it is possible to simultaneously extract the features of the underwater image data at different scales, which helps to capture the performance of underwater features at different resolutions and increases the richness of features. After feature extraction, the size of the feature map is reduced through a pooling operation, effectively reducing the operation parameters in subsequent processing, thereby reducing the computational complexity and resource consumption of the model. Use the Transformer module, especially the multi-head attention mechanism therein, to capture the long-range dependencies in the underwater image data and extract global features.
[0102] In this embodiment, the feature maps are sequentially passed through three Transformer modules for feature extraction to obtain a multi-level feature representation, including: linearly processing the feature maps input to the Transformer module through the attention mechanism to obtain the query vector, key-value vector, and value vector of the feature maps; normalizing the query vector, key-value vector, and value vector through a normalization function to obtain the output of the Transformer module, and obtaining a multi-level feature representation based on the output.
[0103] In this embodiment, the residual idea is adopted to connect the output features of the Transformer module to achieve effective fusion of shallow features and deep features. The attention mechanism formula is used for connection, and the attention mechanism formula is as follows:
[0104]
[0105] Where: Q, K, and V are the input query vector, key-value vector, and value vector respectively, d_k is the dimension of the query vector, key-value vector, and value vector, and softmax represents normalizing the weights so that the sum of the query vector, key-value vector, and value vector equals 1.
[0106] Adopting the residual idea, the output features of the Transformer module are connected to the input features, realizing the effective fusion of shallow and deep features.
[0107] In this embodiment, through the edge fusion module, based on the combination of multi-scale features and edge features, the enhanced features corresponding to each underwater image data are determined, including: for an underwater image data, extracting the multi-scale features and edge features corresponding to the underwater image data; performing weighted combination of the edge features and multi-scale features in the channel dimension to obtain preliminary fusion features; dynamically adjusting the weights of the edge features through the channel cross-attention mechanism to ensure accurate transmission of details such as color and texture; averaging the fusion features of all scales to obtain enhanced features. To calculate the enhanced features through the edge fusion module, it is necessary to calculate the fusion features for each scale. The calculation method of the fusion features is: select the edge features and multi-scale features of the current scale, calculate the eigenvalue after their weighted combination, and adjust the weights through the channel cross-attention mechanism. Finally, average the fusion features of all scales to obtain the final scale-enhanced features.
[0108] In this embodiment, through the spatial information enhancement module, based on the spatial correlation of the enhanced features, the spatial enhanced features corresponding to each underwater image data are determined, including: for an underwater image data, extracting the enhanced features corresponding to the underwater image data; calculating the correlation between the enhanced features at different spatial positions to dynamically adjust the feature weights of different regions of the image; strengthening key detail information through the spatial cross-attention mechanism to reduce interference in smooth regions; averaging the enhanced features at all spatial positions to obtain spatial enhanced features. To calculate the spatial enhanced features through the spatial information enhancement module, it is necessary to calculate the correlation weights for each spatial position. The calculation method of the correlation weights is: select the feature vectors of the enhanced features at different spatial positions, calculate their similarity, and adjust the weights through the spatial cross-attention mechanism. Finally, average the enhanced features at all spatial positions to obtain the final spatial enhanced features.
[0109] In this embodiment, the feature post-processing module determines the final enhanced feature corresponding to each underwater image data through size adjustment and non-linear transformation of the spatial enhanced feature, including: for an underwater image data, extracting the spatial enhanced feature corresponding to the underwater image data; performing normalization processing on the spatial enhanced feature and introducing non-linear transformation through an activation function; restoring the resolution of the feature map through an upsampling layer to ensure better reconstruction of image details in the decoding stage; performing channel compression and combination of features through 1x1 convolution to enhance the expression ability of the model; finally combining all processed features to obtain the final enhanced feature. To calculate the final enhanced feature through the feature post-processing module, size adjustment and non-linear transformation need to be performed on each feature map. The calculation method of the final enhanced feature is: select the spatial enhanced feature, perform normalization, activation function, upsampling, and 1x1 convolution operations in sequence, and finally combine all processed features.
[0110] Through the above steps, the present invention can combine the edge fusion module, the spatial information enhancement module, and the feature post-processing module to comprehensively optimize the performance of the underwater image enhancement model, ensuring that the enhanced image reaches the optimal level in terms of detail restoration, spatial consistency, and visual quality.
[0111] In the embodiment of the present invention, the MS-SSIM loss determines the MS-SSIM loss corresponding to each underwater image data according to the structural similarity between the enhanced image and the target image, including: for an underwater image data, determining the structural similarity between the enhanced image corresponding to the underwater image data and the target image; dividing the logarithm value of the structural similarity by the sum of the structural similarities at all scales corresponding to the underwater image data as the loss value of the multi-scale feature representation; averaging all loss values of the multi-scale feature representation to obtain the MS-SSIM loss. To calculate the loss value through the MS-SSIM loss function, the loss value needs to be calculated for each scale. The calculation method of the loss value is: select the brightness, contrast, and structural components of the enhanced image and the target image at the current scale, calculate the similarity value after their weighted fusion, divide it by the sum of the similarity values of all scales, and finally average all scale loss values to obtain the final MS-SSIM loss. Among them, the loss function is:
[0112]
[0113] In the formula: M represents different scales; u $ , u # respectively represent the means of the predicted image and the true value; σ # , σ $ represent the standard deviations between the predicted image and the true image; σ #$ represents the covariance between the predicted image and the true image; β m , γm represents the relative importance constant between two terms; c 1 , c ( represents the constant term to prevent the divisor from being zero.
[0114] According to the high-level semantic feature differences between the enhanced image and the target image through perceptual loss, determine the perceptual loss corresponding to each underwater image data, including: for an underwater image data, extract the high-level semantic features of the enhanced image and the target image corresponding to the underwater image data; calculate the difference value between the high-level semantic features and use the difference value as the loss value represented by the perceptual feature; average all the loss values represented by the perceptual features to obtain the perceptual loss. In this embodiment, the loss value is calculated through the perceptual loss function, and the loss value needs to be calculated for each sample. The loss value calculation method is: select the high-level feature maps of the enhanced image and the target image in the pre-trained network, calculate the mean square error between their features, and divide by the size of the feature map. Finally, average the loss values of all samples.
[0115] Among them, the loss function is:
[0116] l #-. = |x - y|
[0117] According to the pixel-level differences between the enhanced image and the target image through Charbonnier loss, determine the Charbonnier loss corresponding to each underwater image data, including: for an underwater image data, calculate the pixel-level differences between the enhanced image and the target image corresponding to the underwater image data; input the pixel-level differences into the Charbonnier loss function to calculate its smooth loss value; average all the smooth loss values of the pixel-level differences to obtain the Charbonnier loss. In this embodiment, the loss value is calculated through the Charbonnier loss function, and the loss value needs to be calculated for each pixel point. The loss value calculation method is: select the difference value of the enhanced image and the target image at each pixel point, calculate the square root of the sum of its square and a constant, and divide by the total number of pixels. Finally, average the loss values of all pixel points. Among them, the loss function is:
[0118]
[0119] In the formula: x represents the difference between the predicted image and the Ground truth, and represents a small positive number used for numerical stability.
[0120] Through the above steps, the present invention can combine MS-SSIM loss, perceptual loss and Charbonnier loss to comprehensively optimize the performance of the underwater image enhancement model, ensuring that the enhanced image reaches the optimal level in terms of visual quality, semantic consistency and robustness.
[0121] Working principle: By constructing an attention-guided edge feature fusion module, edge information is extracted using an edge operator, and multi-scale features and edge features of the image are fused through channel cross-attention to enhance the details of objects in the image; by constructing a spatial information enhancement module, the correlation information of different spatial positions of the enhanced features is captured using the feature interaction mechanism of spatial cross-attention to enhance the semantic representation of the features in terms of spatial structure; by constructing a multi-dimensional perception optimization method, the image enhancement effect is improved by combining perception optimization, structure optimization, and anomaly optimization. Perception optimization improves the semantic and visual perception quality of the image, structure optimization enhances the global structure information and local contrast of the image, and anomaly optimization reduces the impact of outliers on the image.
[0122] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed invention.
[0123] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. An underwater image enhancement method based on edge feature attention fusion, characterized in that: The underwater image enhancement method comprises the following steps: S1, extracting multi-level features of the underwater image to obtain the multi-level features of the underwater image; S2, extracting edge information of the underwater image by using an edge detection tool, and fusing the edge information with the multi-level features by using an attention mechanism to obtain redistributed weighted features; S3, by gradually enlarging the module, fusing the multi-level features and the weighted features, restoring the clarity of the underwater image, and generating an enhanced underwater image; S4, calculating the visual difference loss and contrast loss between the enhanced underwater image and the underwater image; S5. According to the visual difference loss and contrast loss, as well as the outlier loss calculated during the training process, a comprehensive loss function is constructed, and the underwater image enhancement model is trained using the loss function to obtain an underwater image enhancement model based on edge feature attention fusion.
2. The underwater image enhancement method according to claim 1, characterized in that: In S1, the underwater image multi-level feature extraction refers to extracting features of different sizes from underwater features of underwater image data step by step to obtain the representation of underwater image features in the underwater image capture image, specifically including: S11, extracting features from the underwater image through a filter to obtain a multi-level detail feature map; S12, dividing the multi-level detail feature map into multiple small blocks; S13, processing the small blocks in turn through feature extraction modules to obtain a multi-level feature representation of the underwater image.
3. The underwater image enhancement method according to claim 2, characterized in that: The feature extraction module uses a transformer module to extract features of a multi-level detail feature map to obtain underwater image features, specifically including: S131, linearly processing the multi-level detail feature map input to the transformer module through the attention mechanism to obtain a query vector, a key-value vector, and a value vector of the multi-level detail feature map; S132. Normalize the query vector, the key-value vector, and the value vector through a normalization function to obtain the output of the transformer module, and obtain the representation of the underwater image features based on the output.
4. The underwater image enhancement method based on edge feature attention fusion according to claim 2, characterized in that: In S2, the edge information and the multi-level features are fused using a multi-feature cross fusion module, specifically including: S21, extracting edge information of the multi-level features by using an edge detection tool; S22, fusing the edge information with the multi-level features through a cross-channel attention mechanism to enhance information interaction between different channels; S23. The long-distance correlation between the multi-level features is captured through a cross-space attention mechanism to enhance the spatial semantic information.
5. The underwater image enhancement method according to claim 4, characterized in that: The cross-channel attention mechanism is implemented by the following steps: S221, splicing the multi-level features and edge information along the channel direction to generate a unified key and value; S222, generating a query vector through a 1x1 convolution operation; S223. Calculate the attention weight through the Softmax function and integrate the feature information of different channels.
6. The underwater image enhancement method according to claim 4, characterized in that: The cross-space attention mechanism is implemented by the following steps: S231, splicing the multi-level features along the channel direction to generate queries and keys; S232, generate a value vector through a 1x1 convolution operation; S233. Calculate the spatial attention weight through the Softmax function to enhance the feature association of different spatial positions.
7. The underwater image enhancement method according to claim 4, characterized in that: In S3, the step-by-step enlargement module is implemented by the following steps: S31, amplifying the output of the final stage of the feature extraction module, and inputting the output into the feature extraction module; S32, fusing the output of the multi-feature cross fusion module with the amplification results of each stage of the step-by-step amplification module step by step; S33, restoring the clarity of the underwater image through the step-by-step enlargement module, and finally generating an enhanced underwater image.
8. The underwater image enhancement method according to claim 7, characterized in that: In S4, the calculation of visual difference loss and contrast loss includes edge feature extraction and edge information extraction; Wherein, the edge feature extraction includes: Use the Sobel operator to calculate the gradient of the underwater image in the horizontal and vertical directions to generate edge features; The edge features are weightedly fused with the multi-level detail feature map output by the transformer module to enhance the details of the underwater image. The edge information extraction comprises: Use edge detection tools to calculate the gradient of underwater images in horizontal and vertical directions to generate edge information; The edge information is weightedly fused with the multi-level detail feature map output by the feature extraction module to enhance the details of the underwater image.
9. The underwater image enhancement method according to claim 1, characterized in that: In S5, the loss function for training the underwater image enhancement model includes: Visual difference loss: evaluates visual similarity by comparing the feature differences between the predicted image and the real image in the pre-trained model; Multi-scale structural similarity loss: Evaluate image quality by multi-scale structural similarity; Smoothness loss: Reduce noise and maintain image sharpness through a smooth absolute error loss.
10. An underwater image enhancement system based on edge feature attention fusion, characterized in that: include: Memory for storing computer programs; A processor, used for implementing the underwater image enhancement method based on edge feature attention fusion as described in any one of claims 1 to claim 9 when executing the computer program.
Citation Information
Patent Citations
Underwater image enhancement method and device
CN118710528A