Underwater video enhancement method based on semantic guidance

The method addresses temporal inconsistency in water under video enhancement by integrating global and regional features with semantic guidance, achieving improved video quality and stability through enhanced frame alignment and attention mechanisms.

CN120318137APending Publication Date: 2025-07-15DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510257034.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing underwater video enhancement methods ignore the time continuity between frames, resulting in poor time consistency, inter-frame flickering and insufficient context information in the video sequence.

Method used

Using a semantic guidance-based underwater video enhancement method, through a multi-step process of global feature enhancement, region enhancement and semantic guidance, the encoder-decoder architecture, packet spatial shift and context information enhancement feature representation are used to combine spatial and pixel attention mechanisms to capture inter-frame timing information and suppress noise.

Benefits of technology

Improves the clarity and stability of underwater videos, reduces inter-frame flickering, enhances the contrast and color saturation of the video, and improves the visual effect and usability of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318137A_ABST
    Figure CN120318137A_ABST
Patent Text Reader

Abstract

The invention provides an underwater video enhancement method based on semantic guidance, and the method comprises the following steps: S1, carrying out the global feature enhancement of a multi-frame video, and obtaining a feature map after global enhancement; s2, performing region enhancement on the context features of the multi-frame video to obtain a feature map after region enhancement; s3, performing semantic guidance on the feature map after region enhancement to obtain a region enhancement feature map after semantic guidance; and S4, fusing the semantically guided region enhanced feature map and the globally enhanced feature map to obtain an output effect of the intermediate frame. According to the global feature enhancement method disclosed by the invention, an encoder-decoder architecture is mainly used, multi-frame time sequence information is fully utilized to perform spatial alignment of adjacent multi-frame features by using grouping spatial shift, context information is fully learned to suppress irrelevant noise, the robustness of feature representation is enhanced, and the effect of global feature enhancement is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and deep learning technology, and in particular to an underwater video enhancement method based on semantic guidance. Background Art

[0002] In the underwater imaging process, due to the particularity of the underwater environment, such as poor light penetration and many suspended particles, even if high-end image and video acquisition equipment is used, the obtained videos and images are often affected by problems such as color difference and blur. Therefore, it is particularly important to develop processing technology that can effectively improve the quality of underwater videos, which can not only improve the clarity and smoothness of the video, but also enhance its contrast and color saturation, so as to better serve marine engineering and scientific research. In order to make full use of underwater video data and improve its availability and accuracy, video enhancement and optimization is an indispensable step.

[0003] At present, underwater video enhancement is mainly divided into two categories: frame-by-frame enhancement and multi-frame enhancement. The frame-by-frame enhancement method mainly uses traditional underwater image enhancement methods to enhance each frame of the video separately. Underwater image enhancement methods are divided into three categories: physical model-based methods, model-free methods, and deep learning-based methods. Model-free methods rely on hardware equipment and optical technology to improve visual effects by adjusting image pixel values. Physical model-based methods restore image features by establishing a physical model of underwater imaging and using model parameters to reverse the degradation effect. The model is usually based on certain assumptions and priors. The parameter estimation algorithm is highly complex and has certain limitations. Deep learning methods enhance images by using multiple deep learning network models to learn image features. These methods can enhance the quality of underwater videos to a certain extent, but this method ignores the temporal continuity between frames, and often leads to problems such as poor temporal consistency, inter-frame flickering, and insufficient context information in video sequences. The multi-frame enhancement method is a deep learning method specifically for underwater video enhancement. It mainly aggregates multi-frame temporal information to improve the visual effect of the video while ensuring the accuracy of alignment. However, due to the small number of existing underwater video datasets, the training of this method faces challenges. Summary of the invention

[0004] In view of this, the purpose of the present invention is to propose an underwater video enhancement method based on semantic guidance to solve the technical problem that the existing underwater video enhancement method ignores the temporal continuity between frames.

[0005] The technical means adopted by the present invention are as follows:

[0006] A semantically guided underwater video enhancement method comprises the following steps:

[0007] S1. Perform global feature enhancement on multiple frames of video to obtain a globally enhanced feature map;

[0008] S2. Perform regional enhancement on the context features of multiple frames of video to obtain a regionally enhanced feature map;

[0009] S3. Perform semantic guidance on the regionally enhanced feature map to obtain a semantically guided regionally enhanced feature map;

[0010] S4. Fuse the semantically guided regionally enhanced feature map with the globally enhanced feature map to obtain the output effect of the intermediate frame.

[0011] Further, S1 specifically includes the following steps:

[0012] S11. Divide multiple frames of video into several different scales;

[0013] S12. Uniformly divide the features of the i-th frame of video with the same scale along the channel dimension to obtain a subset of features;

[0014] S13. Perform a pixel-based spatial offset operation on each subset of features to obtain a spatially offset subset of features;

[0015] S14. Re-concatenate the spatially offset subset of features along the channel dimension to form a spatially offset processed subset of features;

[0016] S15. Filter out non-critical noise in the spatially offset processed subset of features through a context enhancement method to obtain feature maps of different scales;

[0017] S16. Merge the feature maps of different scales through the upsampling operation of the decoder to form a globally enhanced feature map;

[0018] Further, the specific formula of S11 is as follows:

[0019]

[0020] The specific formula of S13 is as follows:

[0021]

[0022] The specific formula of S14 is as follows:

[0023] K′ i = Concat(K′ i,1 ,K′ i,2 ,...,K′ i,8 )

[0024] Where, V iThe original input frame is \(I\), Encoder performs multi-scale encoding operations, Split represents the process of splitting features along the channel dimension, Offest is the process of performing spatial offset operations on the sub-feature sets, Concat represents the process of concatenating features along the channel dimension, and \(K'\) i is the feature frame obtained by grouped spatial shift.

[0025] Further, S15 specifically includes the following steps:

[0026] S151: Use a 3×3 convolutional layer to map the feature set after spatial offset processing to a single-channel mask, calculate the attention mask to capture the context information in the input features. The attention mask is then normalized through the softmax function. Use an upsampling to convert the input feature map into a matrix, and then multiply it by the attention mask to calculate the context feature \(I\) i , and the specific formula is as follows:

[0027] \(I\) i = Upsample(Softmax(Conv(\(K'\) i )))·Conv(\(K'\) i )

[0028] S152: Use the spatial attention and pixel attention mechanisms to implicitly capture the inter-frame temporal information to obtain the feature \(I''\) i , and the specific formula is as follows:

[0029] \(Q1, K1, V1 = Groupconv(Conv(I_i))\)

[0030]

[0031] S153: Add the feature \(I''\) i and \(I\) i to obtain the feature effect The specific formula is as follows:

[0032] \(Q2, K2, V2 = Groupconv(Conv(I'))\) i )

[0033] \(X\) i = Conv(softmax(Norm(Q2)·Norm(K2) T ))·t·V2)+\(I\) i

[0034] where Conv is a 3×3 convolution, softmax is the softmax normalization process, Upsample is the upsampling operation, Groupconv is the group convolution, Norm is the normalization process, and \(I\) iFor the obtained context features, Q1, K1, and V1 are the query, key, and value of the spatial attention operation respectively, Q2, K2, and V2 are the query, key, and value of the pixel attention operation respectively, and I′ i is the feature obtained by spatial attention, and X i is the finally enhanced feature.

[0035] Furthermore, S2 specifically includes the following steps:

[0036] S21. Input the context features into a 1×1 convolution sequence to obtain the feature map N i ;

[0037] S22. Use the context enhancement method of S15 to perform targeted feature enhancement on the local area to obtain the context-enhanced feature map;

[0038] S22. Further process the feature information of the context-enhanced feature map through a feed-forward network, and use the attention mechanism to focus on the detailed information of the key areas;

[0039] S23. Use Softmax normalization to obtain the region-enhanced feature map.

[0040] Furthermore, the specific calculation process of S2 is as follows:

[0041] N' i ' = Feed_forward(N i ')

[0042] The specific steps of S23 are as follows:

[0043] Q, K, V = Groupconv(Conv(N i ”))

[0044]

[0045] where N i ' is the context-enhanced feature map, N i ” is the feature map after passing through the feed-forward network, Feed_forward is the feed-forward network, Conv is the 1×1 convolution, Groupconv is the group convolution, Q represents the query, K represents the key, V represents the value, t represents the learnable parameter, d k represents the dimension of K, and Z k represents the region-enhanced feature map.

[0046] Furthermore, S3 specifically includes the following steps:

[0047] S31. Weight the region-enhanced feature map and the semantic feature map, that is:

[0048]

[0049] Among them, represents pixel-by-pixel multiplication, and Z k is the residual feature after preliminary semantic guidance, and M k is the input multi-frame semantic feature map, and Z k is the feature map after passing through the region enhancement module;

[0050] S32. Integrate the semantic information in the residual feature after preliminary semantic guidance with the feature extraction residual feature, that is:

[0051]

[0052] Among them, represents pixel-by-pixel multiplication, and K i is the original multi-frame input feature map, Conv is a 1×1 convolution, and Z g is the feature map after performing 1×1 convolution, and Z g,k is the output residual feature after re-semantic guidance;

[0053] S33. Fuse the intermediate global enhancement residual feature of the output residual feature after re-semantic guidance with the feature after performing 1×1 convolution on the original input frame to obtain a semantic-guided region enhancement feature map, that is:

[0054]

[0055] S4. Fuse the semantic-guided region enhancement feature map with the globally enhanced feature map Y obtained from the main branch to complement each other and obtain the intermediate frame output effect J, that is:

[0056] J = Z′ g + Y.

[0057] The present invention also provides a storage medium, and the storage medium includes a stored program. Among them, when the program runs, it executes any one of the above-mentioned semantic-guided underwater video enhancement methods.

[0058] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs through the computer program to execute any one of the above-mentioned semantic-guided underwater video enhancement methods.

[0059] Compared with the prior art, the present invention has the following advantages:

[0060] The global feature enhancement method provided by the present invention mainly uses an encoder-decoder architecture. This method uses grouped spatial shifting to make full use of multi-frame temporal information for spatial alignment of adjacent multi-frame features, and fully learns context information to suppress irrelevant noise, enhancing the robustness of feature representation and achieving the effect of global feature enhancement.

[0061] The regional enhancement method provided by the present invention uses the context information and attention mechanism to effectively enhance the local regions of the video by retaining pixel-level dependency information and spatial details, so as to effectively supplement the deficiencies in global feature learning.

[0062] The semantic guidance method provided by the present invention may ignore the influence brought by motion during the enhancement process, which may lead to flickering problems in the enhanced video sequence. Therefore, multi-frame semantic maps are used to accurately identify the key regions of each video frame, so as to more accurately adjust the regional details of the feature map, thereby reducing the instability of the video sequence during the enhancement process. By subjective perception and analyzing the performance results of various indicators such as PSNR, SSIM, MSE, and CDC on the enhanced images, it shows that the present invention has important significance in underwater video enhancement. Brief Description of the Drawings

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a flowchart of the method of the present invention.

[0065] Figure 2 It is a comparison chart of the enhancement effects of the embodiments of the present invention. Detailed Embodiments

[0066] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0067] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0068] As Figure 1 shown, the present invention provides a semantic-guided underwater video enhancement method, which specifically includes the following steps:

[0069] S1. Perform global feature enhancement on multiple frames of video to obtain a globally enhanced feature map;

[0070] First, taking multiple frames of video as input (taking three frames as an example), the video frames are divided into three different scales through an encoder. Then, the video feature of the i-th frame with the same scale (taking K i as an example) is evenly divided along the channel dimension to obtain 8 sub-feature sets where m = 1, 2,..., 8. Subsequently, a pixel-based spatial offset operation is performed on each sub-feature set K i,m to obtain a spatially offset feature set K' i,m . After the offset operation, all the offset feature sets are re-stitched along the channel dimension to form the feature set K' i of the i-th frame after spatial offset processing. The specific process is as follows:

[0071]

[0072] K i ' = Concat(K i ' ,1 , K i ' ,2 ,..., K i ' ,8 )

[0073] where, V i is the original input frame, Encoder() is for performing multi-scale encoding operation, Split() represents the process of splitting features along the channel dimension, Offest() is the process of performing spatial offset operation on sub-feature sets, Concat() represents the process of stitching features along the channel dimension, and K'i The feature frame is obtained by grouped spatial shift.

[0074] Next, through the context enhancement method, non-critical noise is effectively filtered out, and the spatio-temporal correlation in the video is captured and utilized more accurately. The main process is as follows: First, use a 3×3 convolutional layer to map K' i to a single-channel mask to calculate the attention mask to capture the context information in the input features. This mask is then normalized by the softmax function. Use an upsampling to convert the input feature map into a matrix, and then multiply it by the attention mask to calculate the context feature I i , and then further process the context information. To accurately capture and utilize the spatio-temporal correlation in the video, spatial attention and pixel attention mechanisms are used to implicitly capture the inter-frame temporal information to obtain the feature I'' i , so as to utilize the context information to enhance the feature representation of specific pixels and channels. Finally, add it to I i to obtain the final feature effect The specific process is as follows:

[0075] I i = Upsample(Softmax(Conv(K i ')))·Conv(K i ')

[0076] Q1, K1, V1 = Groupconv(Conv(I i ))

[0077]

[0078] Q2, K2, V2 = Groupconv(Conv(I i '))

[0079] X i = Conv(softmax(Norm(Q2)·Norm(K2) T ))·t·V2)+I i

[0080] Among them, Conv() is a 3×3 convolution, soft max() is softmax normalization processing, Upsample() is an upsampling operation, Groupconv() is a group convolution, Norm() is normalization processing, I i is the obtained context feature, Q1, K1, V1 are the query, key, and value of the spatial attention operation respectively, Q2, K2, V2 are the query, key, and value of the pixel attention operation respectively, I' i is the feature obtained by spatial attention, Xi is the finally enhanced feature.

[0081] Finally, the feature maps of different scales are merged through the upsampling operation of the decoder to form the finally globally enhanced feature map Y.

[0082] S2. Regionally enhance the context features of multiple frames of video to obtain the regionally enhanced feature map;

[0083] In the semantic region branch, through the region enhancement module, perform reinforcement learning on local features to effectively supplement the deficiencies in global feature learning. First, I i enter a 1×1 convolution sequence to obtain N i , to better extract and process features. Subsequently, through the fine context feature enhancement module, the specific process is as shown above, to better process context information, so as to perform targeted feature enhancement on local regions while maintaining a global perspective, obtaining the feature N′ i , then further process the feature information through a feed-forward network. Then, in order to better enhance the regional features, the attention mechanism is adopted to focus on the detailed information of key regions, and then Softmax normalization is used to obtain the final regionally enhanced feature map Z k . The specific calculation process is as follows:

[0084] N' i ' = Feed_forward(N i ')

[0085] Q, K, V = Groupconv(Conv(N i ”))

[0086]

[0087] Among them, N′ i is the context-enhanced feature map, N″ i is the feature map after passing through the feed-forward network, Feed_forward() is the feed-forward network, Conv() is a 1×1 convolution, Groupconv() is a group convolution, Q represents the query, K represents the key, V represents the value, t represents the learnable parameter, d k represents the dimension of K, softmax() is the softmax normalization process, and Z k represents the regionally enhanced feature map.

[0088] S3. Semantically guide the regionally enhanced feature map to obtain the output effect of the intermediate frame.

[0089] The feature map Z k that has undergone regional enhancement feature enhancement is passed through the semantic feature map Mk The guidance of k accurately identifies the key areas of each video frame and further enhances them. First, in order to accurately identify the key areas of Z k and alleviate the instability of video movement, we will Z k weighted with the semantic feature map M k , that is:

[0090]

[0091] Among them, represents pixel-by-pixel multiplication, Z k is the residual feature after preliminary semantic guidance, M k is the input multi-frame semantic feature map, Z k is the feature map after passing through the region enhancement module.

[0092] To ensure the accuracy of semantic guidance, we integrate the semantic information in Z' k with the feature extraction residual feature Z g to more accurately adjust the regional details of the feature map, that is:

[0093] Z g = Conv(K i )

[0094]

[0095] Among them, represents pixel-by-pixel multiplication, K i is the original multi-frame input feature map, Conv() is a 1×1 convolution, Z g is the feature map after 1×1 convolution, Z g,k is the residual feature output after re-semantic guidance.

[0096] Finally, combine the intermediate global enhancement residual feature guided by semantic attention with the original feature to fuse, and obtain the final semantic-guided region enhancement feature map Z' g , that is:

[0097]

[0098] S4. Fuse Z' g with the globally enhanced feature map Y obtained from the main branch, complement each other, and obtain the final intermediate frame output effect J, that is:

[0099] J = Z′ g + Y

[0100] Embodiment

[0101] Such as Figure 2As shown in the figure, the present invention provides a comparison chart of the enhancement effects of other algorithms on underwater video frames. It can be seen from the experimental effect diagrams that all six algorithms enhance the underwater video frames to a certain extent. However, the result diagram of the MLLE algorithm has a poor effect on color drift correction. In the UIE algorithm, the problem of color bias is not well solved, and there is still an obvious blue-green bias. Although the UVENet and UVE-Net algorithms solve the problem of color bias, the restored video frames are too dark. The underwater video frames processed by the algorithm of the present invention better solve the problem of color shift compared with other algorithms, and at the same time better retain the detail information and structural information of the background area of the restored image, and are closer to the color of GT.

[0102] In this embodiment, the experimental results of different algorithms are compared from two objective indicators of PSNR and SSIM; PSNR represents the ratio of the maximum possible power of a signal to the power of the destructive noise that affects its representation accuracy. SSIM describes the similarity between two pictures from three aspects: brightness, contrast and structure. The present invention adopts a dual-branch structure, including a multi-scale global feature enhancement branch and a semantic branch, which complement each other to ensure the stability and continuity of the enhanced video. Therefore, the present invention has a greater improvement in the PSNR and SSIM indicators of the original video, and is superior to other underwater image restoration algorithms and existing underwater video enhancement algorithms.

[0103] Table 1 PSNR comparison table of the processing results of the model of the present invention and other advanced algorithms

[0104]

[0105] Table 2 SSIM comparison table of the processing results of the model of the present invention and other advanced algorithms

[0106]

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An underwater video enhancement method based on semantic guidance, characterized in that, It includes the following steps: S1. Perform global feature enhancement on multiple frames of video to obtain a globally enhanced feature map; S2. Perform regional enhancement on the context features of multiple frames of video to obtain a regionally enhanced feature map; S3. Perform semantic guidance on the regionally enhanced feature map to obtain a semantically guided regionally enhanced feature map; S4. Fuse the semantically guided regionally enhanced feature map with the globally enhanced feature map to obtain the output effect of the intermediate frame.

2. The underwater video enhancement method based on semantic guidance according to claim 1, wherein S1 specifically includes the following steps: S11. Divide multiple frames of video into several different scales; S12. Uniformly divide the feature of the i-th frame of video with the same scale along the channel dimension to obtain a sub-feature set; S13. Perform a pixel-based spatial offset operation on each sub-feature set to obtain a spatially offset feature set; S14. Re-stitich the spatially offset feature set along the channel dimension to form a feature set processed by spatial offset; S15. Filter out non-critical noise in the feature set processed by spatial offset through a context enhancement method to obtain feature maps of different scales; S16. Merge the feature maps of different scales through the upsampling operation of the decoder to form a globally enhanced feature map.

3. The underwater video enhancement method based on semantic guidance according to claim 2, wherein The specific formula of S11 is as follows: The specific formula of S13 is as follows: The specific formula of S14 is as follows: K′ i = Concat(K′ i,1 , K′ i,2 ,..., K′ i,8 ) Among them, V i is the original input frame, Encoder performs multi-scale encoding operations, Split represents the process of splitting features along the channel dimension, Offest is the process of performing spatial offset operations on the sub-feature sets, Concat represents the process of concatenating features along the channel dimension, and K' i is the feature frame obtained by grouped spatial shifting.

4. The underwater video enhancement method based on semantic guidance according to claim 3, wherein, S15 specifically includes the following steps: S151. Use a 3×3 convolutional layer to map the spatially offset feature set to a single-channel mask, calculate the attention mask to capture the context information in the input features. The attention mask is then normalized by the softmax function. Use an upsampling to convert the input feature map into a matrix, and then multiply it by the attention mask to calculate the context feature I i , and the specific formula is as follows: I i = Upsample(Soft max(Conv(K′ i )))·Conv(K′ i ) S152. Use the spatial attention and pixel attention mechanisms to implicitly capture the inter-frame temporal information and obtain the feature I". i , and the specific formula is as follows: Q1, K1, V1 = Groupconv(Conv(I i )) S153. Add "Feature I" i to I i to obtain the feature effect The specific formula is as follows: Q2, K2, V2 = Groupconv(Conv(I′ i )) X i = Conv(softmax(Norm(Q2)·Norm(K2) T )·t·V2)+I i Among them, Conv is a 3×3 convolution, softmax is softmax normalization, Upsample is an upsampling operation, Groupconv is a group convolution, Norm is normalization, and I i is the obtained context feature, Q1, K1, and V1 are the query, key, and value of the spatial attention operation respectively, Q2, K2, and V2 are the query, key, and value of the pixel attention operation respectively, and I i ′ is the feature obtained by spatial attention, and X i is the finally enhanced feature.

5. The underwater video enhancement method based on semantic guidance according to claim 1, characterized in that S2 specifically includes the following steps: S21. Input the context features into a 1×1 convolution sequence to obtain the feature map N i ; S22. Use the context enhancement method of S15 to perform targeted feature enhancement on the local area to obtain a context-enhanced feature map; S22. Further process the feature information of the context-enhanced feature map through a feed-forward network, and use the attention mechanism to focus on the detailed information of the key area; S23. Use Softmax normalization to obtain a regionally enhanced feature map.

6. The underwater video enhancement method based on semantic guidance according to claim 5, characterized in that The specific calculation process of S2 is as follows: N” i = Feed_forward(N' i ) The specific steps of S23 are as follows: Q, K, V = Groupconv(Conv(N” i )) Among them, N' i is the feature map after context enhancement, N'' i is the feature map after passing through the feed-forward network, Feed_forward is the feed-forward network, Conv is a 1×1 convolution, Groupconv is a grouped convolution, Q represents the query, K represents the key, V represents the value, t represents the learnable parameter, d k represents the dimension of K, Z k represents the feature map after region enhancement.

7. The semantic-guided underwater video enhancement method according to claim 1, wherein S3 specifically includes the following steps: S31. Weight the regionally enhanced feature map and the semantic feature map, that is: Among them, represents per-pixel multiplication, and Z' k is the residual feature after preliminary semantic guidance, M k is the input multi-frame semantic feature map, and Z k is the feature map after passing through the region enhancement module; S32. Integrate the semantic information in the preliminarily semantically guided residual feature with the feature extraction residual feature, that is: Z g = Conv(K i ) Among them, represents pixel-by-pixel multiplication, K i is the original multi-frame input feature map, Conv is a 1×1 convolution, Z g is the feature map after the 1×1 convolution, Z g,k is the residual feature output after semantic guidance again; S33. Fuse the output residual feature after re-semantic guidance with the feature after 1×1 convolution of the original input frame to obtain the final semantically guided regionally enhanced feature map, that is: S4. Fuse the semantically guided regionally enhanced feature map with the globally enhanced feature map Y obtained from the main branch, and complement each other to obtain the output effect J of the intermediate frame, that is: J = Z' g + Y.

8. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it executes the semantically guided underwater video enhancement method according to any one of claims 1 to 7.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor runs through the computer program to execute the semantically guided underwater video enhancement method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Semantic-driven frequency consistency underwater image enhancement method

    CN121961897A